The RL environments market map for 2026
Most lists of RL environment vendors rank companies against each other as if they were competing for the same order. They are not. A frontier lab buying ten thousand graded coding tasks and an insurer trying to get one claims workflow to run reliably are shopping in different segments, and the same nine names show up on both lists.
This page is a map rather than a ranking. It names the three axes the category actually splits along, places the vendors on each, and links a source for every factual claim. If you want an opinionated evaluation instead - who is better than whom, and for what - that is our best RL environment companies guide, and the two pages are meant to be read together.
Disclosure: this map is written by Pebble, which appears on it. We sell one of the segments below and say so plainly. Corrections to [email protected].
How is the RL environments market segmented?
Three axes, and every vendor's strategy is a position on all three.
- Buyer. Frontier labs buying training data at volume, or enterprises trying to make their own agents work.
- Organizing unit. Environments built around a modality (coding, browser, computer use, spreadsheets) or around a process (claims adjudication, order to cash, onboarding).
- Openness. An open catalog and library anyone can build on, or a closed engagement with a contract and an NDA.
Two of those axes correlate hard. Lab buyers mostly want modality coverage, because a lab is trying to make a general model better at a broad skill. Enterprise buyers mostly want a process, because an enterprise is trying to get one workflow to run. That correlation is the single most useful thing on this map: it explains why the same vendor cannot serve both well, and why the category has produced two almost separate industries under one name.
Who is the buyer, frontier labs or enterprises?
Labs are still the money, and enterprises are the growth story. Epoch AI's January 2026 survey puts typical lab contracts at six to seven figures per quarter, and TechCrunch reported in September 2025 that Anthropic leadership had discussed spending over $1 billion on environments in the following year. That is a small number of buyers writing very large checks.
The lab-facing side is dominated by the human-data incumbents. Scale reports that nearly half of its new data training projects now involve RL environments, in the post that named the category. Surge, reportedly at $1.2 billion in revenue, stood up a dedicated internal environments organization. Mercor, valued at $10 billion, bought its way in by acquiring Sepal AI in February 2026. Mechanize sits in the same segment at the opposite end of the volume curve, founded in April 2025 by three ex-Epoch AI researchers and building a small number of very hard coding environments.
The enterprise-facing side is younger and thinner. Veris raised an $8.5 million seed in June 2025 to sell agent simulation to companies rather than labs. Fleet sells to both, and Sacra reports annualized revenue growing from about $1 million to over $60 million in under a year, which is the clearest evidence that enterprise demand is real. Pebble is enterprise-only.
Is the environment organized by modality or by process?
Modality is the default and process is the white space. A modality vendor sells you a browser gym, a spreadsheet gym, or a terminal gym, and the task inventory inside it is broad and shallow. A process vendor sells you one workflow mapped end to end, and the inventory is narrow and deep.
Modality is the right unit for a lab: general capability comes from breadth. It is the wrong unit for a company that needs its own claims process to work, because nobody's claims process is a modality. The process-organized segment currently holds Pebble, Veris, and the services firms in the second tier below, and it is the segment where environment engineering skill matters most, since the step map has to be negotiated with the people who own the process.
Fleet is the interesting boundary case. Its unit is the application replica rather than the modality or the process, which makes it a supplier to both sides: labs buy the replica to train computer-use agents, enterprises buy it because their work happens in Salesforce.
Is the stack open or closed?
One vendor is genuinely open, a few publish tooling, and the rest are contracts. Prime Intellect runs the Environments Hub, a public catalog built on its open verifiers library, raised $130 million at a $1 billion valuation in July 2026, and trains its own models on the same stack it publishes, including INTELLECT-3. It is the only place on this map where you can read the environments before you buy anything.
HUD is open in the tooling sense: a 15-person Y Combinator W25 company whose hud-python SDK and public benchmarks including OSWorld-Verified are on GitHub. Fleet publishes its Harbor evaluation and RL tooling. Mechanize publishes candid technical work such as GBA Eval without publishing the environments themselves. Scale, Surge, Mercor, Veris, and Pebble are closed engagements.
Openness is not a virtue rating. It is a statement about who does the work. Open means you build and maintain the environment. Closed means someone else does, and you pay for their hours.
Which vendors sit where on the map?
| Vendor | Primary buyer | Organizing unit | Openness | Anchor fact |
|---|---|---|---|---|
| Scale AI | Labs | Modality, plus human data | Closed | Nearly half of new data training projects involve RL environments (Scale) |
| Surge AI | Labs | Expert human judgment | Closed | Dedicated internal environments org, ~$1.2B revenue (TechCrunch) |
| Mercor | Labs | Domain expertise | Closed | Acquired Sepal AI, February 2026 (Orrick) |
| Prime Intellect | Researchers and labs | Community catalog | Open | Environments Hub on the open verifiers library (Prime Intellect) |
| Mechanize | Labs | Hard coding tasks | Closed, publishes research | Founded April 2025 by ex-Epoch AI researchers (the-decoder) |
| Fleet | Both | Application replica | Tooling open | ~$1M to $60M+ annualized in under a year (Sacra) |
| HUD | Agent teams | Benchmarks and SDK | Tooling open | 15-person YC W25 company (YC) |
| Veris AI | Enterprises | Their tools, simulated | Closed | $8.5M seed, June 2025 (BusinessWire) |
| Pebble | Enterprises | One business process | Closed | Per-step checks and customer-owned run records (Pebble) |
Who else is in the RL environments market?
A second tier of services and evaluation firms that added environments to an existing business, all of which publish their offering openly. They matter because they are usually the incumbent already inside the enterprise account.
- Centific sells RL environments as a service across financial services, healthcare, supply chain, and public sector verticals, with licensed practitioners authoring the rubrics, and announced the offering as a launch. It is the closest thing on this map to a direct competitor for the process-organized, enterprise-buyer segment.
- Turing positions environments as the enterprise layer between fine-tuning and deployment, describing each environment as "a self-contained digital twin" of an enterprise system, alongside its datasets, benchmarks, and AI services business.
- Toloka approaches environments from the data-platform side and publishes an explainer on how agents learn to act inside them.
- Patronus AI comes at it from evaluation, describing its platform as one that turns evaluation pipelines into practical RL environments.
Expect more of these. Any firm that already sells annotation, evaluation, or systems integration into an enterprise can add an environments page, and most will.
How fast will the RL environments market consolidate?
Faster than the vendor count suggests, if the investors are right. Wing Venture Capital's Chris Zeoli argued in Who Will Win the RL Environment Market (January 27, 2026) that "roughly 20 seed- to Series A-stage companies" will narrow to "three to five market leaders, with one to two dominant platforms," and that lab spend will concentrate on the few teams that push the frontier rather than on environment volume. The same piece counts "roughly three $1B+ revenue players (Scale, Mercor, Surge)" already established in data labeling.
If that holds, the lab-facing segment consolidates first and hardest, because it has the fewest buyers. The enterprise-facing segment has thousands of potential buyers and no dominant vendor, which is a different market shape and a slower one.
What does this map not tell you?
Which vendor to pick, and what it will cost. A map places vendors, it does not rank them. For an opinionated read on who is better for what, including where we are the wrong choice, see the best RL environment companies in 2026. For the numbers, what RL environments cost collects the public per-task, per-replica, and per-contract benchmarks. If you are specifically leaving the incumbent, Scale AI alternatives for RL environments covers that path.
And if you are in the segment this map calls process-organized and enterprise-buying - one workflow, a check on every step, a dataset and model you own at the end - that is what Pebble does. Write to [email protected] for a scoping call.