guide
Updated July 2026

The 9 best RL environment companies in 2026

RL environment companies build the simulated workspaces where AI agents learn multi-step tasks by trial and error, with an automated grader scoring every attempt. The category went from obscure to overheated in about a year: TechCrunch reported in September 2025 that Anthropic leadership had discussed spending over $1 billion on RL environments in the coming year, and Epoch AI's January 2026 FAQ put typical vendor contracts at six to seven figures per quarter.

That money has pulled in three kinds of vendors: data-labeling incumbents adding environments to existing lab contracts, environment-native startups founded in 2025, and open-source platforms. They are not interchangeable, and most lists in this space pretend they are. Read this one as an evaluation rather than a ranking: the right pick depends entirely on whether you are a frontier lab buying training data or an enterprise trying to get an agent to run a workflow. For the orientation view first - the axes the category splits along and which segment each vendor occupies - start with our RL environments market map. If you are new to the vocabulary (rollouts, verifiers, reward hacking), our glossary covers it.

How we chose

Three criteria: a stated focus you can verify, evidence of real customers or shipped product, and a public technical footprint (SDKs, benchmarks, papers, or published environments). Every factual claim below links to a source.

Disclosure: this list is written by Pebble, and Pebble appears on it. We have tried to describe every competitor the way their own customers would, and we say plainly who each vendor is better than us for. If we got a fact wrong, email [email protected] and we will fix it.

Scale AI

Scale is the data incumbent that now reports nearly half of its new data training projects involve RL environments, per its own category-defining blog post. Scale builds simulated tool-use, computer-use, and coding environments for frontier model developers, layered on top of the largest human-data operation in the industry. The complication is 2025: Meta took a 49% stake for $14.3 billion and hired founder Alexandr Wang, after which Google reportedly planned to cut ties and OpenAI wound down its work with Scale. Best for: large organizations that need human-data pipelines and environments from one vendor and are comfortable with the Meta relationship. If you are not, see our Scale AI alternatives guide.

Surge AI

Surge is the quiet giant of human data, reportedly generating $1.2 billion in revenue from labs including OpenAI, Google, Anthropic, and Meta, and it recently spun up a dedicated internal organization for RL environments per the same TechCrunch report. Surge was bootstrapped and profitable with a small full-time team managing tens of thousands of expert contractors. Its environments work inherits that machine: expert humans defining tasks and grading edge cases at scale. Best for: frontier labs that need expert human judgment folded into environments and evals, at volumes only a handful of vendors can staff.

Mercor

Mercor is the expert-network play, valued at $10 billion, which entered environments directly by acquiring Sepal AI in February 2026. Sepal, a YC-backed data-research company, built training data, expert-graded benchmarks, and RL environments for frontier labs, and Mercor is now pitching domain-specific RL environments for coding, healthcare, and law. The bet is that environment quality bottlenecks on domain experts (doctors, lawyers, accountants) rather than on software. Best for: labs that need deep domain expertise embedded in environment design, especially in regulated fields.

Prime Intellect

Prime Intellect is the open-source pole of the category: its Environments Hub is a public catalog of community-built environments on its open verifiers library, and the company raised $130 million at a $1 billion valuation in July 2026 with backers including Nvidia's NVentures and Intel Capital. It also trains its own models: INTELLECT-3, a 100B+ parameter mixture-of-experts model, was trained with large-scale RL using the same open stack it sells. Best for: researchers and teams that want to build, publish, and train on environments themselves without vendor lock-in. It is a platform, not a service, so plan on doing the work.

Mechanize

Mechanize builds a small number of high-fidelity RL environments and evals for frontier coding agents and sells them to leading labs, founded in April 2025 by ex-Epoch AI researchers Matthew Barnett, Tamay Besiroglu, and Ege Erdil. The company is explicit about its long-term goal of automating the economy, pays engineers $500,000 to build environments, and publishes unusually candid technical work like GBA Eval, which asks whether a coding agent can write a working Game Boy Advance emulator in 24 hours. Best for: frontier labs buying a few extremely hard, extremely well-built coding environments rather than a large catalog.

Fleet

Fleet builds high-fidelity replicas ("gyms") of enterprise software such as Salesforce, Excel, browsers, and desktop workflows so labs and enterprises can train computer-use agents against realistic apps, and Sacra reports it has been in talks to raise at a roughly $750 million valuation with annualized revenue growing from about $1 million to over $60 million in under a year. It ships a Python SDK and the open-source Harbor evaluation and RL tooling. Best for: teams training agents that must operate real enterprise applications, where fidelity of the app replica is the whole game.

HUD

HUD is the evals-and-benchmarks specialist: a 15-person Y Combinator W25 company whose open-source hud-python SDK is pitched as "RL environments + evals for AI agents," with owned public benchmarks including OSWorld-Verified (where the team fixed 300+ issues in the original OSWorld) and SheetBench-50. Strong documentation, a real resources hub, and a define-once-train-anywhere developer experience. Credit where due: HUD's own listicle ranks first for this keyword, and it earned that. Best for: teams evaluating and training computer-use agents who want benchmarks, leaderboards, and an SDK rather than a done-for-you service.

Veris AI

Veris builds simulated training environments for enterprise agent workflows, letting companies train and test agents against high-fidelity simulations of their own tools before production, and it raised an $8.5 million seed led by Decibel and Acrew in June 2025. Where most of this list sells to labs, Veris sells simulation to enterprises. Best for: enterprises that want a platform to simulate and harden their own agents, with early customers in financial services and manufacturing per the same announcement.

Pebble

Pebble (that's us) is a small applied lab in San Francisco, founded in mid-2025, with a few customers and founder-built environments. We sell by business process, not by modality: instead of a coding gym or a browser sandbox, you bring us a process like invoice reconciliation, safety-report triage, or dispute responses. We map the process into steps and define a verification check for each step, computed, classified, or judged against operator-annotated examples. Failed runs stop and resume at the failed step, so a bad run is a caught step, not a silent wrong answer. Run records accumulate into a dataset the customer owns and can use to post-train an open-source model they also own. Engagements ladder from RUN to PROVE to OWN, and everything starts with a scoping call because we have to see the process before we can verify it. There is no public price list, but we published what this class of work costs across the market. Best for: companies that want AI running a specific business process with an audit trail, not a platform to build on. We are the youngest company on this list and the wrong choice if you need a thousand environments by Friday.

Which RL environment company should you pick?

Pick by what you are actually buying, because these nine companies sell five different things. Open-source experimentation and community environments: Prime Intellect. Frontier-scale coding environments: Mechanize. Faithful replicas of enterprise apps: Fleet. Evals, benchmarks, and an SDK for computer-use agents: HUD. Human data and expert grading at lab scale: Surge, Mercor, or Scale, in roughly that order of vendor neutrality after the Meta deal. Simulating your own enterprise agents: Veris.

And if what you want is one business process run by AI with a verification check on every step and a dataset you own at the end, that is the narrow thing Pebble does. Email [email protected] to scope it. If your process does not fit our check taxonomy, we will tell you on the call and point you at whichever company above fits better.