guide
Updated July 2026

Scale AI alternatives for RL environments

Scale AI helped define the RL environments category. Its own blog post reports that nearly half of Scale's new data training projects now involve reinforcement learning environments, built on the biggest human-data operation in the industry. If you are a frontier lab with an existing Scale contract, that gravity is real. This page is for everyone asking what else exists, and it is written by Pebble, one of the alternatives profiled below, so read our entry with that in mind.

Why do teams look for Scale AI alternatives?

Three reasons come up repeatedly: vendor neutrality after the Meta deal, a service model built for labs rather than enterprises, and per-task human-data pricing. On the first: Meta took a 49% stake in Scale for $14.3 billion in June 2025 and hired founder Alexandr Wang to lead its superintelligence effort. Within days, Google, reportedly Scale's largest customer at roughly $200 million a year, planned to cut ties, and OpenAI confirmed it was winding down its Scale work (OpenAI said its pullback predated the deal). If your training priorities would interest Meta, that ownership structure is a legitimate question to ask any vendor, and Scale's competitors will make sure you ask it.

On the second and third: Scale's environments offering grew out of frontier-lab data contracts, and the economics reflect that. Epoch AI's RL environments FAQ documents typical contracts of six to seven figures per quarter and per-task costs of $200 to $2,000, with complex app replicas running around $300,000. Those numbers make sense when you are buying millions of graded tasks to train a frontier model. They make much less sense when you are an enterprise that wants one workflow to run reliably. Different problem, different vendor. Our pricing page shows what process-scoped engagements look like instead.

What should you look for in an alternative?

Four things separate these vendors more than their marketing does: verification depth, dataset ownership, hosting model, and whether they sell by process or by modality.

Verification depth. An environment is only as good as its reward signal, and reward hacking is the top problem practitioners report. Ask exactly how each step gets graded: computed checks, classifiers, human judgment, or LLM judges, and against what ground truth. Our primer on verifiable rewards covers the taxonomy.

Dataset ownership. Every run through an environment produces training data. Ask who owns it. Epoch notes that exclusive deals command 4-5x higher pricing, which tells you how valuable vendors think this data is. If the run records stay with the vendor, you are paying to build someone else's asset.

Hosting model. Can the environment run inside your infrastructure, or only on the vendor's platform? For anything touching real invoices, disputes, or safety reports, this decides the deal.

Process vs modality. Most vendors sell a modality: coding environments, browser environments, computer-use gyms. A few sell an outcome: this specific workflow, running, verified. Neither is wrong, but buying a modality when you need an outcome leaves you doing the integration yourself.

The alternatives

Prime Intellect

Prime Intellect is the open-source alternative: a public Environments Hub of community-built environments on its open verifiers library, backed by a $130 million raise at a $1 billion valuation in July 2026 with investors including Nvidia's NVentures. It proved the stack on its own INTELLECT-3 model. Best for: teams with ML engineers who want to build and train on open infrastructure with zero vendor lock-in.

HUD

HUD is a 15-person YC W25 company with the best developer experience in the category: the open-source hud-python SDK, plus owned benchmarks OSWorld-Verified and SheetBench-50. Best for: teams evaluating and training computer-use agents who want benchmarks and an SDK, not a managed service.

Fleet

Fleet builds high-fidelity replicas of enterprise apps (Salesforce, Excel, browser and desktop workflows) and, per Sacra, grew annualized revenue from roughly $1 million to over $60 million in under a year while discussing a raise near a $750 million valuation. It ships a Python SDK and open-source tooling via its GitHub org. Best for: training agents that must drive real enterprise software, where replica fidelity matters most.

Mechanize

Mechanize sells a small number of very hard, very well-built environments for frontier coding agents, founded in April 2025 by ex-Epoch AI researchers and known for candid technical publishing like GBA Eval. Best for: frontier labs buying depth over breadth in coding environments.

Surge AI

Surge is the closest like-for-like Scale replacement: reportedly $1.2 billion in revenue from labs including OpenAI, Google, Anthropic, and Meta, with a new dedicated RL environments organization per the same report, and no equivalent ownership entanglement. Best for: frontier labs that need expert human data and environments at Scale-class volume from a neutral vendor.

Mercor

Mercor, valued at $10 billion, bought its way into environments by acquiring Sepal AI in February 2026 and is pitching domain-specific environments for coding, healthcare, and law. Best for: labs whose environments need scarce domain experts more than they need software.

Veris AI

Veris raised an $8.5 million seed (Decibel and Acrew, June 2025) to let enterprises train and test their own agents in high-fidelity simulations before production. Best for: enterprises that want a simulation platform for agents their own team is building.

Pebble

Pebble is us: a small applied lab in San Francisco, founded mid-2025, a few customers, founder-built environments. The model is different from everything above. You bring a business process (invoice reconciliation, safety-report triage, dispute responses). We map it into steps and define a verification check for each one: computed, classified, or judged against operator-annotated examples. Failed runs stop and resume at the failed step. Every run record accumulates into a dataset you own, which you can use to post-train an open-source model you also own, and the environment ships as software you host or we host. On the four criteria above: verification is per-step and explicit, the dataset is yours by default, hosting is your choice, and we sell by process, not modality. Best for: companies that want one business process run by AI with an audit trail, and that want to own the resulting data and model. Not for: labs that need hundreds of environments or anyone who wants a self-serve platform. It starts with a scoping call ([email protected]) because we cannot define checks for a process we have not seen.

When is Scale AI still the right choice?

Scale remains the right choice when you are buying frontier-scale human data pipelines and environments as one integrated contract, and the Meta relationship does not concern you. Nobody on this list except Surge can match its combined data-operations capacity, and Scale's environments inherit real infrastructure from years of lab work. Be honest about what you are: if you are a lab training frontier models, Scale and Surge belong on your shortlist. If you are an enterprise that wants a workflow to run with verified steps, per-task data pricing was never built for you, and the alternatives above, ours included, exist precisely because of that gap.

For fuller profiles of every vendor here, see our guide to the best RL environment companies. To see where each of them sits on the buyer, modality-or-process, and open-or-closed axes, see the RL environments market map.