Market Research RL Environments AI Data Economy September 7, 2026

Build the Environments,
Not the Workforce

Market entry assessment, September 2026: six candidate categories, one recommendation.

By Mohinish Shaikh  and  Pragadeesh VS

Published: September 7, 2026  ·  Read time: ~14 minutes

The thesis

The commodity labeling tier is being automated toward zero. The layer above it, reinforcement-learning environments with working verifiers, is the only high-value category where the deliverable is software, the buyer pays per artifact rather than per hour, and two engineers in India can compete on craft instead of headcount.

Recommended entry: RL environments with verifiers. Scope: which of six high-value data categories a two-person, engineering-led startup with an India-based workforce should enter. All figures dated and sourced. Several are estimates or single-source; see the reliability notes at the end.

What the market looks like right now

Demand is enormous and highly concentrated, and it has split cleanly in two: the bottom is collapsing in price while the top is starved of supply.

~$1B

Annual human-data spend per frontier lab (Time, 2025 investigation)

>75%

Of category revenue held by Scale, Surge, Mercor and Handshake (Deedy Das venture map, via Pebblous)

60–90%

Of routine labels now correctly pre-filled by foundation models, compressing that tier (Shaip)

60–80%

Cost advantage of India delivery over US teams for equivalent work (Precise BPO)

Synthetic data is a complement rather than a replacement: cited optimal mixes sit around 60–70% human to 30–40% synthetic, and pure-synthetic training risks model collapse. Machines now handle volume; humans set policy, audit edges, and build the environments the machines train inside. That is where the money moved.

The bench

Six categories measured on the axes that decide whether two engineers can actually win: how badly buyers want it, whether a small team can hold a position, how much capital it eats, how fast it pays, and whether it can be delivered from India without a jurisdictional blocker.

1

RL environments with verifiers

Recommended

Software deliverable, per-artifact pricing, non-personal data, engineering is the moat. Enter here.

Demand
Defensible
Capital-light
Fast revenue
India-ready
2

Computer-use & agentic trajectories

Strong second

Natural bolt-on to environments. Build synthetic capture; keep real-user personal data out of scope.

Demand
Defensible
Capital-light
Fast revenue
India-ready
3

Expert data generation

Big pool, weak position

Biggest revenue pool, but the moat is a 30,000-expert network you cannot rebuild. Viable only as a narrow vertical.

Demand
Defensible
Capital-light
Fast revenue
India-ready
4

Failure-mode data & red-teaming

Skip as primary

Automating fast, and the premium tier is gated to US persons. Skip as a primary business.

Demand
Defensible
Capital-light
Fast revenue
India-ready
5

Indic & multilingual data

Home turf, low ceiling

Home-turf advantage, but the buyer is a thin domestic budget reached through slow government tenders.

Demand
Defensible
Capital-light
Fast revenue
India-ready
6

Physical AI & robotics data

Avoid

Real demand, wrong profile: capital-heavy, ops-heavy, geographically anchored, slow. Avoid.

Demand
Defensible
Capital-light
Fast revenue
India-ready

Each gauge is five segments; more filled is better. Scores are judgements drawn from the evidence in the dossiers below, not published metrics.

The dossiers

What actually ships, who buys it, what it costs, who already owns the position, and what would stop you delivering it from India.

1. RL environments with verifiers Recommended

What ships

An action space plus surrounding state (file systems, simulated apps, environment variables), usually delivered as a Docker container, plus tasks (a prompt and a grader). Graders are unit tests, state inspection, programmatic checks, or LLM-judges against rubrics. A commercial example: a repo snapshot, task spec, and a verifier scoring each attempt on a continuous 0–1 reward, shipped as "the container and the reward entrypoint," plus reward-labeled agent trajectories. Coding environments usually bundle one task each; a computer-use replica (an Excel, Bloomberg or Airbnb clone) can carry hundreds.

Who buys

Frontier labs (Anthropic; OpenAI and xAI named via job postings), neolabs such as Cursor, and product partnerships (Benchling with Anthropic for biology; OpenAI with Shopify and Stripe for shopping). Environment vendors also subcontract build capacity overseas.

Price anchors

Contracts often six to seven figures per quarter; deal sizes of $300k–$500k cited by one neolab; website replicas around $20k; a complex product replica such as Slack around $300k; per-task $200–$2,000, rarely to $20k for hard software engineering. Exclusive deals cost 4–5× non-exclusive. (Epoch AI survey of 18 insiders, Jan 2026; SemiAnalysis)

Size signal

The Information reported in September 2025 that Anthropic discussed spending "at the level of $1 billion per year on training environments alone." For scale, OpenAI's R&D consumed roughly $19B of about $34B total 2025 spending, and Greg Brockman testified it would spend around $50B on compute in 2026, triple 2025.

Who's there

Specialists: Mechanize (SF, ~$9.1M raised, working with Anthropic, $500k engineer salaries), Halluminate (YC S25, narrowed to financial-services environments, ~$1.3M ARR reported 2025), Prime Intellect (open Environments Hub, 500+ community environments, Verifiers library), HUD, and others catalogued on rl-list.com. The large data vendors (Mercor, Surge, Turing, Handshake) are moving in with distribution, but it is an open question whether labeling operations translate into sound verifier engineering. The field is fragmented and early.

Barriers

Capital: negligible. Skill: high. Reward-hacking resistance takes many iterations; difficulty must be calibrated to roughly a 2–3% minimum pass rate with a smooth gradient and compositional skills. Every interviewee named the same hard problem: scaling task volume without quality collapse. That is a management problem a software-minded team can partly convert into a tooling problem, which is exactly the opening.

Automation risk

Low, and inverted: environments are themselves the machinery that generates synthetic data, so the synthetic shift increases demand here. Real risks are labs in-housing (already "substantially more in-housing") and a flood of vibe-coded clones at the low end: "a large amount of useless bad environments out there."

India delivery

Clear. Environments are non-personal software artifacts, outside the DPDP Act's personal-data scope. Commercial frontier work is open to India-based delivery, and environment vendors already hire overseas developers to replicate site UIs. A US wrapper additionally opens anything government-adjacent.

2. Computer-use & long-horizon agentic trajectories Strong second

A natural bolt-on to environments rather than a separate business: the same engineering that builds an environment also produces the recorded trajectories labs want. Build synthetic capture rather than instrumenting real users, and keep real-user personal data out of scope entirely, because that is what keeps the category clean under the DPDP Act and simple to sell across borders. Full dossier in the interactive assessment.

3. Expert data generation Big pool, weak position

What ships

Credentialed experts author original problems, worked solutions, grading rubrics, evaluations and reasoning traces: production, not labeling. Domains in demand: medicine, law, senior software engineering, advanced mathematics, quantitative finance, accounting, PhD sciences.

The model

Mercor, Surge, Handshake, Micro1, Turing and AfterQuery place vetted experts onto lab projects. Mercor runs about a 35% take rate; The Information, citing internal documents, put gross revenue at roughly $614M in H1 2026 and a $2B annualized gross run-rate by June (up from about $760M at end-2025 and $500M in September 2025), with 91% of revenue from foundation-model companies such as OpenAI and Anthropic. Sacra estimates contractors keep 60–70% of top line, implying H1 net revenue of roughly $180M–$250M. It distributes about $1.5M per day to 30,000+ contractors across 45+ countries, India its largest talent source.

Rates paid

Mercor average around $85/hr (advertised $81–$114). Tiers: $12–$25/hr entry generalist; $25–$53/hr specialized RLHF, coding and translation; $75–$200+/hr credentialed experts. Medical MDs $100–$210/hr with a top decile near $250/hr; offensive-security $200–$250/hr; law, finance and quant in between. (Contractor-reported, 2026)

Scale of rivals

Sacra estimates Handshake reached $1.1B annualized gross revenue in April 2026, up 349% year over year from about $245M, with roughly $450M net after contractor payouts; The Information put its AI-training gross ARR near $1B, up from $550M in January 2026.

Why it's a weak fit

The moat is the credentialed-expert network and the vetting engine (Mercor's APEX AI interview), not technology, and not replicable by two engineers. You would be reselling the same India expert labour the incumbents already recruit directly, with no differentiation.

India delivery

Open but undifferentiated. Only worth pursuing as a jurisdiction-specific vertical where you hold real advantage (Indian law, ICAI accounting, Indian medical boards) rather than head-on.

4. Failure-mode data & red-teaming Skip as primary

Automating fast, and the premium tier (the defence- and government-adjacent work that pays $200–$250 an expert hour) is gated to US persons. That combination closes the high end to an India-based team while the low end erodes. Skip it as a primary business; revisit only as an add-on to environment work. Full dossier in the interactive assessment.

5. Indic & multilingual data Home turf, low ceiling

Genuine home-turf advantage, but the buyer is a thin domestic budget reached through slow government tenders (IndiaAI Mission, CPP portal empanelment, BharatGen, Sarvam). Defensibility is real; the ceiling and the sales cycle are the problem. A reasonable second line once delivery credibility exists, not a first revenue source. Full dossier in the interactive assessment.

6. Physical AI & robotics data Avoid

Real demand, wrong profile. Teleoperation episodes price at $8–$40 each, and earning that requires rigs, physical space, operators, and geographic anchoring: capital-heavy, ops-heavy and slow, the exact inverse of a two-person software team's advantages. Avoid. Full dossier in the interactive assessment.

Price anchors across the market

The single most important structural fact for an India-based team: categories priced per artifact preserve the cost arbitrage, categories priced per expert hour have already competed it away.

Unit Price Category Source
Complex product replica~$300,000RL environment (e.g. Slack clone)Epoch AI, Jan 2026
Quarterly contract$300k–$500kRL environments, neolab deal sizeEpoch AI, Jan 2026
Website replica ("UI gym")~$20,000RL environmentEpoch AI / SemiAnalysis
Single task$200–$2,000RL environment task (rarely to $20k)Epoch AI, Jan 2026
Exclusivity premium4–5×RL environmentsEpoch AI, Jan 2026
Compute burned per task~$2,400RL training, the reason labs pay upMechanize estimate
Expert hour, medical MD$100–$250Expert data generationContractor-reported, 2026
Expert hour, offensive security$200–$250Expert data / red-teamingContractor-reported, 2026
Expert hour, marketplace average~$85Mercor blended rateContractor-reported, 2026
Red-team findingup to $100,000Bug bounty maximumOpenAI programme
Teleoperation episode$8–$40Robotics dataRobotics Center of Silicon Valley; Dexset
India commodity labeling hour$5–$10The tier to stay out ofMarket rate, 2026
India salaried annotator, annual~₹9.3 lakh≈ $11,000SalaryExpert, 2026

Why margin has to come from software

The one hard net-margin figure available in this market is sobering, and it should shape the whole business model.

Invisible Technologies EBITDA

~11%

Mercor take rate on gross

~35%

Contractor share of marketplace top line

60–70%

India cost advantage vs US delivery

60–80%

Invisible Technologies runs roughly 11% EBITDA on about $134M of revenue (Sacra). Human-in-the-loop is a thin-margin business unless automation is layered on top. Note too that headline "ARR" figures across this sector are usually gross marketplace volume before contractor payouts. Turing's economics are explicitly described as a staffing spread, earning the difference between what clients pay and what talent receives.

Your arbitrage survives only where pricing is per artifact. An India engineering team building environments sold at $20k apiece into US per-environment pricing captures the gap. The same team reselling expert hours does not.

Three ways in, ranked

1

Subcontract environment and coding-gym capacity to US vendors

The clearest structural opening in the market: environment companies already hire overseas developers to replicate site UIs (SemiAnalysis), and OpenAI reportedly bought hundreds of roughly $20k UI gyms. Be that development shop: for environment vendors first, then directly for labs.

Why: Fastest to revenue, engineering-native, non-personal data, no expert network required. Price per environment, not per hour.

2

Sell environments and evaluations directly to applied-AI companies

Agent startups and neolabs need training and evaluation environments but cannot build their own. Smaller deals, but you own the customer relationship and the roadmap.

Why: Fallback if labs accelerate in-housing, and the natural home for a vertical specialisation later.

3

Sell to Indian sovereign model programmes

IndiaAI Mission RFEs, CPP portal empanelment, and direct contracts with BharatGen, Sarvam and others.

Why not first: Slow tender cycles and thin budgets make this a poor first revenue source, though a reasonable second line once you have delivery credibility.

Structure the company as a US wrapper over India engineering, the Turing pattern (Palo Alto incorporation, India-heavy delivery). India-domiciled entities do not appear to win direct frontier RL contracts at scale today, and the wrapper also opens government-adjacent work later. iMerit made the same move: India-founded, now headquartered in Los Gatos, climbing from annotation into frontier expert-data tiers.

What you must clear before serious work

Certifications

Data residency limits

India: DPDP Act

Worker classification

Twelve months, three stages

Months 0–2 Build proof in public

Pick RL environments and build a portfolio labs can see

  • Ship three to five high-quality coding and enterprise-workflow environments publicly on Prime Intellect's Environments Hub and claim bounties, typically $1,000–$5,000 each. The field openly rewards people who create benchmarks that actually get used, and this is your cheapest credential and your lead-generation channel at once.
  • Instrument everything as reusable tooling from day one: a verifier framework, a reward-hacking test harness, a difficulty-calibration pipeline. Your defensibility is software, not headcount.

Gate to advance: two environments adopted or downloaded, or one paid bounty, within eight weeks.

Months 2–6 First revenue

Convert the portfolio into subcontracted contracts

  • Approach environment vendors (Halluminate, HUD, and similar shops) and agentic applied-AI startups as a subcontracted development partner for UI and coding gyms, anchoring on the roughly $20k-per-environment reference price.
  • Price per environment and per task, never per hour. That is the only way the India cost structure becomes margin instead of a discount you hand to the buyer.
  • Incorporate the Delaware C-corp over the India engineering entity.
  • Employ your first India engineers under the new Labour Codes with proper appointment letters.

Gate to advance: first paid environment contract of $20k–$100k by month six. No contract by then means switch to path 2 before considering a category change.

Months 6–12 Productise

Specialise, then certify

  • Start SOC 2 Type II and ISO 27001 the moment a lab or vendor deal requires it. Expect six to twelve months elapsed and low-to-mid five figures in audit fees. Sequence ISO 27001 first to shorten any later ISO 42001 path.
  • Specialise into a vertical where a verifier moat is real: long-horizon multi-application enterprise workflows (the explicit 2026 growth area), or a regulated coding and finance gym where domain-informed verifiers are hard to fake.

Gate to advance: a recurring or exclusive environment contract. Exclusivity prices at 4–5× non-exclusive.

What would change the plan

How much to trust these numbers

  • Revenue and pricing figures are largely estimates or single-source. RL environment prices come from Epoch AI's anonymised interviews and SemiAnalysis. Mercor, Surge and Handshake revenues are gross marketplace volume, not net, and mix Sacra estimates, reporting by The Information, and company disclosures. Treat everything here as directional.
  • The Anthropic ~$1B environment figure and the OpenAI compute numbers are reported discussions and projections, not audited spend.
  • Contractor pay rates are aggregator- and contractor-reported. No tier-one vendor publishes buyer-side pricing.
  • The China data-pipeline reporting is single-source and its primary page was inaccessible. Reported, not confirmed.
  • Some sources are vendor or SEO content rather than primary reporting. Validate any specific figure against a real contract before committing capital to it.
  • Forecasts are forecasts. "India as a training ground for physical AI" and the $10B India data market by 2030 are Counterpoint projections, not realised figures.
  • The market is moving fast. Epoch itself expects the RL-environment picture to look quite different within a year.

Compiled September 2026 from Epoch AI, SemiAnalysis, Sacra, The Information, Time, Contrary Research, Pebblous, The Business Research Company, Counterpoint, SalaryExpert, Precise BPO, Shaip, AI4Bharat, and company and government sources including MeitY and the IndiaAI Mission.

About the authors

Mohinish Shaikh and Pragadeesh VS work on AI software, media and investing at Serverlessvc.com.

Prepared as a market-entry assessment. Not investment advice. Figures are dated and, in several cases, estimates or single-source; see the reliability notes above.