Build the Environments,
Not the Workforce
Market entry assessment, September 2026: six candidate categories, one recommendation.
By Mohinish Shaikh and Pragadeesh VS
Published: September 7, 2026 · Read time: ~14 minutes
The thesis
The commodity labeling tier is being automated toward zero. The layer above it, reinforcement-learning environments with working verifiers, is the only high-value category where the deliverable is software, the buyer pays per artifact rather than per hour, and two engineers in India can compete on craft instead of headcount.
Recommended entry: RL environments with verifiers. Scope: which of six high-value data categories a two-person, engineering-led startup with an India-based workforce should enter. All figures dated and sourced. Several are estimates or single-source; see the reliability notes at the end.
What the market looks like right now
Demand is enormous and highly concentrated, and it has split cleanly in two: the bottom is collapsing in price while the top is starved of supply.
Annual human-data spend per frontier lab (Time, 2025 investigation)
Of category revenue held by Scale, Surge, Mercor and Handshake (Deedy Das venture map, via Pebblous)
Of routine labels now correctly pre-filled by foundation models, compressing that tier (Shaip)
Cost advantage of India delivery over US teams for equivalent work (Precise BPO)
Synthetic data is a complement rather than a replacement: cited optimal mixes sit around 60–70% human to 30–40% synthetic, and pure-synthetic training risks model collapse. Machines now handle volume; humans set policy, audit edges, and build the environments the machines train inside. That is where the money moved.
The bench
Six categories measured on the axes that decide whether two engineers can actually win: how badly buyers want it, whether a small team can hold a position, how much capital it eats, how fast it pays, and whether it can be delivered from India without a jurisdictional blocker.
RL environments with verifiers
RecommendedSoftware deliverable, per-artifact pricing, non-personal data, engineering is the moat. Enter here.
Computer-use & agentic trajectories
Strong secondNatural bolt-on to environments. Build synthetic capture; keep real-user personal data out of scope.
Expert data generation
Big pool, weak positionBiggest revenue pool, but the moat is a 30,000-expert network you cannot rebuild. Viable only as a narrow vertical.
Failure-mode data & red-teaming
Skip as primaryAutomating fast, and the premium tier is gated to US persons. Skip as a primary business.
Indic & multilingual data
Home turf, low ceilingHome-turf advantage, but the buyer is a thin domestic budget reached through slow government tenders.
Physical AI & robotics data
AvoidReal demand, wrong profile: capital-heavy, ops-heavy, geographically anchored, slow. Avoid.
Each gauge is five segments; more filled is better. Scores are judgements drawn from the evidence in the dossiers below, not published metrics.
The dossiers
What actually ships, who buys it, what it costs, who already owns the position, and what would stop you delivering it from India.
1. RL environments with verifiers Recommended
What ships
An action space plus surrounding state (file systems, simulated apps, environment variables), usually delivered as a Docker container, plus tasks (a prompt and a grader). Graders are unit tests, state inspection, programmatic checks, or LLM-judges against rubrics. A commercial example: a repo snapshot, task spec, and a verifier scoring each attempt on a continuous 0–1 reward, shipped as "the container and the reward entrypoint," plus reward-labeled agent trajectories. Coding environments usually bundle one task each; a computer-use replica (an Excel, Bloomberg or Airbnb clone) can carry hundreds.
Who buys
Frontier labs (Anthropic; OpenAI and xAI named via job postings), neolabs such as Cursor, and product partnerships (Benchling with Anthropic for biology; OpenAI with Shopify and Stripe for shopping). Environment vendors also subcontract build capacity overseas.
Price anchors
Contracts often six to seven figures per quarter; deal sizes of $300k–$500k cited by one neolab; website replicas around $20k; a complex product replica such as Slack around $300k; per-task $200–$2,000, rarely to $20k for hard software engineering. Exclusive deals cost 4–5× non-exclusive. (Epoch AI survey of 18 insiders, Jan 2026; SemiAnalysis)
Size signal
The Information reported in September 2025 that Anthropic discussed spending "at the level of $1 billion per year on training environments alone." For scale, OpenAI's R&D consumed roughly $19B of about $34B total 2025 spending, and Greg Brockman testified it would spend around $50B on compute in 2026, triple 2025.
Who's there
Specialists: Mechanize (SF, ~$9.1M raised, working with Anthropic, $500k engineer salaries), Halluminate (YC S25, narrowed to financial-services environments, ~$1.3M ARR reported 2025), Prime Intellect (open Environments Hub, 500+ community environments, Verifiers library), HUD, and others catalogued on rl-list.com. The large data vendors (Mercor, Surge, Turing, Handshake) are moving in with distribution, but it is an open question whether labeling operations translate into sound verifier engineering. The field is fragmented and early.
Barriers
Capital: negligible. Skill: high. Reward-hacking resistance takes many iterations; difficulty must be calibrated to roughly a 2–3% minimum pass rate with a smooth gradient and compositional skills. Every interviewee named the same hard problem: scaling task volume without quality collapse. That is a management problem a software-minded team can partly convert into a tooling problem, which is exactly the opening.
Automation risk
Low, and inverted: environments are themselves the machinery that generates synthetic data, so the synthetic shift increases demand here. Real risks are labs in-housing (already "substantially more in-housing") and a flood of vibe-coded clones at the low end: "a large amount of useless bad environments out there."
India delivery
Clear. Environments are non-personal software artifacts, outside the DPDP Act's personal-data scope. Commercial frontier work is open to India-based delivery, and environment vendors already hire overseas developers to replicate site UIs. A US wrapper additionally opens anything government-adjacent.
2. Computer-use & long-horizon agentic trajectories Strong second
What ships
Multi-app, multi-step interaction traces: task prompt, screen states and screenshots, DOM or accessibility trees, mouse and keyboard actions, timestamps, intermediate steps, intent and state labels, completion status, often reward labels. A commercial spec to benchmark against: Datoric's 250,000 real-world computer-use traces covering roughly 15,000 hours of screen activity, rights-cleared with consent documentation and chain of custody. Open datasets (AgentNet, ScaleCUA, OSWorld, Toucan's 1.5M MCP tool-agent trajectories) show the schema and the collection pipelines.
Who buys
Foundation-model labs and browser or computer-use agent companies, plus the same environment vendors (Halluminate sells human data services alongside its sandboxes).
Price anchors
Per-item pricing is opaque. Work prices near agentic rates: roughly $75–$300+ per hour for skilled trajectory work, with per-task overlapping the $200–$2,000 environment band. What buyers actually pay a premium for is rights clearance and consent provenance.
The hard part
Real capture records whatever is on screen: personal data, credentials, third-party information. Consent documentation and chain of custody are simultaneously the moat and the liability, and this is where DPDP and cross-border transfer rules bite hardest if collection happens in India.
Automation risk
Moderate to high. Automated web-exploration pipelines (InSTA's 150k-site Playwright trajectories; Toucan's synthetic MCP traces) scale cheaply but noisily. Human demonstration stays higher quality and more expensive; hybrid is the norm. More commoditizable than verifier engineering.
India delivery
Partial. Fine for synthetic and scripted capture and for building the recording and labeling tooling. Risky for real-user capture involving personal data. Build the pipeline in India; leave personal-data capture out of scope.
3. Expert data generation Big pool, weak position
What ships
Credentialed experts author original problems, worked solutions, grading rubrics, evaluations and reasoning traces: production, not labeling. Domains in demand: medicine, law, senior software engineering, advanced mathematics, quantitative finance, accounting, PhD sciences.
The model
Mercor, Surge, Handshake, Micro1, Turing and AfterQuery place vetted experts onto lab projects. Mercor runs about a 35% take rate; The Information, citing internal documents, put gross revenue at roughly $614M in H1 2026 and a $2B annualized gross run-rate by June (up from about $760M at end-2025 and $500M in September 2025), with 91% of revenue from foundation-model companies such as OpenAI and Anthropic. Sacra estimates contractors keep 60–70% of top line, implying H1 net revenue of roughly $180M–$250M. It distributes about $1.5M per day to 30,000+ contractors across 45+ countries, India its largest talent source.
Rates paid
Mercor average around $85/hr (advertised $81–$114). Tiers: $12–$25/hr entry generalist; $25–$53/hr specialized RLHF, coding and translation; $75–$200+/hr credentialed experts. Medical MDs $100–$210/hr with a top decile near $250/hr; offensive-security $200–$250/hr; law, finance and quant in between. (Contractor-reported, 2026)
Scale of rivals
Sacra estimates Handshake reached $1.1B annualized gross revenue in April 2026, up 349% year over year from about $245M, with roughly $450M net after contractor payouts; The Information put its AI-training gross ARR near $1B, up from $550M in January 2026.
Why it's a weak fit
The moat is the credentialed-expert network and the vetting engine (Mercor's APEX AI interview), not technology, and not replicable by two engineers. You would be reselling the same India expert labour the incumbents already recruit directly, with no differentiation.
India delivery
Open but undifferentiated. Only worth pursuing as a jurisdiction-specific vertical where you hold real advantage (Indian law, ICAI accounting, Indian medical boards) rather than head-on.
4. Failure-mode data & red-teaming Skip as primary
What ships
Documented model failures, jailbreaks, prompt-injection exploits, hard-case mining, and severity-rubric-calibrated adversarial datasets with audit trails for EU AI Act Article 55 and NIST AI 600-1 compliance. Kili Technology runs private programmes with 2,000+ specialists across 40+ languages.
Price anchors
Crowd red-teaming $30–$120/hr (entry $20–$30; specialist or direct-to-lab $150–$300). Per finding: a few hundred dollars up to $100,000 (OpenAI's maximum ~$100k); Gray Swan Arena pools $40k–$300k+ per challenge. Engagements: one-time audits $8k–$25k, comprehensive $50k–$150k, continuous $5k–$20k per month.
Market size
The Business Research Company projects growth from $1.75B in 2025 to $2.26B in 2026 at a 28.8% CAGR, reaching $6.17B by 2030.
Why to skip
Agentic red-teaming is automating quickly: XBOW topped HackerOne's US leaderboard in 2025 and valid AI-generated vulnerability reports rose 210% year over year. Meanwhile the highest-value work (weapons-uplift evaluation, classified and federal red teams) is gated to US persons.
India delivery
Commercial tiers only. The premium tier is jurisdictionally closed. A crowded, partly-automated market with its best segment out of reach.
5. Indic & multilingual data Home turf, low ceiling
What ships
Low-resource language corpora, code-mixed text in Hinglish and Tanglish (native script, romanized, and code-switched), ASR and TTS speech corpora, and multilingual safety, red-teaming and evaluation sets. Shaip already packages Indian-language ASR/TTS and code-mixed corpora; AI4Bharat maintains open Indic catalogues; academic Hinglish sets (L3Cube HingCorpus, PHINC, MUTANT) establish the schema.
Two buyer pools
Frontier labs closing multilingual coverage and safety gaps; and India's sovereign programmes: the IndiaAI Mission (launched March 2024), BharatGen, Sarvam (selected April 2025 to build a sovereign LLM; shipped Sarvam-M 24B and open-sourced 30B/105B models trained on IndiaAI compute, explicitly handling native, romanized and code-mixed input), Krutrim, Gnani and Soket.
How work is won
Government empanelment through MeitY and IndiaAI Independent Business Division RFEs on the CPP portal (36-month engagements), the AIKosh datasets platform, and direct contracts with sovereign model builders. This is tender mechanics, not a frontier-lab sales motion.
The ceiling
The domestic budget is thin next to frontier spend, and Indic data is not where the billion-dollar budgets concentrate. Sarvam itself leans heavily on synthetic generation (a 2T-token Indic corpus for Sarvam-1), which compresses demand for human-collected Indic text.
India delivery
Best of the six. But slow tenders and a low ceiling make it capital-inefficient as a first move for a two-person team that needs revenue quickly. Keep as an optional later line.
6. Physical AI & robotics data Avoid
What ships
Teleoperation episodes with synchronized joint angles, gripper forces, camera frames and task context; bimanual manipulation; egocentric video; sensor and LiDAR data. Delivered as HDF5, RLDS or LeRobot formats with calibration metadata.
Who buys
Robotics foundation-model companies (Physical Intelligence, Skild, Figure, 1X, Agility), plus autonomous-vehicle and humanoid programmes.
Unit economics
Cost per episode $8–$40 by complexity, view count and QA. All-in operator cost $28–$60/hr; trained operators produce 25–40 usable episodes per hour; 10–30% QA rejection. A 50,000-episode dataset works out to roughly $67k of collection labour before QA.
Capital
The heaviest of the six. Rigs: ALOHA bimanual around $20k, GELLO about $300 per arm, UMI handheld (robot-free), VR via Quest 3, plus robots, lab space and operator training.
Automation risk
High and active. Simulation is running at roughly eight sim samples to one teleop sample, and robot-free capture (XRZero) claims about one-twentieth the cost, eroding pure teleoperation economics.
India delivery
Possible, wrong profile. Counterpoint projects India could become a training ground for physical AI, but this is capital-heavy, physically anchored, annotation-operations-intensive rather than engineering-intensive, and slow to first revenue: the inverse of this team's strengths.
Price anchors across the market
The single most important structural fact for an India-based team: categories priced per artifact preserve the cost arbitrage, categories priced per expert hour have already competed it away.
| Unit | Price | Category | Source |
|---|---|---|---|
| Complex product replica | ~$300,000 | RL environment (e.g. Slack clone) | Epoch AI, Jan 2026 |
| Quarterly contract | $300k–$500k | RL environments, neolab deal size | Epoch AI, Jan 2026 |
| Website replica ("UI gym") | ~$20,000 | RL environment | Epoch AI / SemiAnalysis |
| Single task | $200–$2,000 | RL environment task (rarely to $20k) | Epoch AI, Jan 2026 |
| Exclusivity premium | 4–5× | RL environments | Epoch AI, Jan 2026 |
| Compute burned per task | ~$2,400 | RL training, the reason labs pay up | Mechanize estimate |
| Expert hour, medical MD | $100–$250 | Expert data generation | Contractor-reported, 2026 |
| Expert hour, offensive security | $200–$250 | Expert data / red-teaming | Contractor-reported, 2026 |
| Expert hour, marketplace average | ~$85 | Mercor blended rate | Contractor-reported, 2026 |
| Red-team finding | up to $100,000 | Bug bounty maximum | OpenAI programme |
| Teleoperation episode | $8–$40 | Robotics data | Robotics Center of Silicon Valley; Dexset |
| India commodity labeling hour | $5–$10 | The tier to stay out of | Market rate, 2026 |
| India salaried annotator, annual | ~₹9.3 lakh | ≈ $11,000 | SalaryExpert, 2026 |
Why margin has to come from software
The one hard net-margin figure available in this market is sobering, and it should shape the whole business model.
Invisible Technologies EBITDA
Mercor take rate on gross
Contractor share of marketplace top line
India cost advantage vs US delivery
Invisible Technologies runs roughly 11% EBITDA on about $134M of revenue (Sacra). Human-in-the-loop is a thin-margin business unless automation is layered on top. Note too that headline "ARR" figures across this sector are usually gross marketplace volume before contractor payouts. Turing's economics are explicitly described as a staffing spread, earning the difference between what clients pay and what talent receives.
Your arbitrage survives only where pricing is per artifact. An India engineering team building environments sold at $20k apiece into US per-environment pricing captures the gap. The same team reselling expert hours does not.
Three ways in, ranked
Subcontract environment and coding-gym capacity to US vendors
The clearest structural opening in the market: environment companies already hire overseas developers to replicate site UIs (SemiAnalysis), and OpenAI reportedly bought hundreds of roughly $20k UI gyms. Be that development shop: for environment vendors first, then directly for labs.
Why: Fastest to revenue, engineering-native, non-personal data, no expert network required. Price per environment, not per hour.
Sell environments and evaluations directly to applied-AI companies
Agent startups and neolabs need training and evaluation environments but cannot build their own. Smaller deals, but you own the customer relationship and the roadmap.
Why: Fallback if labs accelerate in-housing, and the natural home for a vertical specialisation later.
Sell to Indian sovereign model programmes
IndiaAI Mission RFEs, CPP portal empanelment, and direct contracts with BharatGen, Sarvam and others.
Why not first: Slow tender cycles and thin budgets make this a poor first revenue source, though a reasonable second line once you have delivery credibility.
Structure the company as a US wrapper over India engineering, the Turing pattern (Palo Alto incorporation, India-heavy delivery). India-domiciled entities do not appear to win direct frontier RL contracts at scale today, and the wrapper also opens government-adjacent work later. iMerit made the same move: India-founded, now headquartered in Los Gatos, climbing from annotation into frontier expert-data tiers.
What you must clear before serious work
Certifications
- SOC 2 Type II: the baseline security assurance for enterprise and lab procurement. No AI-specific controls, but a standard gate.
- ISO 27001: information security management, usually required in parallel. Sequence this first: holding it cuts the ISO 42001 timeline to about 5–9 months versus 9–14 greenfield.
- ISO 42001: AI management system, increasingly a line item on 2026 vendor questionnaires. Audit fees roughly $25k–$75k for a 51–200-person firm on the accredited path; SMB audit-only $7.5k–$25k. Timeline 3–6 months if mature, 9–14 if starting cold.
Data residency limits
- Regulated and government-adjacent data can require US-based workforces, dedicated no-retention environments, and US-persons-only access (FedRAMP, CMMC, DFARS 252.239-7010, and GSA's 2026 draft AI-safeguarding clause).
- ITAR deemed-export rules make offshore access to controlled technical data a licensing question, not a staffing choice.
- These close defence-adjacent and classified work to an India-based team, but not commercial frontier RL and coding environments.
India: DPDP Act
- Rules notified 13 November 2025. A permissive blacklist model under Section 16(1) and Rule 15: data may transfer anywhere except government-restricted countries. No standard-contractual-clause or adequacy regime.
- Government retains discretion to restrict specific corridors or categories, and Significant Data Fiduciary designation could trigger localisation.
- Crucially, DPDP governs personal data. Synthetic RL environments, code, and non-personal agentic traces largely fall outside its scope, another argument for the recommended category.
Worker classification
- The US is in a litigation wave: class actions against Scale AI (including a PAGA suit alleging roughly $15/hr effective pay against California minimum wage, and a traumatic-content claim), Surge AI (Clarkson, San Francisco Superior Court, May 2025: misclassification, unpaid training, sub-minimum effective wage) and Mercor (Texas, May 2026: ERISA and Internal Revenue Code claims). The pattern is operating like an employer while classifying workers as contractors.
- India's four Labour Codes, in force 21 November 2025, formally recognise gig, platform and aggregator workers, mandate appointment letters for all hires, and fund social security through aggregator contributions of 1–2% of turnover, while preserving non-employee status. Karnataka (2025) and Rajasthan have dedicated gig-worker acts.
- Employing a small India engineering team, rather than gig-contracting it, sidesteps this risk entirely.
Twelve months, three stages
Pick RL environments and build a portfolio labs can see
- Ship three to five high-quality coding and enterprise-workflow environments publicly on Prime Intellect's Environments Hub and claim bounties, typically $1,000–$5,000 each. The field openly rewards people who create benchmarks that actually get used, and this is your cheapest credential and your lead-generation channel at once.
- Instrument everything as reusable tooling from day one: a verifier framework, a reward-hacking test harness, a difficulty-calibration pipeline. Your defensibility is software, not headcount.
Gate to advance: two environments adopted or downloaded, or one paid bounty, within eight weeks.
Convert the portfolio into subcontracted contracts
- Approach environment vendors (Halluminate, HUD, and similar shops) and agentic applied-AI startups as a subcontracted development partner for UI and coding gyms, anchoring on the roughly $20k-per-environment reference price.
- Price per environment and per task, never per hour. That is the only way the India cost structure becomes margin instead of a discount you hand to the buyer.
- Incorporate the Delaware C-corp over the India engineering entity.
- Employ your first India engineers under the new Labour Codes with proper appointment letters.
Gate to advance: first paid environment contract of $20k–$100k by month six. No contract by then means switch to path 2 before considering a category change.
Specialise, then certify
- Start SOC 2 Type II and ISO 27001 the moment a lab or vendor deal requires it. Expect six to twelve months elapsed and low-to-mid five figures in audit fees. Sequence ISO 27001 first to shorten any later ISO 42001 path.
- Specialise into a vertical where a verifier moat is real: long-horizon multi-application enterprise workflows (the explicit 2026 growth area), or a regulated coding and finance gym where domain-informed verifiers are hard to fake.
Gate to advance: a recurring or exclusive environment contract. Exclusivity prices at 4–5× non-exclusive.
What would change the plan
- Labs accelerate in-housing environment construction. Pivot to applied-AI companies and neolabs that cannot build their own.
- Commodity vibe-coded environments crater per-unit prices. Move up-market into verifier robustness and reward-hacking QA as a service.
- US wrapper plus certification proves too slow or costly for two people. Fall back to a specialised expert-data vertical using India's domain supply (Indian law, medicine, accounting) and accept lower margins.
- Cross-border training-data flows come under export scrutiny. Watch this: Forbes reported in August 2026 a roughly $500M US-to-China training-data trade, with AfterQuery earning about $50M from Chinese labs. Single-source, but it signals a regime risk for any cross-border data vendor.
How much to trust these numbers
- Revenue and pricing figures are largely estimates or single-source. RL environment prices come from Epoch AI's anonymised interviews and SemiAnalysis. Mercor, Surge and Handshake revenues are gross marketplace volume, not net, and mix Sacra estimates, reporting by The Information, and company disclosures. Treat everything here as directional.
- The Anthropic ~$1B environment figure and the OpenAI compute numbers are reported discussions and projections, not audited spend.
- Contractor pay rates are aggregator- and contractor-reported. No tier-one vendor publishes buyer-side pricing.
- The China data-pipeline reporting is single-source and its primary page was inaccessible. Reported, not confirmed.
- Some sources are vendor or SEO content rather than primary reporting. Validate any specific figure against a real contract before committing capital to it.
- Forecasts are forecasts. "India as a training ground for physical AI" and the $10B India data market by 2030 are Counterpoint projections, not realised figures.
- The market is moving fast. Epoch itself expects the RL-environment picture to look quite different within a year.
Compiled September 2026 from Epoch AI, SemiAnalysis, Sacra, The Information, Time, Contrary Research, Pebblous, The Business Research Company, Counterpoint, SalaryExpert, Precise BPO, Shaip, AI4Bharat, and company and government sources including MeitY and the IndiaAI Mission.
About the authors
Mohinish Shaikh and Pragadeesh VS work on AI software, media and investing at Serverlessvc.com.
Prepared as a market-entry assessment. Not investment advice. Figures are dated and, in several cases, estimates or single-source; see the reliability notes above.