YouTube · 23 Aug 2026
View the original on youtu.beTopicRL environments for AI training
RL environments for AI training
By Sequoia Capital · YouTube
Source: https://youtu.be/a00xIn5kwhM?si=gKYi_gOcize7VQc8
AI generated summary of the linked source. Posted anonymously. Transcripts are published only for public links, never for uploads, documents or pasted text.
Share
The AI link is plain markdown. Paste it into any assistant and it can read the whole public entry.
Summary
Brendan from Mercor breaks down how RL environments work: worlds, apps, and tasks built by domain experts to train frontier AI models. He walks through the shift from crowdsourced behavior cloning data in 2020 to today's high-skill expert-built environments, using a legal data room example to show how this drives real model performance gains.
Core ideas
- From crowdsourcing to expertise: The data market moved from low-skilled annotators doing behavior cloning and RLHF preference labeling in 2020 to high-skilled professionals (lawyers, doctors, bankers, engineers) building frontier evals and RL environments.
- Three parts of an RL environment: Worlds (messages, docs, sheets mimicking real projects), apps (high-fidelity clones of Salesforce, ServiceNow, Microsoft 365 accessible via MCP, CLI, or Kua), and tasks (prompts plus verifiers like rubrics or unit tests).
- Humans grade the homework: Models can't reliably self-assess mistakes outside clean domains like math. Experts build rubrics the way a professor grades an essay, because a model can't accurately judge its own slide deck.
- Massive scale up: Mercor logged 2.5 million expert hours in Q2 alone, growing to cover categories across the whole economy, referencing GDPval's 205 BLS job domains as the target distribution to replicate.
- Real performance gains: Post-training GLM 4.7 on just 1,800 Apex Agents tasks (about 500k compute) pushed corporate law performance from 4.7% to 26.6%, and the gains generalized to other benchmarks like GDPval that weren't in the training set.
- Three data business models: Custom per-task pricing (from $50 to $10,000 a task, sometimes $2,000 flat for specific shapes), off-the-shelf data sets sold to multiple labs, and hourly expert staffing (Mercor's smallest and shrinking focus).
- Quality means two things: Realism (does the environment reflect the true distribution of a real job) and verifier accuracy (does the rubric score trajectories the way a human stack-rank would).
- Synthetic data is often misunderstood: RLVR itself is a bet on synthetic trajectories, and models help experts build environments faster, but humans remain essential because you need judgment beyond a model's frontier to grade it reliably.
- What comes next: Ultra-long-horizon tasks (100 to 1,000 human-hours) and "virtual co-worker" scenarios that test social and multi-person interaction, since 60-70% of real jobs involve working with other people but almost no evals measure that.
Quotes
“RLVR is a bet on synthetic data.”
Brendan
“It's as if you would be asking a human to grade their own homework.”
Brendan
“We have tasks that range from $50 to $10,000.”
Brendan
“Data's often the most differentiating factor.”
Brendan
“You need something that has capabilities beyond the frontier of that model to do so reliably.”
Brendan
“There's this giant realism gap associated with how you actually measure how well agents engage in social interaction.”
Brendan
Resources
Links to the tools, products, repos or reading this entry mentions. Found by a web search of what the entry names, so you can go and use them.
Anyone can run this once. After that the links are saved on this entry and shown to everyone.
Remix this knowledge
Turn this entry into something you can post. Pick a format and get a fresh, original rewrite.
Remixes are AI generated, original wording, and free to reuse.
Talk it through
Open a round table on this entry and argue with other people about what it means.
Start a round tableDrop your own
Paste any video, podcast or audio link and get the transcript, the core ideas and the quotes back in one pass.
Try Drop