YouTube · 23 Aug 2026
View the original on youtube.comTopicContinual learning for AI agents
Continual learning for AI agents
By Sequoia Capital · YouTube
Source: https://www.youtube.com/watch?v=eYrMF9Cht8A&list=PLaqC3GACblSs&index=1
AI generated summary of the linked source. Posted anonymously. Transcripts are published only for public links, never for uploads, documents or pasted text.
Share
The AI link is plain markdown. Paste it into any assistant and it can read the whole public entry.
Summary
Arjun (Trajectory) argues today's AI models are getting smarter but never gain experience, staying stuck at "day one on the job" performance. His company builds infrastructure for continual learning: capturing agent interactions, turning them into training signal, and routing feedback to either model weights or harness/context depending on what kind of knowledge it is.
Core ideas
- The experience gap: Models keep improving on IQ but act like it's always their first day. A genius with zero experience underperforms someone less smart but seasoned, and closing that gap is Trajectory's whole thesis.
- Wasted signal: Hundreds of millions of tokens of real agent work get generated daily and thrown away instead of being used to make agents better over time.
- Four wishes framework: Arjun frames the unsolved problems as four wishes: traceability, evals, harness design, and model flexibility, each blocking real continual learning.
- Traceability means more than logging: Companies need to trace entire trees including sub-agent calls, and capture corrective behavior (edits, undos, retries) not just thumbs up/down, which is too noisy.
- Evals should mirror production: Training, eval, and real product use should converge, meaning evals drawn from actual traffic, tasks that are replayable, and grading through the real harness.
- Harness as orchestrator, not gatekeeper: Older harnesses were built to stop agents from breaking; now harnesses should expose primitives (search, tools, private data) and let capable agents orchestrate them, with informative tool responses instead of vague "done" messages.
- Model vs harness routing decision: Global, stable facts (a tool repeatedly failing) should train into the model; user-specific or ephemeral preferences (a company got delisted, a user hates a sub-agent) should live in context or per-org harness settings.
- Privacy without raw data training: Instead of training directly on customer data, Trajectory samples distributions and synthetically generates comparable data, checking whether synthetic data is "on distribution" without ever training on the real thing.
- Continual learning as a systems problem: Rather than treating it as pure real-time weight updates, Trajectory treats intelligence as a system of components (weights, harness, context) and optimizes which part should update, similar to abstracting away RAM vs disk decisions.
Quotes
“We're living in an incredible time in human history here.”
Arjun
“It always feels like when you're talking to them, it's their first day on the job.”
Arjun
“If you ask like six researchers what continual learning is, you're going to get like seven answers, probably.”
Arjun
“Today's agents, they're still slow, expensive, error prone over time. You implement them, they probably don't get better.”
Arjun
“It's the corrective behavior, the edits, the undos and the retries that both need to be elicited from the user, but also captured.”
Arjun
“We're very much in a let the agents cook world.”
Arjun
Resources
Links to the tools, products, repos or reading this entry mentions. Found by a web search of what the entry names, so you can go and use them.
Remix this knowledge
Turn this entry into something you can post. Pick a format and get a fresh, original rewrite.
Remixes are AI generated, original wording, and free to reuse.
Talk it through
Open a round table on this entry and argue with other people about what it means.
Start a round tableDrop your own
Paste any video, podcast or audio link and get the transcript, the core ideas and the quotes back in one pass.
Try Drop