YouTube · 23 Aug 2026
View the original on youtube.comTopicBuilding and evaluating AI agent harnesses
Building and evaluating AI agent harnesses
By Sequoia Capital · YouTube
Source: https://www.youtube.com/watch?v=HI2q3ci3Iuc&list=PLaqC3GACblSs&index=2
AI generated summary of the linked source. Posted anonymously. Transcripts are published only for public links, never for uploads, documents or pasted text.
Share
The AI link is plain markdown. Paste it into any assistant and it can read the whole public entry.
Summary
Harrison (LangChain co-founder/CEO) breaks down agents into three owned parts: model, context, and harness, then argues that harnesses matter most for control and differentiation. He walks through how harnesses work under the hood, when to build custom versus use off-the-shelf ones like Claude Code or Codex, and how evals plus observability (via tools like Harbor and LangSmith) create a data flywheel to keep improving agents.
Core ideas
- Three parts of intelligence: An agent is made of a model, context, and a harness. Owning your intelligence means owning all three, not just picking a good model.
- Harness as orchestrator: The harness's main job is feeding the right context to the model at the right time, running the loop, calling tools, and handling responses.
- Simple loop, complex wrapping: Nearly every agent today is just an LLM running in a loop calling tools. Advanced harnesses like Deep Agents add sandboxes, file systems, sub-agents, and summarization through "middleware" hooks around that same core loop.
- Cognitive architectures are fading but not dead: Bespoke, staged pipelines (common in 2023-2024 when models were weaker) still get used for domains needing tight control, like financial services.
- In distribution vs out of distribution: Off-the-shelf harnesses (Claude Code, Codex) work best when your task matches what the model was RL'd on. Legal AI, for example, is out of distribution overall but file editing itself is in distribution, so a custom harness should still reuse the model's native edit-file tool.
- Model profiles trick: Deep Agents switches between different file-edit implementations depending on which underlying model is running, to stay in distribution on small sub-tasks.
- Evals define "good": Quoting Satya, private evals set the internal standard for quality, and owning your traces, feedback, and institutional context builds a compounding "hill climbing machine."
- Harbor as emerging standard: An open-source eval runner (from the Terminal Bench 2 team) that packages tasks into environment, solution, test, and instruction files, run in sandboxes and scored automatically.
- Data flywheel: Build agent, collect traces, curate feedback (explicit or synthetic via cheap fine-tuned judge models), then use that data to update the harness, model, or context. LangChain's "Engine" agent automates parts of this by scanning traces and proposing fixes.
Quotes
“The main job of a harness is to bring context to the model at the right point in time.”
Harrison
“When agents mess up, they mess up because an LLM call goes wrong. Why might it go wrong? It might go wrong for one of two reasons. One, the model is not good enough. Two, the context that the LLM received isn't good enough.”
Harrison
“The more in distribution you are of what the models are trained on, then the better the off-the-shelf harness will be.”
Harrison
“Create your private evals because eval defines what good looks like inside the organization.”
Harrison
“Run agent, get traces, see patterns, fix.”
Harrison
“I don't know is the answer to the fast moving space, that's why evals and observability are important.”
Harrison
Resources
Links to the tools, products, repos or reading this entry mentions. Found by a web search of what the entry names, so you can go and use them.
Anyone can run this once. After that the links are saved on this entry and shown to everyone.
Remix this knowledge
Turn this entry into something you can post. Pick a format and get a fresh, original rewrite.
Remixes are AI generated, original wording, and free to reuse.
Talk it through
Open a round table on this entry and argue with other people about what it means.
Start a round tableDrop your own
Paste any video, podcast or audio link and get the transcript, the core ideas and the quotes back in one pass.
Try Drop