Do not make this a framework tutorial.
The learner has already built the important runtime ideas. Now use a higher-level abstraction and interrogate it.
First define the layers
MODEL
Reasoning and generation capability.
HARNESS
The machinery that turns a model into an agent:
planning loop, tools, filesystem/workspace, subagents,
context-management conventions.
ORCHESTRATION RUNTIME
Execution state, persistence, interrupts, durable continuation,
streaming and workflow control.
APPLICATION GUARANTEES
Evidence rules, business state, idempotency, worker ownership,
authorization, budgets, completion policy and evals.
Frameworks can span several layers, but the distinction helps engineers ask the right questions.
Where LangGraph fits in this course
Use LangGraph as the orchestration runtime for:
explicit graph/state transitions
checkpoint persistence
durable execution semantics
interrupt/resume
streaming
Do not say “LangGraph makes everything durable” without naming the exact behavior being relied on.
Where Deep Agents fits
Deep Agents is a higher-level harness on top of LangGraph. Its current design includes capabilities around:
planning
filesystem-backed work/context
subagents
context management
memory/backends
This can replace a large amount of custom agent-loop code.
The production question is not:
Is Deep Agents more powerful?
It is:
Which of our invariants still pass after the port?
Port the Vendor Review Agent
Give the agent a narrow tool set:
discover_vendor_pages
fetch_vendor_page
store_evidence
query_verified_findings
request_review
Working artifacts might be:
/plan.md
/progress.md
/open_questions.md
Specialized subagents:
security_researcher
pricing_researcher
Working files are not automatically system-of-record state
A plan file can be useful for model coherence.
It should not become the authoritative source for:
who currently owns the run
whether a reviewer approved
whether an external publish succeeded
which tenant may see the run
Keep high-consequence business truth in structured durable state with deterministic validation.
Subagents: why they can help
Do not teach “more agents = better.”
A subagent is useful when it creates a meaningful boundary:
main supervisor
↓ delegate one bounded research problem
security subagent
↓ receives security-specific context/tools
↓
returns structured result
↓
main supervisor
Result schema:
class ResearchResult(TypedDict):
requirement: str
finding: str | None
evidence_ids: list[str]
unknowns: list[str]
The main agent gets the result rather than an entire noisy internal transcript.
This is context isolation.
Parallel/async subagents introduce runtime questions
If work can run concurrently or in the background, explicitly ask:
What is the subtask identity?
Where is its partial progress stored?
Who owns it?
How is it cancelled?
What if the parent process disappears?
What if a result arrives twice?
How is its budget bounded?
How are concurrent findings merged?
A framework may answer some of these. Your application still needs the answers.
Harness comparison
Run the same evaluation set against:
A. explicit LangGraph Vendor Review Agent
B. Deep Agents Vendor Review Agent
Compare:
verified outcome quality
evidence correctness
trajectory length
context size
model calls
cost
crash recovery
security invariants
duplicate work
implementation complexity
The purpose is not to declare one universally superior.
It is to teach an engineering method for evaluating harnesses.
Framework portability checklist
For any future agent framework, answer:
Where is durable state stored?
What identifies one logical run?
Where are checkpoints written?
What re-executes after process failure?
How are external writes made safe?
How does unfinished work become runnable?
How is concurrent ownership controlled?
How do human waits resume?
How does cancellation propagate?
How are tool permissions enforced outside the model?
How are model/tool traces inspected?
How do we run offline evals?
If one answer is unclear, build a failure lab instead of trusting a marketing phrase.
FAILURE LAB 12: Replace the Harness
Port the agent, then run the existing invariants unchanged.
The important outcome is not “the Deep Agent completed the task.”
It is a comparison report showing exactly which guarantees remained application-level and which implementation burden moved into the framework.
Check your understanding
Answer before moving on. If one is fuzzy, the relevant section is a scroll away.
- What is a harness?
- What is the difference between harness features and application guarantees?
- Why can a filesystem artifact be useful without becoming the source of truth?
- What production questions appear when subagents become asynchronous?
- How would you evaluate a new framework without relying on a demo?
Exit criteria
Observable conditions, not “I understand it”. Check them off; progress is saved in your browser.
- The port reuses your application tables as the source of truth
- Idempotency keys and the lease scheduler are wired at the tool boundary of the harness
- The comparison matrix is complete, with verification notes per row
- The four-lab subset is green on both implementations
Primary sources
- Deep Agents overview and context engineering: read the mechanisms, then map each to the module where you built it by hand.
- Anthropic on effective harnesses for long-running agents: the industry statement of the harness/runtime split this module makes you feel.