We intentionally build a fragile system first.
Define the contract
vendor = {
"name": "Acme",
"url": "http://localhost:8001/acme/",
}
CHECKLIST = [
"product",
"customers",
"pricing",
"security",
"developer_experience",
]
Define three tools
def discover_pages(root_url: str) -> list[str]:
...
def fetch_page(url: str) -> str:
...
def extract_finding(page_text: str, requirement: str) -> dict:
...
Only extract_finding fundamentally needs model reasoning.
The course should repeatedly reinforce:
Use deterministic software where the rule is known. Use the model where semantic judgment is genuinely useful.
First state object
state = {
"run_id": "run_123",
"vendor": vendor,
"remaining": list(CHECKLIST),
"findings": [],
"visited_urls": [],
"pages_fetched": 0,
}
At this point, state means simply:
Information the program needs to know in order to decide what to do next.
It lives only in process memory.
The loop
while state["remaining"]:
requirement = choose_requirement(state)
url = choose_page(requirement, state)
page = fetch_page(url)
finding = extract_finding(page, requirement)
state["visited_urls"].append(url)
state["pages_fetched"] += 1
if finding["supported"]:
state["findings"].append(finding)
state["remaining"].remove(requirement)
report = create_report(state)
Operationally:
process memory
│
├─ choose
├─ fetch
├─ model
├─ mutate state
└─ repeat
Hidden assumptions
This small loop assumes:
process remains alive
memory remains available
network calls return
model calls return
repeating a tool is harmless
context remains useful
only one worker executes this run
the completion rule is trustworthy
The rest of the course removes these assumptions one at a time.
Deterministic fixtures
Do not make reliability labs depend on the live internet.
Use fixture vendors and a failure profile:
routes:
/flaky/pricing:
responses:
- status: 503
- status: 503
- status: 200
fixture: pricing.html
This ensures every learner can reproduce the same fault.
FAILURE LAB 01: Kill the Agent
Run until:
Product complete
Customers complete
Pricing in progress
Security pending
Developer experience pending
Stop Python.
Restart it.
What does the program know?
Nothing about the previous run.
It lost:
run identity
vendor input
discovered pages
visited pages
findings
current progress
remaining work
cost already spent
Measure the failure
Run again and record:
pages repeated
model calls repeated
time repeated
estimated cost repeated
Reliability has a cost dimension.
The question that leads to Module 02
Do not ask:
Which database should I use?
Ask:
Which facts must survive for another process to continue correctly?
Check your understanding
Answer before moving on. If one is fuzzy, the relevant section is a scroll away.
- What does “state” mean in this module?
- Which parts of the system actually need model reasoning?
- What assumptions does an in-memory loop make?
- Why use fixture websites?
- Which exact information disappeared after the kill?
Exit criteria
Observable conditions, not “I understand it”. Check them off; progress is saved in your browser.
- The naive loop runs end to end against the fixture server
- Step profile recorded: latency and call counts per tool across a full run
- You killed the process and documented, in resume_answer.md, exactly what was lost and what it cost to repeat
- You can recite the loop’s assumption list (memory, one process, calls return, tools re-runnable, context fits) without the page open
Primary sources
- LangGraph overview: read only the first section now. The point is to see that the framework you meet in Module 03 is built around exactly the deterministic-control-plus-LLM-decisions mix you just wrote by hand.