Separate request handling from execution
Avoid this architecture:
POST /run
↓
keep HTTP request open
↓
execute for 40 minutes
↓
return final response
Use:
POST /runs
↓
validate input
↓
create durable run
↓
return 202 Accepted + run_id
The API has accepted work. It has not promised the work is complete.
Execution happens asynchronously
Browser
↓
FastAPI
↓ creates run row
PostgreSQL
↓
Scheduler/worker claims runnable run
↓
LangGraph executes
↓
checkpoints + events + artifacts persist
For a small free deployment, API and worker loops may share one service/container. Keep their responsibilities conceptually separate.
Run lifecycle
Example non-terminal states:
pending
running
retry_wait
awaiting_review
publishing
Terminal states:
completed
completed_with_unknowns
rejected
cancelled
failed
Explicit terminal reasons matter for operations and evals.
Startup reconciliation
This section must be much clearer than typical tutorials.
When a worker starts:
1. connect to durable stores
2. verify database/schema compatibility
3. establish worker identity
4. begin claim loop
5. discover pending runs
6. discover expired leases/orphans
7. reclaim eligible work
8. resume from durable runtime state
This is the bridge between:
“My checkpoint is still in PostgreSQL.”
and:
“My product actually continued the job.”
A checkpoint does not schedule itself.
Health versus readiness
Health
Answers:
Is this process alive?
Readiness
Answers:
Can this process safely accept or execute work?
Readiness may require:
database reachable
required migrations applied
secrets/config present
scheduler started
dependency initialization complete
A process can be healthy but not ready.
Deployment compatibility
A long-running job may outlive a software release.
Every run should record at least:
code_version
state_schema_version
When code changes, classify:
Compatible
New code can safely resume old state.
Migratable
Old state must be transformed before resume.
Breaking
Drain/pin/terminate old runs explicitly rather than hoping they work.
Redeploy timeline
Worker A owns run_123
↓
deploy begins
↓
Worker A terminated
↓
lease remains temporarily
↓
Worker B starts new version
↓
startup claim loop runs
↓
A's lease expires
↓
B claims run with new lease generation
↓
loads durable checkpoint
↓
continues
Now every concept from Modules 03 and 05 connects to deployment.
Public/free hosting
A free web host and serverless Postgres can be used for the learning deployment.
Be explicit in the copy:
Free hosting is a reproducible learning/portfolio environment, not a production SLA.
Process spin-down and restarts are useful failure injectors because the architecture should treat local memory as disposable.
Do not claim an in-process FastAPI background task is a durable queue.
If the course uses a small DB-backed worker loop for simplicity, say so and explain what would change in a larger production deployment.
Abuse and cost controls
A public agent endpoint can spend real money.
At minimum discuss:
authentication or abuse-resistant access
per-user/tenant concurrency
maximum pages per run
maximum model steps
maximum cost
request-size limit
allowed URL/domain policy
Runbook
Require the learner to answer:
How do I find a stuck run?
How do I cancel it?
How do I inspect ownership?
How do I see the last useful progress?
How do I investigate an ambiguous side effect?
How do I find its trace?
How do I disable new run creation?
How do I roll back a bad deploy?
A service is not fully built if nobody can operate it.
FAILURE LAB 13: The Redeploy
Restart/redeploy while the run is:
- in an ordinary research step;
- waiting for review;
- owned by a worker;
- near/inside an ambiguous publish.
Verify:
committed state survives
runnable work is discovered automatically
expired ownership is reclaimed
human wait survives
publish does not duplicate
code/schema version is visible
Normal recovery should require no manual database editing.
Check your understanding
Answer before moving on. If one is fuzzy, the relevant section is a scroll away.
- Why should API request handling be decoupled from long execution?
- What is startup reconciliation?
- Why does a checkpoint not automatically resume itself?
- What is health versus readiness?
- Why should a run record code and schema versions?
Exit criteria
Observable conditions, not “I understand it”. Check them off; progress is saved in your browser.
- Public URL live on the free tier, deployed from main
- The redeploy lab passed, both mid-research and mid-review
- Migrations are additive and run before the new code starts
- Security minimums applied: rate limit, allow-list, escaped rendering, untrusted-content fences
- deploy_checklist.md records what you verified, including the redeploy lab
Primary sources
- Render free instance limitations and Neon free-plan limits: the two pages that decide whether Lesson 14.2’s constraints still hold. Hosting limits change without notice; verify both before relying on this module’s numbers.