Checkpointing made work resumable. It did not decide who should resume it.
The race
With one worker:
Worker A → run_123
With two workers, naive polling can do this:
Worker A Worker B
│ │
│ SELECT pending run │ SELECT pending run
│ → run_123 │ → run_123
│ │
│ execute │ execute
Both believe they own the same job.
You now pay for duplicate model calls and duplicate tool work. Module 04’s idempotency may protect the final write, but ownership should stop the duplicate work earlier.
What a scheduler actually needs to answer
Before introducing queue products, teach the underlying questions:
Which runs are ready?
Which worker owns each run?
When does that ownership expire?
How does an abandoned run become available again?
Claim
A claim is an atomic state change that grants a worker authorization to execute a run.
Conceptually:
worker-7 owns run_123
The selection and claim cannot be separate raceable operations.
Why SELECT then UPDATE is unsafe
Bad:
run = select_pending_run()
claim(run)
Two processes can select the same row before either updates it.
This is a race condition: correctness depends on timing between concurrent operations.
Atomic PostgreSQL claim
Use a transaction/row lock pattern such as:
WITH candidate AS (
SELECT run_id
FROM runs
WHERE
status = 'pending'
OR (
status = 'running'
AND lease_expires_at < now()
)
ORDER BY created_at
FOR UPDATE SKIP LOCKED
LIMIT 1
)
UPDATE runs
SET
status = 'running',
lease_owner = $1,
lease_expires_at = now() + interval '60 seconds'
WHERE run_id IN (SELECT run_id FROM candidate)
RETURNING *;
Explain the SQL rather than presenting it as magic.
FOR UPDATE locks the selected row inside the transaction.
SKIP LOCKED tells another worker to skip rows already claimed by a concurrent transaction rather than waiting and then taking the same work.
Lease
A permanent lock would strand work when a worker dies.
A lease is temporary ownership.
worker-7 owns run_123 until 16:05:00
The worker must renew it before expiry.
If it disappears, the lease eventually expires and another worker can reclaim the run.
Heartbeat
A heartbeat is periodic evidence that a worker is alive.
UPDATE runs
SET
heartbeat_at = now(),
lease_expires_at = now() + interval '60 seconds'
WHERE
run_id = $1
AND lease_owner = $2;
Important distinction:
heartbeat = worker appears alive
progress = useful business state changed
A worker can heartbeat while stuck forever. Module 09 will detect that.
Orphan
An orphaned run is unfinished work whose previous ownership is no longer valid.
Example:
status = running
owner = worker-7
lease_expires_at = 15:00
now = 15:07
The run is eligible for recovery.
Fencing tokens: deeper production correctness
A lease timestamp alone has a subtle race.
Timeline:
Worker A owns run
↓
A pauses for 90 seconds
↓
lease expires
↓
Worker B claims run
↓
A wakes up and continues
If A can still commit authoritative writes, you have split ownership.
Use an ownership generation/fencing token:
A owns generation 18
B reclaims and receives generation 19
Every authoritative worker mutation requires the current generation:
UPDATE runs
SET ...
WHERE run_id = $1
AND lease_generation = $2;
A stale worker with generation 18 can no longer commit after generation 19 exists.
This is a valuable advanced detail because it explains what a robust lease is protecting against.
Cancellation
Cancellation should be durable too.
UPDATE runs
SET cancel_requested = true
WHERE run_id = $1;
The worker checks at safe boundaries:
if run.cancel_requested:
transition_to_cancelled()
return
Do not make “cancel” depend only on killing one process; another worker could otherwise resume it.
Backpressure
Long-running work can be created faster than workers can process it.
Add limits such as:
max active runs
max active runs per tenant
max concurrent model calls
max run cost
Scheduling is partly a cost-control problem.
FAILURE LAB 05: Two Workers, One Run
- Start two workers with broken select-then-update claiming.
- Observe duplicate execution.
- Replace it with atomic claim.
- Prove at most one current owner.
- Kill that worker.
- Wait for lease expiry.
- Verify a second worker reclaims and resumes.
- Advanced: let the stale first worker wake up and prove its old fencing token cannot commit.
Check your understanding
Answer before moving on. If one is fuzzy, the relevant section is a scroll away.
- Why does checkpointing not solve ownership?
- What is a race condition?
- Why must a claim be atomic?
- Why does a lease expire?
- Why can a fencing token be stronger than only checking time?
Exit criteria
Observable conditions, not “I understand it”. Check them off; progress is saved in your browser.
- Claiming is one atomic statement with FOR UPDATE SKIP LOCKED
- Heartbeat extends the lease inside the step loop, not in a background thread
- Every commit is fenced with owner and lease predicates, and a zero-row write aborts the worker
- An orphaned run is reclaimed within one lease interval, automatically
- Cancellation is honored at the next step boundary via the status column
Primary sources
- Temporal’s AI cookbook shows what claim, lease, heartbeat and reclaim look like when a durable workflow engine owns them; compare each concept to the table you just built.