Three audiences
Design observability for different users.
Reviewer/user
Needs:
progress
evidence
unknowns
what is waiting on me
Operator
Needs:
is the run alive?
is it making progress?
who owns it?
is the lease healthy?
can I cancel/recover it?
Engineer
Needs:
model inputs/outputs
tool calls
latency
exceptions
state transitions
retries
versions
A raw trace is not a good end-user progress UI.
Logs, traces, metrics, and audit events
Explain the difference.
Logs
Diagnostic records.
worker started
timeout calling vendor
retry scheduled
Traces
A causal tree/timeline for one execution.
run_123
├─ plan next action
│ └─ model call
├─ fetch pricing
│ └─ HTTP call
└─ verify evidence
└─ model call
Metrics
Aggregated numerical behavior across many runs.
p95 duration
run failure rate
median cost
orphan count
Audit events
Business-significant facts that must be independently inspectable.
approval requested
approval received
publish attempted
publish reconciled
run cancelled
Do not depend on parsing debug logs to reconstruct an approval audit trail.
Application event table
CREATE TABLE run_events (
event_id BIGSERIAL PRIMARY KEY,
run_id TEXT NOT NULL,
ts TIMESTAMPTZ NOT NULL DEFAULT now(),
event_type TEXT NOT NULL,
node TEXT,
attempt INT,
payload JSONB NOT NULL
);
Useful events:
run_created
work_claimed
node_started
node_completed
node_failed
retry_scheduled
evidence_added
checkpoint_committed
approval_requested
approval_received
cancel_requested
side_effect_started
side_effect_reconciled
run_terminal
Correlation
Every important model/tool/runtime record should include:
run_id
Often also:
thread_id
worker_id
code_version
model_id
Without correlation IDs, production debugging becomes timestamp archaeology.
Alive versus making progress
This must be a major concept.
A recent heartbeat means:
worker appears alive
It does not mean:
run is accomplishing useful work
Track:
last_heartbeat_at
last_progress_at
Define progress as a meaningful state change, for example:
new evidence
requirement verified
contradiction found
human decision received
Silent loop example
planner chooses page A
planner chooses page B
planner chooses page A
planner chooses page B
...
The worker can heartbeat forever.
The HTTP service is healthy.
Nothing throws an exception.
The run is still bad.
Add a progress watchdog:
if decisions_since_progress > 8:
emit("suspected_loop")
pause_or_replan()
Operator dashboard
The top of the page should answer operational questions quickly:
Run: run_123
Status: researching
Owner: worker-2
Lease remaining: 38s
Last heartbeat: 3s ago
Last useful progress: 42s ago
Current requirement: security
Coverage: 3 / 5
Unknowns: 1
Errors: 2 transient
Model calls: 17
Cost: $0.14 / $0.40
Pages fetched: 11
[Cancel] [Pause] [Open trace]
Then show a compact event timeline and evidence-backed progress.
Cost attribution
Long-running systems must answer:
What did this run cost?
What caused the cost?
Track cost by at least:
run
model
operation type
code version
For multi-tenant systems, also tenant/user.
SLO thinking
Introduce service-level indicators without inventing universal thresholds.
Examples:
accepted runs reaching a terminal state
orphan recovery time
approval-to-resume latency
duplicate side-effect count
time to first useful progress
Choose thresholds from business consequence, not tutorial convention.
FAILURE LAB 09: Silent Run
Create a planner loop with:
process alive = yes
heartbeats = yes
new durable progress = no
A passing implementation surfaces:
alive = true
making_progress = false
suspected_loop = true
The operator should discover this from the dashboard/events without reading source code.
Check your understanding
Answer before moving on. If one is fuzzy, the relevant section is a scroll away.
- Why are logs and audit events different?
- What does a trace show that a metric does not?
- Why is a heartbeat not proof of progress?
- What belongs in an operator view?
- How can a run be unhealthy while every process is healthy?
Exit criteria
Observable conditions, not “I understand it”. Check them off; progress is saved in your browser.
- run_events is append-only, and every node writes started plus finished or failed
- Every LLM decision event records the model’s stated reason
- The dashboard answers the operator questions live, with a working Cancel
- The three alerts exist as queries, tested against seeded events
- A colleague explained a run from the database in under a minute, timed