Incident Recovery
Agents will fail. This runbook is the sequence you follow when a production agent goes wrong — the same discipline incident responders use, applied to LLM systems.The Runbook
1. Stop the bleeding
Freeze the affected session so the agent can’t keep mutating state while you investigate:2. Reconstruct what happened
Do not trust the agent’s summary. Replay the actual decision trace:3. Find the divergence point
Identify the exact turn where state stopped matching reality. This is your rollback target:4. Roll back to the last good state
5. Prevent recurrence
Check what enabled the failure:- Missing checkpoint? Add a checkpoint before the risky operation.
- No validation? Add a schema check on tool output.
- No permission check? Enforce an allowlist (see HITL).
- State drift? Sync with the source of truth (see Failure Modes).
6. Resume
Severity Matrix
Post-Incident Review
- What failure mode was it? (hallucination, drift, tool failure, loop…)
- Where was the missing checkpoint?
- Add the scenario to your failure-injection tests (see Templates).
Next Steps
- Failure Modes: the 7 ways agents fail
- Rollback Strategies: technique comparison
- Replay & Audit: debugging production runs