Skip to main content

Incident Recovery

Agents will fail. This runbook is the sequence you follow when a production agent goes wrong — the same discipline incident responders use, applied to LLM systems.

The Runbook

1. Stop the bleeding

Freeze the affected session so the agent can’t keep mutating state while you investigate:
Redirect new traffic to a safe version or the human queue.

2. Reconstruct what happened

Do not trust the agent’s summary. Replay the actual decision trace:

3. Find the divergence point

Identify the exact turn where state stopped matching reality. This is your rollback target:

4. Roll back to the last good state

Rollback is versioned and auditable — the rollback itself becomes a new version, so you can inspect or undo it.

5. Prevent recurrence

Check what enabled the failure:
  • Missing checkpoint? Add a checkpoint before the risky operation.
  • No validation? Add a schema check on tool output.
  • No permission check? Enforce an allowlist (see HITL).
  • State drift? Sync with the source of truth (see Failure Modes).
Patch the agent, test the scenario in staging with injected failures, then resume.

6. Resume


Severity Matrix


Post-Incident Review

  • What failure mode was it? (hallucination, drift, tool failure, loop…)
  • Where was the missing checkpoint?
  • Add the scenario to your failure-injection tests (see Templates).

Next Steps