Long-Running Agent Pattern
Long-running agents handle tasks that span hours, days, or weeks: monitoring pipelines, scheduled jobs, approval workflows, and research agents that run overnight. The core challenge is that infrastructure fails more often than your task is short. StateBase makes long-running agents safe by persisting every step. If your process crashes, restarts, or gets redeployed, the agent resumes from its last committed state.Why Long-Running Agents Fail
The Core Pattern
The golden rule: checkpoint before every side effect, resume from the last checkpoint.Heartbeat Checkpoints
For very long tasks, commit a “heartbeat” on a timer so a crash never loses more than N minutes of work:Idempotency
Long-running work retries. Make every step idempotent so replay is safe:Waiting for External Events
Long-running agents often block on humans, APIs, or schedules. Persist the pending intent so a restart can pick up where the wait ended:Best Practices
- Checkpoint before every side effect — never after a risk without a save point
- Make steps idempotent — replay must be safe
- Use heartbeats for very long loops
- Persist the “intent” of waits, not just the last message
- Monitor via Traces —
sb.traces.list()gives a full audit of the run
Next Steps
- Checkpoints & Rollbacks: granular recovery
- Multi-Agent Pattern: coordinate several long-running agents
- Incident Recovery: what to do when a run goes wrong