Skip to main content

Long-Running Agent Pattern

Long-running agents handle tasks that span hours, days, or weeks: monitoring pipelines, scheduled jobs, approval workflows, and research agents that run overnight. The core challenge is that infrastructure fails more often than your task is short. StateBase makes long-running agents safe by persisting every step. If your process crashes, restarts, or gets redeployed, the agent resumes from its last committed state.

Why Long-Running Agents Fail


The Core Pattern

The golden rule: checkpoint before every side effect, resume from the last checkpoint.
If the process dies at any point, restarting the function rehydrates state and skips already-completed items.

Heartbeat Checkpoints

For very long tasks, commit a “heartbeat” on a timer so a crash never loses more than N minutes of work:

Idempotency

Long-running work retries. Make every step idempotent so replay is safe:

Waiting for External Events

Long-running agents often block on humans, APIs, or schedules. Persist the pending intent so a restart can pick up where the wait ended:
When the webhook or callback arrives, rehydrate and continue.

Best Practices

  • Checkpoint before every side effect — never after a risk without a save point
  • Make steps idempotent — replay must be safe
  • Use heartbeats for very long loops
  • Persist the “intent” of waits, not just the last message
  • Monitor via Traces — sb.traces.list() gives a full audit of the run

Next Steps