Cost vs. Safety
Every checkpoint, turn, and memory you persist costs a little money — and costs you nothing when disaster strikes. The art is deciding where to spend.The Tradeoff
There is no free lunch: you can only roll back to what you saved.
What Actually Costs Money
- State storage — every checkpoint is a version of your state
- Turn storage — every logged interaction
- Memory embeddings — semantic search indexing
- API calls — retrieving large contexts
- Store only state, not raw transcripts. If you don’t need full transcripts for auditing, persist state and a summary.
- Sliding-window turns.
get_context(turn_limit=10)reads 10 turns but state history stays intact. - Summarize old turns into memory instead of keeping them verbatim.
- TTL on caches. Don’t checkpoint throwaway data (greetings, transient lookups).
The 80/20 Rule
Most production agents get 90% of the safety for 20% of the cost by checkpointing only on commitment:- Before calling a paid/external API
- Before executing a user-visible action
- After a state mutation that later steps depend on
- When the user confirms or changes a decision
Budgeting Model
Estimate with real numbers:Recommendation
Start per-turn. Ship. Then measure how often you actually roll back. If rollbacks are rare and your context is expensive to retrieve, move to per-milestone. You can always add checkpoints later — you can’t restore what was never saved.Next Steps
- When to Checkpoint: the critical-path principle
- Rollback Strategies: how to recover
- Reliability Guarantees: what happens when things go wrong