Agent Application DevelopmentAccount
Knowledge catalogChoose core direction and segmented content
knowledge unit 30AdvancedImplementationAbout 14 minutes

Understand → Implement → Debug → Design

Checkpoints, replay, and external side-effect boundaries

Distinguish state recovery, history replay and external side effects, design stable operation keys, execution ledger and fault injection verification.

LangGraphCheckpointIdempotentrestore

Knowledge content check2026-10-03 · Check the source of the original question2026-10-02

Which step do you want to learn from this knowledge point?

Select the starting point based on the current basis, or you can go deeper one by one. When you encounter an unfamiliar concept, go back to the core principles first; use the knowledge exercises to check your understanding when you are finished.

Understand first

New to this knowledge point

Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.

Start with core principles →

Implement next

Prepare to write the principles into code

Understand implementation steps and boundaries, complete small tasks, and check results against acceptance requirements.

View the code example →

Debug failures

Need to handle failures and changes in conditions

Follow the continuous questioning to locate the failure premise, and then compare the migration cases to explain how the plan should be adjusted.

Continue to delve deeper into the problem →

Compare designs

Need to design or review plans

Combine engineering deductions and senior self-evaluation standards to explain the applicable conditions, costs and alternatives of the plan.

Analyze engineering scenarios →
Knowledge unit directory

LEARN · PRACTICE · REFLECT

Knowledge learning and personal records

My notes and review ↗

First read along the principles, Q&A and migration cases. When you need to check your understanding, switch to reinforcement exercises or start personal recording.

Answers and personal notes

Each modified commit will be kept as an independent history. Your level of mastery is up to you to evaluate yourself against the standards.

Core concept · Checkpoints, replay, and external side-effect boundaries

Understand the core principles first

Preparatory concepts:Local transactions and remote calls, Idempotent operation keys, Running state machine

Checkpoint persists the calculation state and cannot turn remote actions and local submissions into one transaction. When the result is unknown, it must be checked along with the stable operating identity, and replay cannot pretend to be read-only history.

Recovery restores computation, not the world

A saved checkpoint can restore the next node and earlier results, while tickets, emails, or payments live elsewhere. Remote success before local recording creates an unresolved window. Frequent checkpoints reduce recomputation but cannot establish the outcome of a lost remote response.

Separate operations from attempts

Persist logical intent and operation_id before calling a tool, reusing that ID on retries. Three intended tickets need three operation IDs; repeated attempts for one ticket retain one ID. Distinguish confirmed success, confirmed failure, and unknown. Reconcile unknown outcomes through the target or a unique business key. Local state alone cannot guarantee one effect if the target supports neither idempotency nor queryable receipts.

Replay executes again

LangGraph replay from a historical checkpoint re-executes subsequent nodes, including models, APIs, and interrupts. New exploratory branches need isolated write tools or explicit new authorized intents. APIs must distinguish recovery of an original action from creation of a new business action.

Durability modes change local guarantees

Synchronous persistence commits before the next step; asynchronous persistence may not finish before process exit. Exit persistence does not save every intermediate step. Choose by acceptable recomputation and cost; synchronous writes still cannot create a remote transaction. Inject faults before a call, between remote success and local commit, and after commit. Check business records and operation identities rather than final prose. These tests are instructional designs.

Check understanding with a question

How to recover after Agent crashes? Why can't Checkpoint guarantee that the tool will only be executed once?

Checkpoints persist state and recovery position without guaranteeing one external effect. Recovery or explicit historical replay can repeat nodes. Protect email and ticket writes with stable operation_id and target idempotency or reconciliation. Record intent, responses, and local commits separately. Inject crashes before calls, after remote success before local commit, and after commit; inspect business effects as well as recovery state.

Implementation and trade-offs

State recovery and historical playback solve different problems

Checkpoint must at least record the run_id, status version, submitted node results, next step location and referenced business resources. Continue from the persistent state after the process is restarted to avoid redoing already submitted calculations. LangGraph's current documentation clearly states that when specifying historical checkpoint playback, previous nodes will be skipped and subsequent nodes will be re-executed; this includes model calls, API requests and interruptions. Playback cannot be understood as playing a read-only video. Restoring the original task, creating a new branch and re-executing the business operation should be clearly distinguished at the product and interface levels.

The real difficulty is the failure window between the two systems

Assume that the tool submits the work order successfully, but the network is disconnected and the caller cannot get the response, and the local checkpoint still shows that it is not completed. Retrying directly will risk duplicating the work order; skipping it directly will not be sure whether it is really successful. The engineering solution is to generate stable keys for business operations, such as run_id plus logical node name and input version; retry the same key, and change the key only for real new operations. When the target system supports idempotency keys, repeated submissions will return the same business result; otherwise, the unique business key will be used to query and reconcile. The execution ledger must distinguish between prepared, confirmed, unknown, and failed. Unknown is not automatically regarded as failure. If the target neither provides idempotency nor can be queried, it can only wait for verification and cannot promise end-to-end exactly-once.

Persistence mode and deployment version are also recovery contracts

Synchronous checkpoint will wait for writing before entering the next step. Asynchronous mode has better throughput but there are windows that have not been written before the process exits. Exit mode cannot guarantee recovery after a crash. The choice should be combined with the task value, allowed recalculation range and storage overhead; synchronous placement still cannot turn the remote API into the same transaction. The schema_version, tool contract version and resource reference are saved in the state, and migration or incompatible versions are rejected before recovery. Do not hand over the old state directly to the new node with changed semantics. Use object references and checksums for large files to avoid duplicating the entire binary content at each step.

Acceptance requires observing business results rather than just the final answer

Set three failure points: before the tool is called, the remote end succeeds but the local result is not written, and after the result is written. Restart and recover at each point, check business operation keys, target record number, checkpoint and final task status. Also covers duplicate restore requests, concurrent restores of the same run, and revocation of permissions during restore. The following SQLite example only demonstrates the business unique key and status submission within the local transaction; the remote system must provide additional idempotency or reconciliation capabilities.

code example

Simulate repeated recovery with unique action key

The Python standard library works. Business side effects are in the same local database transaction; it does not mean that email, payment or remote API can be submitted atomically with checkpoint.

import sqlite3, tempfile
from pathlib import Path
with tempfile.TemporaryDirectory() as folder:
    path = Path(folder) / 'run.db'
    db = sqlite3.connect(path)
    db.executescript('CREATE TABLE effects(op TEXT PRIMARY KEY, result TEXT); CREATE TABLE checkpoints(run TEXT PRIMARY KEY, state TEXT);')
    for attempt in range(2):
        with db:
            db.execute('INSERT OR IGNORE INTO effects VALUES (?, ?)', ('run-1:create-ticket:v1', 'ticket-42'))
            db.execute('INSERT OR REPLACE INTO checkpoints VALUES (?, ?)', ('run-1', 'done'))
        db.close()
        db = sqlite3.connect(path)
    print(db.execute('SELECT COUNT(*) FROM effects').fetchone()[0])
    print(db.execute('SELECT state FROM checkpoints').fetchone()[0])
    db.close()

expected output

1
done

Engineering deduction

scene
Hypothetical engineering scenario: The research agent crashes after creating a task ticket.
design decisions
The combination of run and logical operations is used as the idempotency key of the work order. Unknown responses first check the target system and then restore the node.
Verify target
The acceptance goal is to still locate the same work order after repeated recovery and continue to generate reports; this is a design acceptance requirement and does not claim the real company effect.
applicable boundary
When the target system does not have an idempotent or query interface, manual verification is required; checkpoint itself cannot supplement this capability.

Continuous questions and answers

Continue reading along with the premises and constraints of the problem. Understand the reference answers first, then try to put away the answers and explain the cause and effect and trade-offs in your own words.

Draw inferences from one example: If the conditions change, how to deduce it?

First find out the conditions for change, and then determine which premises in the original plan still hold true. The following cases are teaching deductions to facilitate the transfer of principles to new problems.

Read-only research tasks

Changing conditions:The tool does not generate business writes, but the data may change

Extended question:Is replaying completely risk-free?

Derivation and reference solutions

There is no risk of write side effects, but the changed data will be rechecked, tokens will be spent, or rate limiting will be triggered, and the results may not be consistent with those at the time. Save source versions or snapshots and model configurations when reproduction is needed; explicitly use new data when current answers are needed. Just because a computation is repeatable does not mean that the input world remains unchanged.

The principles that remain unchanged:Checkpoint saves the calculation state, and the external world and costs still have independent life cycles.

The remote end does not support idempotency

Changing conditions:There is no idempotency key in the write API, but you can query by business number

Extended question:Can one still design recoverable processes?

Derivation and reference solutions

First persist the stable business number, query whether it exists after timeout, and record the receipt after confirmation. If the query is not found, visibility delay must also be considered, and wait according to the target consistency contract. If the query cannot be uniquely identified or the result is unknown for a long time, pause for manual verification and local replay cannot ensure non-duplication.

The principles that remain unchanged:Recovery must prove external results, and an unknown state cannot automatically equate to failure.

Easy to make mistakes

  • Treat checkpoints as external side-effect transactions
  • Generate new random idempotency keys with each retry
  • Use a memory saver to claim that the process can be restored after restarting
  • Ignore current permissions and code version when replaying old state

References

It is designed based on public technical information; the reference materials support the technical mechanism, and the scenarios and scoring standards are designed by this website and do not represent the original interview questions of a certain company. New Q&A and migration cases are added for principle explanation, and source verification and case operation verification are recorded separately.

Check how far you understand

After reading, you can explain the principles, boundaries, and trade-offs against these standards. It is up to you to evaluate your mastery; if further verification is needed, complete the small tasks below.

Basic standards met
Distinguish between saved state and external side-effect guarantees.
Intermediate and advanced signals
Clarify the recovery starting point, completion record and idempotent operation identity.
Senior criteria
Prove replay boundaries and handle unknown results through downtime before and after submission.

View verification records for independent examples

Continue to do advanced research experiments

After the process really exits, how can it continue?

From the idempotent counterexample of two databases, continue to verify the lease, checkpoint and independent server receipt. Unzip the reliability experiment v3 and execute it in a separate directory.

Read full text and fault analysis → · Download Reliability Experiment v3 ↓

python3 cli.py submit --db crash.sqlite
python3 cli.py run --db crash.sqlite --lease-seconds 2 --fault after_collect
python3 cli.py inspect --db crash.sqlite
# 首次运行预期退出码 75;等待至少 2 秒后分别执行
python3 cli.py run --db crash.sqlite
python3 evaluate.py --db crash.sqlite

Keep evidence and check item by item

  • First inspect shows running, collect checkpoint saved.
  • After the lease expires, succeeded, generation=2, collect is only submitted once.
  • Press README and run after_effect to check that there is only one receipt in the independent publisher database.

Fixed collect → draft → verify → publish flow; verifying local persistence protocol with real process exit does not prove that any remote service executes exactly once.

Hands-on verificationComplete on demand · Suggestions15 minutes

After the tool is submitted and before the Checkpoint is saved, the system crashes and the deduction resumes.

Expand acceptance requirements and checkpoints
  • External facts are not lost
  • Repeat requests separate from repeat effects
  • Restore the settled budget without repeated consumption

Key inspections

  • Able to distinguish graph status, execution logs and external business status
  • Ability to describe remote success but locally unrecorded failure windows
  • Can explain stable idempotency keys, query reconciliation, and manual intervention when idempotency is not possible
  • Ability to design recovery-compatible state versions and fault injection acceptance