Understand first
New to this knowledge point
Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.
Start with core principles →Understand → Implement → Debug → Design
Distinguish state recovery, history replay and external side effects, design stable operation keys, execution ledger and fault injection verification.
Knowledge content check2026-10-03 · Check the source of the original question2026-10-02
It is recommended to understand first:
Tool calls: structure, authorization, and business contracts →Select the starting point based on the current basis, or you can go deeper one by one. When you encounter an unfamiliar concept, go back to the core principles first; use the knowledge exercises to check your understanding when you are finished.
New to this knowledge point
Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.
Start with core principles →Prepare to write the principles into code
Understand implementation steps and boundaries, complete small tasks, and check results against acceptance requirements.
View the code example →Need to handle failures and changes in conditions
Follow the continuous questioning to locate the failure premise, and then compare the migration cases to explain how the plan should be adjusted.
Continue to delve deeper into the problem →Need to design or review plans
Combine engineering deductions and senior self-evaluation standards to explain the applicable conditions, costs and alternatives of the plan.
Analyze engineering scenarios →LEARN · PRACTICE · REFLECT
First read along the principles, Q&A and migration cases. When you need to check your understanding, switch to reinforcement exercises or start personal recording.
Can be practiced directly. After logging in, answers, favorites, and notes will be saved to your account.
Log in and saveEach modified commit will be kept as an independent history. Your level of mastery is up to you to evaluate yourself against the standards.
Core concept · Checkpoints, replay, and external side-effect boundaries
Preparatory concepts:Local transactions and remote calls, Idempotent operation keys, Running state machine
Checkpoint persists the calculation state and cannot turn remote actions and local submissions into one transaction. When the result is unknown, it must be checked along with the stable operating identity, and replay cannot pretend to be read-only history.
A saved checkpoint can restore the next node and earlier results, while tickets, emails, or payments live elsewhere. Remote success before local recording creates an unresolved window. Frequent checkpoints reduce recomputation but cannot establish the outcome of a lost remote response.
Persist logical intent and operation_id before calling a tool, reusing that ID on retries. Three intended tickets need three operation IDs; repeated attempts for one ticket retain one ID. Distinguish confirmed success, confirmed failure, and unknown. Reconcile unknown outcomes through the target or a unique business key. Local state alone cannot guarantee one effect if the target supports neither idempotency nor queryable receipts.
LangGraph replay from a historical checkpoint re-executes subsequent nodes, including models, APIs, and interrupts. New exploratory branches need isolated write tools or explicit new authorized intents. APIs must distinguish recovery of an original action from creation of a new business action.
Synchronous persistence commits before the next step; asynchronous persistence may not finish before process exit. Exit persistence does not save every intermediate step. Choose by acceptable recomputation and cost; synchronous writes still cannot create a remote transaction. Inject faults before a call, between remote success and local commit, and after commit. Check business records and operation identities rather than final prose. These tests are instructional designs.
Checkpoints persist state and recovery position without guaranteeing one external effect. Recovery or explicit historical replay can repeat nodes. Protect email and ticket writes with stable operation_id and target idempotency or reconciliation. Record intent, responses, and local commits separately. Inject crashes before calls, after remote success before local commit, and after commit; inspect business effects as well as recovery state.
Checkpoint must at least record the run_id, status version, submitted node results, next step location and referenced business resources. Continue from the persistent state after the process is restarted to avoid redoing already submitted calculations. LangGraph's current documentation clearly states that when specifying historical checkpoint playback, previous nodes will be skipped and subsequent nodes will be re-executed; this includes model calls, API requests and interruptions. Playback cannot be understood as playing a read-only video. Restoring the original task, creating a new branch and re-executing the business operation should be clearly distinguished at the product and interface levels.
Assume that the tool submits the work order successfully, but the network is disconnected and the caller cannot get the response, and the local checkpoint still shows that it is not completed. Retrying directly will risk duplicating the work order; skipping it directly will not be sure whether it is really successful. The engineering solution is to generate stable keys for business operations, such as run_id plus logical node name and input version; retry the same key, and change the key only for real new operations. When the target system supports idempotency keys, repeated submissions will return the same business result; otherwise, the unique business key will be used to query and reconcile. The execution ledger must distinguish between prepared, confirmed, unknown, and failed. Unknown is not automatically regarded as failure. If the target neither provides idempotency nor can be queried, it can only wait for verification and cannot promise end-to-end exactly-once.
Synchronous checkpoint will wait for writing before entering the next step. Asynchronous mode has better throughput but there are windows that have not been written before the process exits. Exit mode cannot guarantee recovery after a crash. The choice should be combined with the task value, allowed recalculation range and storage overhead; synchronous placement still cannot turn the remote API into the same transaction. The schema_version, tool contract version and resource reference are saved in the state, and migration or incompatible versions are rejected before recovery. Do not hand over the old state directly to the new node with changed semantics. Use object references and checksums for large files to avoid duplicating the entire binary content at each step.
Set three failure points: before the tool is called, the remote end succeeds but the local result is not written, and after the result is written. Restart and recover at each point, check business operation keys, target record number, checkpoint and final task status. Also covers duplicate restore requests, concurrent restores of the same run, and revocation of permissions during restore. The following SQLite example only demonstrates the business unique key and status submission within the local transaction; the remote system must provide additional idempotency or reconciliation capabilities.
The Python standard library works. Business side effects are in the same local database transaction; it does not mean that email, payment or remote API can be submitted atomically with checkpoint.
import sqlite3, tempfile
from pathlib import Path
with tempfile.TemporaryDirectory() as folder:
path = Path(folder) / 'run.db'
db = sqlite3.connect(path)
db.executescript('CREATE TABLE effects(op TEXT PRIMARY KEY, result TEXT); CREATE TABLE checkpoints(run TEXT PRIMARY KEY, state TEXT);')
for attempt in range(2):
with db:
db.execute('INSERT OR IGNORE INTO effects VALUES (?, ?)', ('run-1:create-ticket:v1', 'ticket-42'))
db.execute('INSERT OR REPLACE INTO checkpoints VALUES (?, ?)', ('run-1', 'done'))
db.close()
db = sqlite3.connect(path)
print(db.execute('SELECT COUNT(*) FROM effects').fetchone()[0])
print(db.execute('SELECT state FROM checkpoints').fetchone()[0])
db.close()
expected output
1
doneContinue reading along with the premises and constraints of the problem. Understand the reference answers first, then try to put away the answers and explain the cause and effect and trade-offs in your own words.
Level 1When the same node is cycled three times, how to distinguish three legal operations from retries of one operation?
After node recovery encounters a cycle, the idempotent identity must express the difference between legal multiple attempts and the same retry.
Persist the logical loop generation as intent_generation, for example, use run, node, business item ID, input revision and the generation to identify the intent_id; three legal operations of the same business item and the same input version still need three generations. A legal operation corresponds to a stable operation_id, and the attempt_id also records the number of retries: the next legitimate generation is incremented, and network retries and recovery use the original generation. You cannot just use the node name, nor can you generate random keys every time the process resumes. Generation allocation and intent submission are completed atomically, preventing two Workers from independently inventing the next operation.
Follow this answer further
Level 2The loop count is in memory and will be reset to zero after a machine crash. What will be the consequences?
After the parent asks to establish an operational identity, the persistence of the identity source becomes a new point of failure.
Counting reuse may treat the second valid operation as the first retry, or regenerate the key causing duplication. The cycle generation and business item identity should be persisted before the tool is called, and only the committed intent will be read during recovery; if there is no committed intent, it will be written and created once based on conditions. The number of business operations cannot be deduced from the number of recoveries.
Follow this answer further
Level 3Both Workers read the next generation as 2, how to avoid creating two intents?
The persistence generation also faces concurrent allocation and needs to be protected by local uniqueness and remote idempotency.
Use unique constraints and conditions to update the database, atomically assign logical generations and write intentions; only the worker whose claim succeeds will execute. The second Worker reads the existing intent and reconciles it without adding a new key. External tools still need to be idempotent by the same operation_id, since old Workers with expired leases may arrive late to execute.
Level 1The tool has been successful but the local database cannot be written. How can I enter unknown and make the reconciliation?
Local failure after remote success requires acknowledgment that the record itself may fail.
If the local database cannot be written to, do not claim that the unknown outcome has been persisted there. The prepared intent committed before the call provides a recovery clue: treat every intent without a confirmed receipt as potentially unknown, stop new writes, and wait for storage to recover before reconciliation. An independent reliable log can supplement this record, but an in-memory log cannot replace it. Retain the operation key even when the remote outcome is known; never retry blindly.
Level 1The old state cannot be deserialized after the upgrade. How to choose between rollback and migration?
Restoring the contract not only relies on placing the order, but also relies on the code to continue to understand the original state.
First, pause the affected recovery and determine whether it can be safely migrated based on the state_schema and workflow version. Deterministic migration retains the original snapshot, operation ledger, and approval constraints; tasks with ambiguous semantics are completed by the old executor or processed manually. Before rolling back, confirm that the new state can be read by the old code. Rolling back the code cannot reverse the side effects that have occurred.
First find out the conditions for change, and then determine which premises in the original plan still hold true. The following cases are teaching deductions to facilitate the transfer of principles to new problems.
Changing conditions:The tool does not generate business writes, but the data may change
Extended question:Is replaying completely risk-free?
There is no risk of write side effects, but the changed data will be rechecked, tokens will be spent, or rate limiting will be triggered, and the results may not be consistent with those at the time. Save source versions or snapshots and model configurations when reproduction is needed; explicitly use new data when current answers are needed. Just because a computation is repeatable does not mean that the input world remains unchanged.
The principles that remain unchanged:Checkpoint saves the calculation state, and the external world and costs still have independent life cycles.
Changing conditions:There is no idempotency key in the write API, but you can query by business number
Extended question:Can one still design recoverable processes?
First persist the stable business number, query whether it exists after timeout, and record the receipt after confirmation. If the query is not found, visibility delay must also be considered, and wait according to the target consistency contract. If the query cannot be uniquely identified or the result is unknown for a long time, pause for manual verification and local replay cannot ensure non-duplication.
The principles that remain unchanged:Recovery must prove external results, and an unknown state cannot automatically equate to failure.
It is designed based on public technical information; the reference materials support the technical mechanism, and the scenarios and scoring standards are designed by this website and do not represent the original interview questions of a certain company. New Q&A and migration cases are added for principle explanation, and source verification and case operation verification are recorded separately.
After reading, you can explain the principles, boundaries, and trade-offs against these standards. It is up to you to evaluate your mastery; if further verification is needed, complete the small tasks below.
View verification records for independent examples
From the idempotent counterexample of two databases, continue to verify the lease, checkpoint and independent server receipt. Unzip the reliability experiment v3 and execute it in a separate directory.
Read full text and fault analysis → · Download Reliability Experiment v3 ↓
python3 cli.py submit --db crash.sqlite
python3 cli.py run --db crash.sqlite --lease-seconds 2 --fault after_collect
python3 cli.py inspect --db crash.sqlite
# 首次运行预期退出码 75;等待至少 2 秒后分别执行
python3 cli.py run --db crash.sqlite
python3 evaluate.py --db crash.sqliteFixed collect → draft → verify → publish flow; verifying local persistence protocol with real process exit does not prove that any remote service executes exactly once.
After the tool is submitted and before the Checkpoint is saved, the system crashes and the deduction resumes.