Understand first
New to this knowledge point
Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.
Start with core principles →Understand → Implement → Debug → Design
Handle the uncertain results of external side effects with stable operation flags, parameter summaries, and result queries.
Knowledge content check2026-10-03 · Check the source of the original question2026-10-02
It is recommended to understand first:
Tool calls: structure, authorization, and business contracts →Select the starting point based on the current basis, or you can go deeper one by one. When you encounter an unfamiliar concept, go back to the core principles first; use the knowledge exercises to check your understanding when you are finished.
New to this knowledge point
Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.
Start with core principles →Prepare to write the principles into code
Understand implementation steps and boundaries, complete small tasks, and check results against acceptance requirements.
View the code example →Need to handle failures and changes in conditions
Follow the continuous questioning to locate the failure premise, and then compare the migration cases to explain how the plan should be adjusted.
Continue to delve deeper into the problem →Need to design or review plans
Combine engineering deductions and senior self-evaluation standards to explain the applicable conditions, costs and alternatives of the plan.
Analyze engineering scenarios →LEARN · PRACTICE · REFLECT
First read along the principles, Q&A and migration cases. When you need to check your understanding, switch to reinforcement exercises or start personal recording.
Can be practiced directly. After logging in, answers, favorites, and notes will be saved to your account.
Log in and saveEach modified commit will be kept as an independent history. Your level of mastery is up to you to evaluate yourself against the standards.
Core concept · Idempotency and compensation when outcomes are unknown
Preparatory concepts:Network timeout, unique constraint, Business operation identification
The caller does not get the response and cannot distinguish between unexecuted and executed. Safe retry relies on the same business intent being recognized and deduplicated on the execution side; local checkpoints cannot span the atomicity gap between remote side effects and local records.
A request may never arrive, be rejected, or execute before losing its response. The caller may see an exception in each case. Retrying the last can duplicate payments or messages. Writes need an unknown-outcome state outside generic failed-request retry handling.
A UUID is suitable only if reused for every attempt of one business action. Compare stored parameters and results, rejecting one key used for different intent. Atomic constraints handle concurrent requests; checking absence before execution alone creates a race.
Stripe documents parameter comparisons and key retention for its APIs; other services can differ. Beyond retention, an old key may represent a new request, requiring business-state reconciliation. An outbox prevents lost local intent without granting exactly-once delivery to a non-idempotent mail service. Compensation is a separate action that can also fail.
Saves a prepared intent with stable business key K and parameter digest H.
Receives K/H and atomically commits the business effect and queryable receipt R.
The response is lost. The caller knows only that a timeout occurred; the outcome stays unknown.
Looks up or retries using the same K. Identical parameters return R; conflicting parameters are rejected.
Persists R before marking the outcome confirmed; verifies that the provider has only one effect.
Safe retries depend on provider deduplication/lookup, retention periods, and current authorization. If neither capability exists, pause for human verification. A local checkpoint cannot supply a remote guarantee.
A timeout means no observed response, not necessarily no execution. Reuse one business operation ID as the idempotency key, storing parameter digests and results. Identical parameters return the existing result; conflicting parameters are rejected. Use target idempotency or reconcile status before manual handling or compensation. Checkpoints and deduplication serve different purposes; local persistence alone cannot guarantee exactly one external effect.
The client makes a publish request, the server completes the publish, but the response is lost on the network. At this time, the client sees a timeout and re-executing may produce a second article. Mapping timeouts uniformly as failed will hide the fact. The running record should allow outcome_unknown, save the operation ID and remote request identification, and go through the query confirmation process. Read-only queries are generally easier to retry; whether a write request is retryable depends on its idempotent design, not on whether the model deems it worth retrying.
The key can be composed of task ID, action phase and target version, the key is that the same logical action remains unchanged across resumes and retries. Creating a new random key every time you try will lose the deduplication effect. The server creates unique constraints based on tenant and operation keys, and stores normalized parameter summaries, status, and results. If the same key and the same parameters are used, the old result will be replayed, and if the same key and different parameters are used, a conflict will be reported. Concurrent execution requires competing for the same key in a database transaction or remote atomic operation. You cannot first check whether it exists and exit the transaction before writing. Logging attempt, business recording operation, do not mix the two into one ID.
The following SQLite code puts local business writing and result storage into the same transaction to demonstrate local deduplication. For sending letters or payments, database submission and remote execution cannot be merged by this transaction. Prioritize using the remote idempotent API and pass the same business key there; otherwise, using persistent records to be sent and retryable delivery, you still need to acknowledge the risk of remote duplication. If the remote end can query by business number, verify first and then try again; high-impact actions that cannot be deduplicated or queried enter manual verification, and compensation actions must be recorded separately, and rollback cannot be pretended to be successful.
Failure is caused before execution, after business writing, and before response after submission, and the same request is sent twice. Assert that only one business record appears locally, the replay results are consistent, and a conflict occurs when parameters are modified. Concurrent scenarios verify unique constraints separately, and cross-process scenarios use a persistent database. LangGraph interrupt recovery will rerun the node, which further illustrates that the tool gateway must independently protect the side effects. Acceptance focuses on business results rather than "the function is called once"; production code also requires expiration policies, audits, permission re-verification and remote reconciliation.
Python 3 standard library demonstration, using in-memory SQLite; real recovery requires a file or server database, and external actions require remote idempotent capabilities.
import hashlib
import json
import sqlite3
db = sqlite3.connect(":memory:")
db.executescript("""
CREATE TABLE publications (
id INTEGER PRIMARY KEY AUTOINCREMENT,
tenant TEXT NOT NULL,
title TEXT NOT NULL
);
CREATE TABLE operations (
tenant TEXT NOT NULL,
op_key TEXT NOT NULL,
payload_hash TEXT NOT NULL,
result_id INTEGER NOT NULL,
PRIMARY KEY (tenant, op_key)
);
""")
def publish(tenant, op_key, title):
payload = json.dumps({"title": title}, sort_keys=True).encode()
digest = hashlib.sha256(payload).hexdigest()
db.execute("BEGIN IMMEDIATE")
try:
old = db.execute(
"SELECT payload_hash, result_id FROM operations WHERE tenant=? AND op_key=?",
(tenant, op_key),
).fetchone()
if old:
if old[0] != digest:
raise ValueError("IDEMPOTENCY_CONFLICT")
db.commit()
return old[1]
cur = db.execute(
"INSERT INTO publications(tenant,title) VALUES (?,?)", (tenant, title)
)
result_id = cur.lastrowid
db.execute("INSERT INTO operations VALUES (?,?,?,?)",
(tenant, op_key, digest, result_id))
db.commit()
return result_id
except Exception:
db.rollback()
raise
print(publish("t1", "task-7:publish:v3", "Agent recovery"))
print(publish("t1", "task-7:publish:v3", "Agent recovery"))
print("rows", db.execute("SELECT COUNT(*) FROM publications").fetchone()[0])
try:
publish("t1", "task-7:publish:v3", "Changed title")
except ValueError as exc:
print(str(exc))
db.close()
expected output
1
1
rows 1
IDEMPOTENCY_CONFLICTContinue reading along with the premises and constraints of the problem. Understand the reference answers first, then try to put away the answers and explain the cause and effect and trade-offs in your own words.
Level 1Why does it not work if I generate a new UUID every time I try again?
The identification must correspond to the business life cycle and cannot only correspond to network requests.
The new UUID will make the remote end think that this is a new business intention, so the deduplication record of the original request will not be hit. The key should be generated when the operation is first created, persisted, and used by retries, recoveries, and re-deliveries; a separate attempt_id is recorded for each network attempt. Different business actions still use different keys, and stability does not mean that the entire system reuses the same value.
Follow this answer further
Level 2The task has been suspended for two days and the remote idempotency key has expired. Can I still try again with the original key?
A stable key still depends on the lifetime of the executor's deduplication records, including recovery of long-running tasks.
Safety cannot be assumed. To check the remote retention period, first query the original action according to the business number or receipt; if it is completed, fill in the local record; if it is not completed and is reliably confirmed, submit it. Keep the local long-term ledger and associate it with the remote state, and pause for verification while the outcome is unknown. Stripe's term is its product strategy and cannot be generalized to all services.
Follow this answer further
Level 3The local ledger says that it is being executed, but I don’t know whether the original consumer is still alive. Can it be resent after seizing the lock?
Local mutual exclusion and remote deduplication are two different consistency issues.
Locks or leases only resolve who can continue processing, but cannot prove that the remote end is not executing. The new consumer first checks the same business operation and then continues according to the remote idempotent semantics; use an incrementing claim generation to prevent the old consumer from resubmitting after recovery. The execution end rejects the old version request, but the downstream still needs to face the window if the constraint is not recognized.
Level 1How to deal with the same idempotency key but different recipients?
The same key must agree with the graph, otherwise the deduplication will cause different businesses to be folded incorrectly.
Reject the request as an idempotency conflict, save the original parameter digest and clearly indicate the intended change. If the user really requests to change the recipient, first check whether the old letter has been sent, then create a new action and new key, and recheck the authorization and content; the result of the original key cannot be overwritten quietly, otherwise it will not be auditable, and the old approval may be used for the new target.
Level 1The remote end does not support idempotency and has no status query. Will you automatically retry?
When the mechanism is not available, adjust delivery semantics instead of promising non-existent guarantees.
High-impact write operations are not automatically retried, and are marked as unknown and entered for manual verification or clear business reconciliation. Scenarios with low impact and allowing duplication can accept at-least-once delivery by the predetermined policy, but the possibility of duplication must be stated and duplication prevention cannot be claimed to be successful. Decisions are based on action cost and verifiability and cannot be ad hoc optimistic judgments by the model.
First find out the conditions for change, and then determine which premises in the original plan still hold true. The following cases are teaching deductions to facilitate the transfer of principles to new problems.
Changing conditions:Writes with side effects become queries without side effects
Extended question:Do we still need idempotency keys?
It is generally not necessary to prevent query results from being duplicated, but deadlines, backoffs, and call budgets are still used to avoid amplifying the load. The query may read different versions. When checking payment, it must be judged based on the same business number, rather than treating the new query results as proof of failure of the original request. Read-only and easy to retry does not mean consistency is not important.
The principles that remain unchanged:Whether it is safe to retry depends on the operational semantics, and the timeout itself does not describe the business results.
Changing conditions:Create action becomes compensating action
Extended question:Does deletion mean it has never been published?
Not equivalent, the user may have read it and the external cache is still there. Treat revocation as a new operation, check permissions and target versions, save two receipts for release and revocation, and explain the scope of recovery to the outside world. The revocation timeout also enters unknown processing, and a compensation request is not evidence that compensation succeeded.
The principles that remain unchanged:Side effects cannot be erased from the ledger, and the results of compensation must be independently verified.
It is designed based on public technical information; the reference materials support the technical mechanism, and the scenarios and scoring standards are designed by this website and do not represent the original interview questions of a certain company. New Q&A and migration cases are added for principle explanation, and source verification and case operation verification are recorded separately.
After reading, you can explain the principles, boundaries, and trade-offs against these standards. It is up to you to evaluate your mastery; if further verification is needed, complete the small tasks below.
View verification records for independent examples
From the idempotent counterexample of two databases, continue to verify the lease, checkpoint and independent server receipt. Unzip the reliability experiment v3 and execute it in a separate directory.
Read full text and fault analysis → · Download Reliability Experiment v3 ↓
python3 cli.py submit --db crash.sqlite
python3 cli.py run --db crash.sqlite --lease-seconds 2 --fault after_collect
python3 cli.py inspect --db crash.sqlite
# 首次运行预期退出码 75;等待至少 2 秒后分别执行
python3 cli.py run --db crash.sqlite
python3 evaluate.py --db crash.sqliteFixed collect → draft → verify → publish flow; verifying local persistence protocol with real process exit does not prove that any remote service executes exactly once.
Demonstrates a recovery process when publishing is successful but the response is lost.