Agent Application DevelopmentAccount
Knowledge catalogChoose core direction and segmented content
knowledge unit 11AdvancedImplementationAbout 15 minutes

Understand → Implement → Debug → Design

Idempotency and compensation when outcomes are unknown

Handle the uncertain results of external side effects with stable operation flags, parameter summaries, and result queries.

Idempotentside effectsTry againNot sure about the result

Knowledge content check2026-10-03 · Check the source of the original question2026-10-02

Which step do you want to learn from this knowledge point?

Select the starting point based on the current basis, or you can go deeper one by one. When you encounter an unfamiliar concept, go back to the core principles first; use the knowledge exercises to check your understanding when you are finished.

Understand first

New to this knowledge point

Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.

Start with core principles →

Implement next

Prepare to write the principles into code

Understand implementation steps and boundaries, complete small tasks, and check results against acceptance requirements.

View the code example →

Debug failures

Need to handle failures and changes in conditions

Follow the continuous questioning to locate the failure premise, and then compare the migration cases to explain how the plan should be adjusted.

Continue to delve deeper into the problem →

Compare designs

Need to design or review plans

Combine engineering deductions and senior self-evaluation standards to explain the applicable conditions, costs and alternatives of the plan.

Analyze engineering scenarios →
Knowledge unit directory

LEARN · PRACTICE · REFLECT

Knowledge learning and personal records

My notes and review ↗

First read along the principles, Q&A and migration cases. When you need to check your understanding, switch to reinforcement exercises or start personal recording.

Answers and personal notes

Each modified commit will be kept as an independent history. Your level of mastery is up to you to evaluate yourself against the standards.

Core concept · Idempotency and compensation when outcomes are unknown

Understand the core principles first

Preparatory concepts:Network timeout, unique constraint, Business operation identification

The caller does not get the response and cannot distinguish between unexecuted and executed. Safe retry relies on the same business intent being recognized and deduplicated on the execution side; local checkpoints cannot span the atomicity gap between remote side effects and local records.

Distinguish three possible outcomes

A request may never arrive, be rejected, or execute before losing its response. The caller may see an exception in each case. Retrying the last can duplicate payments or messages. Writes need an unknown-outcome state outside generic failed-request retry handling.

Keep idempotency keys stable

A UUID is suitable only if reused for every attempt of one business action. Compare stored parameters and results, rejecting one key used for different intent. Atomic constraints handle concurrent requests; checking absence before execution alone creates a race.

Define retention and remote limits

Stripe documents parameter comparisons and key retention for its APIs; other services can differ. Beyond retention, an old key may represent a new request, requiring business-state reconciliation. An outbox prevents lost local intent without granting exactly-once delivery to a non-idempotent mail service. Compensation is a separate action that can also fail.

Check understanding with a question

The tool retries after timeout. How to prevent repeated sending of letters, repeated orders or repeated publishing?

A timeout means no observed response, not necessarily no execution. Reuse one business operation ID as the idempotency key, storing parameter digests and results. Identical parameters return the existing result; conflicting parameters are rejected. Use target idempotency or reconcile status before manual handling or compensation. Checkpoints and deduplication serve different purposes; local persistence alone cannot guarantee exactly one external effect.

Implementation and trade-offs

Why does normal retry cause errors?

The client makes a publish request, the server completes the publish, but the response is lost on the network. At this time, the client sees a timeout and re-executing may produce a second article. Mapping timeouts uniformly as failed will hide the fact. The running record should allow outcome_unknown, save the operation ID and remote request identification, and go through the query confirmation process. Read-only queries are generally easier to retry; whether a write request is retryable depends on its idempotent design, not on whether the model deems it worth retrying.

Idempotent keys must express business intent

The key can be composed of task ID, action phase and target version, the key is that the same logical action remains unchanged across resumes and retries. Creating a new random key every time you try will lose the deduplication effect. The server creates unique constraints based on tenant and operation keys, and stores normalized parameter summaries, status, and results. If the same key and the same parameters are used, the old result will be replayed, and if the same key and different parameters are used, a conflict will be reported. Concurrent execution requires competing for the same key in a database transaction or remote atomic operation. You cannot first check whether it exists and exit the transaction before writing. Logging attempt, business recording operation, do not mix the two into one ID.

How to handle cross-system boundaries

The following SQLite code puts local business writing and result storage into the same transaction to demonstrate local deduplication. For sending letters or payments, database submission and remote execution cannot be merged by this transaction. Prioritize using the remote idempotent API and pass the same business key there; otherwise, using persistent records to be sent and retryable delivery, you still need to acknowledge the risk of remote duplication. If the remote end can query by business number, verify first and then try again; high-impact actions that cannot be deduplicated or queried enter manual verification, and compensation actions must be recorded separately, and rollback cannot be pretended to be successful.

Verify with fault injection

Failure is caused before execution, after business writing, and before response after submission, and the same request is sent twice. Assert that only one business record appears locally, the replay results are consistent, and a conflict occurs when parameters are modified. Concurrent scenarios verify unique constraints separately, and cross-process scenarios use a persistent database. LangGraph interrupt recovery will rerun the node, which further illustrates that the tool gateway must independently protect the side effects. Acceptance focuses on business results rather than "the function is called once"; production code also requires expiration policies, audits, permission re-verification and remote reconciliation.

code example

Deduplication of business operations within SQLite transactions

Python 3 standard library demonstration, using in-memory SQLite; real recovery requires a file or server database, and external actions require remote idempotent capabilities.

import hashlib
import json
import sqlite3

db = sqlite3.connect(":memory:")
db.executescript("""
CREATE TABLE publications (
    id INTEGER PRIMARY KEY AUTOINCREMENT,
    tenant TEXT NOT NULL,
    title TEXT NOT NULL
);
CREATE TABLE operations (
    tenant TEXT NOT NULL,
    op_key TEXT NOT NULL,
    payload_hash TEXT NOT NULL,
    result_id INTEGER NOT NULL,
    PRIMARY KEY (tenant, op_key)
);
""")

def publish(tenant, op_key, title):
    payload = json.dumps({"title": title}, sort_keys=True).encode()
    digest = hashlib.sha256(payload).hexdigest()
    db.execute("BEGIN IMMEDIATE")
    try:
        old = db.execute(
            "SELECT payload_hash, result_id FROM operations WHERE tenant=? AND op_key=?",
            (tenant, op_key),
        ).fetchone()
        if old:
            if old[0] != digest:
                raise ValueError("IDEMPOTENCY_CONFLICT")
            db.commit()
            return old[1]
        cur = db.execute(
            "INSERT INTO publications(tenant,title) VALUES (?,?)", (tenant, title)
        )
        result_id = cur.lastrowid
        db.execute("INSERT INTO operations VALUES (?,?,?,?)",
                   (tenant, op_key, digest, result_id))
        db.commit()
        return result_id
    except Exception:
        db.rollback()
        raise

print(publish("t1", "task-7:publish:v3", "Agent recovery"))
print(publish("t1", "task-7:publish:v3", "Agent recovery"))
print("rows", db.execute("SELECT COUNT(*) FROM publications").fetchone()[0])
try:
    publish("t1", "task-7:publish:v3", "Changed title")
except ValueError as exc:
    print(str(exc))
db.close()

expected output

1
1
rows 1
IDEMPOTENCY_CONFLICT

Engineering deduction

scene
Assume an engineering scenario: the article publishing API times out, and the Agent re-requests publishing after recovery.
design decisions
The operation key is bound to the task and draft version; retrying with the same key returns the original release result; parameter changes require new approval and new operations.
Verify target
Demonstration results: Two identical calls return the same ID, the database has only one publishing record, and different parameters are rejected.
applicable boundary
The code only proves deduplication within a single database transaction; it does not prove exactly-once external network actions, nor does it implement distributed delivery.

Continuous questions and answers

Continue reading along with the premises and constraints of the problem. Understand the reference answers first, then try to put away the answers and explain the cause and effect and trade-offs in your own words.

Draw inferences from one example: If the conditions change, how to deduce it?

First find out the conditions for change, and then determine which premises in the original plan still hold true. The following cases are teaching deductions to facilitate the transfer of principles to new problems.

Read-only query timeout

Changing conditions:Writes with side effects become queries without side effects

Extended question:Do we still need idempotency keys?

Derivation and reference solutions

It is generally not necessary to prevent query results from being duplicated, but deadlines, backoffs, and call budgets are still used to avoid amplifying the load. The query may read different versions. When checking payment, it must be judged based on the same business number, rather than treating the new query results as proof of failure of the original request. Read-only and easy to retry does not mean consistency is not important.

The principles that remain unchanged:Whether it is safe to retry depends on the operational semantics, and the timeout itself does not describe the business results.

Revoke published content

Changing conditions:Create action becomes compensating action

Extended question:Does deletion mean it has never been published?

Derivation and reference solutions

Not equivalent, the user may have read it and the external cache is still there. Treat revocation as a new operation, check permissions and target versions, save two receipts for release and revocation, and explain the scope of recovery to the outside world. The revocation timeout also enters unknown processing, and a compensation request is not evidence that compensation succeeded.

The principles that remain unchanged:Side effects cannot be erased from the ledger, and the results of compensation must be independently verified.

Easy to make mistakes

  • Treat the local timeout as if the remote end has not been executed
  • Check first and then write but without transaction and unique constraints
  • Think that saving checkpoints can ensure that external side effects only happen once

References

It is designed based on public technical information; the reference materials support the technical mechanism, and the scenarios and scoring standards are designed by this website and do not represent the original interview questions of a certain company. New Q&A and migration cases are added for principle explanation, and source verification and case operation verification are recorded separately.

Check how far you understand

After reading, you can explain the principles, boundaries, and trade-offs against these standards. It is up to you to evaluate your mastery; if further verification is needed, complete the small tasks below.

Basic standards met
Know that timeouts may occur after external commits.
Intermediate and advanced signals
Use stable business idempotency keys, intent ledgers, and receipt reconciliation.
Senior criteria
Unknown final states, manual processing and fault injection when idempotent are explicitly not supported.

View verification records for independent examples

Continue to do advanced research experiments

After the process really exits, how can it continue?

From the idempotent counterexample of two databases, continue to verify the lease, checkpoint and independent server receipt. Unzip the reliability experiment v3 and execute it in a separate directory.

Read full text and fault analysis → · Download Reliability Experiment v3 ↓

python3 cli.py submit --db crash.sqlite
python3 cli.py run --db crash.sqlite --lease-seconds 2 --fault after_collect
python3 cli.py inspect --db crash.sqlite
# 首次运行预期退出码 75;等待至少 2 秒后分别执行
python3 cli.py run --db crash.sqlite
python3 evaluate.py --db crash.sqlite

Keep evidence and check item by item

  • First inspect shows running, collect checkpoint saved.
  • After the lease expires, succeeded, generation=2, collect is only submitted once.
  • Press README and run after_effect to check that there is only one receipt in the independent publisher database.

Fixed collect → draft → verify → publish flow; verifying local persistence protocol with real process exit does not prove that any remote service executes exactly once.

Hands-on verificationComplete on demand · Suggestions15 minutes

Demonstrates a recovery process when publishing is successful but the response is lost.

Expand acceptance requirements and checkpoints
  • Don’t blindly re-post
  • Business key stable across retries
  • Not declared successful or not executed when it cannot be verified

Key inspections

  • Distinguish between unexecuted, executed and unknown results
  • Duplicate requests with the same key and different parameters with the same key are clearly handled
  • It can explain that local transactions cannot guarantee that remote side effects are exactly once