Agent Application DevelopmentAccount
Knowledge catalogChoose core direction and segmented content

SYSTEMATIC LEARNING / FOUR-LEVEL COURSE

Idempotency, unknown outcomes, and task recovery

Use separate caller and service databases to simulate a successful commit with a lost response, and examine what checkpoints can guarantee.

Learning objectives: Distinguish execution attempts from business operations, recover unknown outcomes using stable operation identities, and explain the limits when a remote service cannot deduplicate.

Content checked: 2026-10-04 · Each level has independent explanations, tasks and inspections

Choose a starting point based on your familiarity with this topic. Current level: Design · Explain the trade-offs. After completing the task, continue to the next level. Reading and self-checks alone do not establish mastery.

On this level

Review the prerequisites

Suitable for: Cross-system task recovery and approval need to be designed.

Idempotent
When the same business operation is requested repeatedly, no additional repeated business effects will be generated; the scope of the guarantee depends on the server implementation and validity period.
operating identity
A stable identifier representing the same business intent, which is different from the identifier for each network attempt.
Checkpoint
The calculation state saved by the caller does not automatically equal the fact that the external service has occurred.
Unknown result
The caller does not get a reliable receipt and cannot assert success or failure.

How does the mechanism work?

  1. persistence intent

    Save the stable operation identity and parameters before initiating the action.

  2. Submitted by service provider

    The service party executes and removes duplication with its own contract.

  3. Caller Check

    Response is lost when querying or retrying along the same identity.

  4. Update task status

    Confirm after obtaining a verifiable receipt. If it cannot be verified, the status will remain unknown.

Design · Explain the trade-offs

Choose recovery guarantees based on the remote service capabilities

Objectives of this level: Able to write out the guarantee scope, deduplication validity period, verification path and manual upgrade conditions.

Inventory provider guarantees

A stable idempotency identity permits retries only within that service's contract. Business-ID queries allow reconciliation. A service with neither requires acceptance of unknown outcomes and duplicate risk. A local state machine cannot establish end-to-end exactly-once for every provider.

Define identity scope and lifetime

Include the appropriate tenant and business scope. Use an argument digest to detect changed intent. Deduplication retention must cover allowed retries and recovery: three days of approval waiting with one day of provider retention leaves a gap. Define status queries, rejection, or manual verification after expiry in advance.

Treat recovery as a current decision

Recheck permissions, approved content, tool versions, and resource state. Earlier approval does not authorize changed content, and old permissions may no longer apply. Account for durable ledgers, reconciliation requests, conflicts, and manual work. Choose mechanisms according to the loss and reversibility of side effects.

Run experiments and observe counterexamples

Two local SQLite files simulate independent submission and response loss; there is no real remote end, process termination or network failure, and it does not prove end-to-end exactly-once.

Python 3.10+ · Runs by default using only the standard library · Runs on your computer

  1. Observe the differences between the caller's prepared and the server's submitted
  2. Get the original receipt along the same operation identity
  3. Use the same identity to change content and verify conflicts
Downloadidempotency_recovery.py ↓
python3 idempotency_recovery.py
View the entry-point script
"""Two local SQLite files model independent caller/provider commits.

Not a real remote service, production queue or end-to-end exactly-once proof.
"""
import json
import sqlite3
import tempfile
from pathlib import Path


def provider(path, operation, payload):
    with sqlite3.connect(path) as db:
        db.execute("CREATE TABLE IF NOT EXISTS effects (operation TEXT PRIMARY KEY, payload TEXT NOT NULL, receipt TEXT NOT NULL)")
        # Serializes the read/check/write in this local demonstration.
        db.execute("BEGIN IMMEDIATE")
        existing = db.execute("SELECT payload,receipt FROM effects WHERE operation=?", (operation,)).fetchone()
        if existing:
            if existing[0] != payload:
                raise ValueError("same operation with different payload")
            return existing[1]
        receipt = "receipt:" + operation
        db.execute("INSERT INTO effects VALUES(?,?,?)", (operation, payload, receipt))
        return receipt


def demo():
    with tempfile.TemporaryDirectory() as folder:
        remote, local = Path(folder) / "provider.db", Path(folder) / "caller.db"
        with sqlite3.connect(local) as db:
            db.execute("CREATE TABLE intents(operation TEXT PRIMARY KEY, payload TEXT, status TEXT, receipt TEXT)")
            db.execute("INSERT INTO intents VALUES('publish-1','report-v1','prepared',NULL)")
        first = provider(remote, "publish-1", "report-v1")
        # Simulated response loss: provider committed, caller did not get receipt.
        with sqlite3.connect(local) as db:
            before = db.execute("SELECT status FROM intents").fetchone()[0]
        second = provider(remote, "publish-1", "report-v1")
        with sqlite3.connect(local) as db:
            db.execute("UPDATE intents SET status='confirmed', receipt=?", (second,))
        with sqlite3.connect(remote) as db:
            effects = db.execute("SELECT COUNT(*) FROM effects").fetchone()[0]
        assert before == "prepared" and first == second and effects == 1
        return dict(caller_before_recovery=before, same_receipt=first == second,
                    provider_effects=effects, caller_after_recovery="confirmed")


if __name__ == "__main__":
    print(json.dumps(demo(), sort_keys=True))

Expected output when running locally

{"caller_after_recovery": "confirmed", "caller_before_recovery": "prepared", "provider_effects": 1, "same_receipt": true}
  • There is only one business effect for the server
  • Consent diagram retry receipt consistent
  • Remote capabilities and deduplication deadlines need to be verified separately
View the running environment, output and verification records →

Continue to do advanced research experiments

After the process really exits, how can it continue?

From the idempotent counterexample of two databases, continue to verify the lease, checkpoint and independent server receipt. Unzip the reliability experiment v3 and execute it in a separate directory.

Read full text and fault analysis → · Download Reliability Experiment v3 ↓

python3 cli.py submit --db crash.sqlite
python3 cli.py run --db crash.sqlite --lease-seconds 2 --fault after_collect
python3 cli.py inspect --db crash.sqlite
# 首次运行预期退出码 75;等待至少 2 秒后分别执行
python3 cli.py run --db crash.sqlite
python3 evaluate.py --db crash.sqlite

Keep evidence and check item by item

  • First inspect shows running, collect checkpoint saved.
  • After the lease expires, succeeded, generation=2, collect is only submitted once.
  • Press README and run after_effect to check that there is only one receipt in the independent publisher database.

Fixed collect → draft → verify → publish flow; verifying local persistence protocol with real process exit does not prove that any remote service executes exactly once.

Acceptance task for this level

Write separate recovery plans for report publishing and non-queryable email services to clarify what they can guarantee.

Check each item after completion

  • Distinguish between remote deduplication, query and neither capabilities.
  • Explain deduplication deadlines, identity conflicts and recovery permissions
  • Write down the final state after manual verification and compensation failure

Save your own processes, code and results. Acceptance requirements are provided here, and course mastery status will not be automatically graded or saved at this time.

Hide the answer and check your understanding

Why can't the experiment in this lesson prove that any third-party interface is only executed once?

Further explanations and practice

When encountering unfamiliar principles, first read the implementation, continuous questioning and migration cases, and then independently explain the premise and boundaries. Answers and notes are saved to the original account record.

All linked explanations and exercises (5 )

Sources and verification scope

The principles are based on public information; the numbers, cases and tasks are the teaching design of this website. Offline experiments verify the range noted on this page, and the learning effect still needs to be judged through independent tasks and feedback.