Agent Application DevelopmentAccount
Knowledge catalogChoose core direction and segmented content

SYSTEMATIC LEARNING / FOUR-LEVEL COURSE

Idempotency, unknown outcomes, and task recovery

Use separate caller and service databases to simulate a successful commit with a lost response, and examine what checkpoints can guarantee.

Learning objectives: Distinguish execution attempts from business operations, recover unknown outcomes using stable operation identities, and explain the limits when a remote service cannot deduplicate.

Content checked: 2026-10-04 · Each level has independent explanations, tasks and inspections

Choose a starting point based on your familiarity with this topic. Current level: Debugging · Diagnose failures. After completing the task, continue to the next level. Reading and self-checks alone do not establish mastery.

On this level

Review the prerequisites

Suitable for: It is necessary to troubleshoot duplicate publishing, recovery of lost state, or concurrent takeover.

Idempotent
When the same business operation is requested repeatedly, no additional repeated business effects will be generated; the scope of the guarantee depends on the server implementation and validity period.
operating identity
A stable identifier representing the same business intent, which is different from the identifier for each network attempt.
Checkpoint
The calculation state saved by the caller does not automatically equal the fact that the external service has occurred.
Unknown result
The caller does not get a reliable receipt and cannot assert success or failure.

How does the mechanism work?

Response loss: recover with the same operation identity
  1. Caller

    Saves a prepared intent with stable business key K and parameter digest H.

  2. Provider

    Receives K/H and atomically commits the business effect and queryable receipt R.

  3. Network

    The response is lost. The caller knows only that a timeout occurred; the outcome stays unknown.

  4. Recovering worker

    Looks up or retries using the same K. Identical parameters return R; conflicting parameters are rejected.

  5. Caller

    Persists R before marking the outcome confirmed; verifies that the provider has only one effect.

Safe retries depend on provider deduplication/lookup, retention periods, and current authorization. If neither capability exists, pause for human verification. A local checkpoint cannot supply a remote guarantee.

Step-by-step explanation
  1. persistence intent

    Save the stable operation identity and parameters before initiating the action.

  2. Submitted by service provider

    The service party executes and removes duplication with its own contract.

  3. Caller Check

    Response is lost when querying or retrying along the same identity.

  4. Update task status

    Confirm after obtaining a verifiable receipt. If it cannot be verified, the status will remain unknown.

Debugging · Diagnose failures

Check replay boundaries across failure windows

Objectives of this level: Observable invariants can be defined for different failure points.

Inject failures around the commit boundary

Failure before execution should leave no provider effect. Failure after provider commit but before caller confirmation should be reconciled using the original identity. Recovery after confirmation should create no new operation. Inspect business rows, operation identity, receipts, and caller state rather than a final completed message.

A lease does not deduplicate effects

Leases select the takeover worker, but an expired worker can still send requests. Resource-side tokens may block certain stale writes. Business deduplication identity still needs an independent design. Repeated recovery and approval must refer to the same logical action; a new lease creates no new business intent.

Preserve unknown outcomes

Without provider deduplication or status queries, unlimited retries cannot establish certainty. Retain the unknown state and evidence, then use manual reconciliation or a permitted follow-up action. Compensation can fail too and needs its own record. Requesting cancellation does not establish restoration.

Run experiments and observe counterexamples

Two local SQLite files simulate independent commits and response loss. No real remote service, process termination, or network failure is involved; this does not establish end-to-end exactly-once execution.

Python 3.10+ · Runs by default using only the standard library · Runs on your computer

  1. Observe the differences between the caller's prepared and the server's submitted
  2. Get the original receipt along the same operation identity
  3. Use the same identity to change content and verify conflicts
Downloadidempotency_recovery.py ↓
python3 idempotency_recovery.py
View the entry-point script
"""Two local SQLite files model independent caller/provider commits.

Not a real remote service, production queue or end-to-end exactly-once proof.
"""
import json
import sqlite3
import tempfile
from pathlib import Path


def provider(path, operation, payload):
    with sqlite3.connect(path) as db:
        db.execute("CREATE TABLE IF NOT EXISTS effects (operation TEXT PRIMARY KEY, payload TEXT NOT NULL, receipt TEXT NOT NULL)")
        # Serializes the read/check/write in this local demonstration.
        db.execute("BEGIN IMMEDIATE")
        existing = db.execute("SELECT payload,receipt FROM effects WHERE operation=?", (operation,)).fetchone()
        if existing:
            if existing[0] != payload:
                raise ValueError("same operation with different payload")
            return existing[1]
        receipt = "receipt:" + operation
        db.execute("INSERT INTO effects VALUES(?,?,?)", (operation, payload, receipt))
        return receipt


def demo():
    with tempfile.TemporaryDirectory() as folder:
        remote, local = Path(folder) / "provider.db", Path(folder) / "caller.db"
        with sqlite3.connect(local) as db:
            db.execute("CREATE TABLE intents(operation TEXT PRIMARY KEY, payload TEXT, status TEXT, receipt TEXT)")
            db.execute("INSERT INTO intents VALUES('publish-1','report-v1','prepared',NULL)")
        first = provider(remote, "publish-1", "report-v1")
        # Simulated response loss: provider committed, caller did not get receipt.
        with sqlite3.connect(local) as db:
            before = db.execute("SELECT status FROM intents").fetchone()[0]
        second = provider(remote, "publish-1", "report-v1")
        with sqlite3.connect(local) as db:
            db.execute("UPDATE intents SET status='confirmed', receipt=?", (second,))
        with sqlite3.connect(remote) as db:
            effects = db.execute("SELECT COUNT(*) FROM effects").fetchone()[0]
        assert before == "prepared" and first == second and effects == 1
        return dict(caller_before_recovery=before, same_receipt=first == second,
                    provider_effects=effects, caller_after_recovery="confirmed")


if __name__ == "__main__":
    print(json.dumps(demo(), sort_keys=True))

Expected output when running locally

{"caller_after_recovery": "confirmed", "caller_before_recovery": "prepared", "provider_effects": 1, "same_receipt": true}
  • There is only one business effect for the server
  • Consent diagram retry receipt consistent
  • Remote capabilities and deduplication deadlines need to be verified separately
View the running environment, output and verification records →

Continue to do advanced research experiments

After the process really exits, how can it continue?

From the idempotent counterexample of two databases, continue to verify the lease, checkpoint and independent server receipt. Unzip the reliability experiment v3 and execute it in a separate directory.

Read full text and fault analysis → · Download Reliability Experiment v3 ↓

python3 cli.py submit --db crash.sqlite
python3 cli.py run --db crash.sqlite --lease-seconds 2 --fault after_collect
python3 cli.py inspect --db crash.sqlite
# 首次运行预期退出码 75;等待至少 2 秒后分别执行
python3 cli.py run --db crash.sqlite
python3 evaluate.py --db crash.sqlite

Keep evidence and check item by item

  • First inspect shows running, collect checkpoint saved.
  • After the lease expires, succeeded, generation=2, collect is only submitted once.
  • Press README and run after_effect to check that there is only one receipt in the independent publisher database.

Fixed collect → draft → verify → publish flow; verifying local persistence protocol with real process exit does not prove that any remote service executes exactly once.

Acceptance task for this level

List the three failure points and acceptance conditions before execution, after submission by the service party, and after confirmation by the caller.

Check each item after completion

  • Assert effect quantity and receipt consistency
  • Add same-key parameter conflicts and duplicate recovery samples
  • Keep the outcome unknown when verification is unavailable; do not fabricate completion.

Save your own processes, code and results. Acceptance requirements are provided here, and course mastery status will not be automatically graded or saved at this time.

Hide the answer and check your understanding

After taking over the Worker, should I change the business operation keys?

Further explanations and practice

When encountering unfamiliar principles, first read the implementation, continuous questioning and migration cases, and then independently explain the premise and boundaries. Answers and notes are saved to the original account record.

All linked explanations and exercises (5 )

Sources and verification scope

The principles are based on public information; the numbers, cases and tasks are the teaching design of this website. Offline experiments verify the range noted on this page, and the learning effect still needs to be judged through independent tasks and feedback.