01 · Why one retry can create two reports
Suppose you are building an internal technical research assistant. It reads two approved sources, generates a summary with source IDs, and saves it through a report service. A task may take tens of minutes; scheduling messages can arrive more than once, and model requests can time out. The business requirements are to retain completed work, explain unknown outcomes, and avoid saving the same report twice.
| Time | Worker A's action | Established fact |
|---|---|---|
| T0 | Reads sources and generates a summary | The result exists only in memory and is lost if the process exits |
| T1 | The report service saves successfully | Report R1 exists remotely |
| T2 | The process terminates before saving the receipt locally | The local task still shows publish as incomplete |
| T3 | Worker B takes over and retries publish | A new operation key on each attempt may create R2 |
Even deterministic model output leaves a commit gap between T1 and T2. A local transaction controls only the local database; it cannot roll back a write already committed by another service. Recovery must answer two separate questions: where execution resumes, and how the business service handles duplicate write requests.
The approach: Checkpoints identify the next step, lease generations reject stale workers' writebacks, and stable operation keys with durable remote receipts handle retries. Each mechanism covers a different failure window.
02 · A verifiable example: four steps and two databases
The shared example is a fixed four-step research workflow. It validates the agent harness: scheduling, state, tool results, and recovery boundaries. By default, the draft step uses a deterministic stand-in for reproducibility; an HTTP model is optional. The model does not dynamically plan a tool route, so these results do not demonstrate an autonomous agent's planning ability.
| Step | Input → output | Can it run again? |
|---|---|---|
| collect | Two frozen sources + valid memory → context snapshot | Yes; it reads only local data |
| draft | Context snapshot → summary / citations / language | Skip committed results; repeating an uncommitted call may incur another model charge |
| verify | Draft + allowed source set → shape and citation-ID checks | Yes; a failed check blocks publish |
| publish | Committed draft → simulated report-service receipt | Requests may be retried; one operation key retains one business record |
The run creates lab.sqlite for tasks, steps, events, and memory, and lab.sqlite.publisher.sqlite for the simulated external report service. Using two independent databases deliberately prevents the remote write and local checkpoint from sharing a transaction. The webpage's recovery results were read from these databases; the source material is synthetic.
One example shared across all three articles
Python 3.10+ · Standard library · Offline by default · Source code, 36 tests, and run records included
Download the complete lab ZIP · Run instructions
- Create an empty directory, extract the package, and enter the directory containing cli.py.
- Run the commands below. The default path needs no pip install, account, or API key.
- Inspect the four checkpoints with inspect, then run evaluate.py for mechanical acceptance checks.
First run: the normal path
python3 cli.py memory-put
python3 cli.py submit
python3 cli.py run
python3 cli.py inspect
python3 evaluate.py
python3 -m unittest discover -s . -p test_lab.py -v
run should report succeeded. The evaluation JSON reports passed=true, while semantic_support and model_quality remain not_scored. These omissions are explicit: an existing citation ID establishes no semantic support, and an offline test double provides no measurement of live model quality.
03 · Database fields tied to specific failure modes
runtime.py · The schema used by the lab
SCHEMA = '''
CREATE TABLE IF NOT EXISTS jobs(
id TEXT PRIMARY KEY, spec TEXT NOT NULL, input_hash TEXT NOT NULL,
status TEXT NOT NULL, generation INTEGER NOT NULL DEFAULT 0, owner TEXT,
lease_until REAL NOT NULL DEFAULT 0, attempts INTEGER NOT NULL DEFAULT 0,
available_at REAL NOT NULL DEFAULT 0, scope TEXT NOT NULL,
memory_rev INTEGER NOT NULL, memory_expires REAL NOT NULL,
deadline REAL NOT NULL, error TEXT);
CREATE TABLE IF NOT EXISTS checkpoints(
job_id TEXT NOT NULL REFERENCES jobs(id), step TEXT NOT NULL, payload TEXT NOT NULL,
PRIMARY KEY(job_id,step));
CREATE TABLE IF NOT EXISTS events(
seq INTEGER PRIMARY KEY AUTOINCREMENT, job_id TEXT NOT NULL REFERENCES jobs(id),
at REAL NOT NULL, generation INTEGER NOT NULL, kind TEXT NOT NULL, detail TEXT NOT NULL);
'''
| Field or constraint | Purpose | Failure if omitted |
|---|---|---|
| jobs.id + input_hash | Repeated submissions with one task ID must retain the same input and scope | An old checkpoint is applied to a new question |
| generation + owner + lease_until | Identify the current authorized worker and limit its execution window | A stale process overwrites a newer worker's result |
| attempts + available_at + deadline | Bound claims, backoff, and total execution time | Retry storms or tasks that never terminate |
| checkpoints primary key (job_id,step) | Retain one committed result per step | Recovery cannot select the authoritative output |
| memory_rev + memory_expires | Detect invalid context before recovery | Revoked memory continues to affect output |
| events.seq + generation | Retain ordered, correlated execution evidence | Final status alone cannot explain failure |
input_hash is computed from canonical JSON containing sources, provider, model, endpoint, workflow version, and scope. Resubmitting the same task ID with a different digest fails. Model credentials are excluded from spec. Recovery does not automatically load updated sources: create a new task ID explicitly. Changes to step semantics require a new report/v1 version and a new task under this lab's migration rule.
At larger scale, keep large documents outside jobs.spec. Store object-storage URIs, content hashes, and authorization snapshots in the task, and verify hashes when fetching content. This lab inlines two short synthetic sources to simplify reproduction.
04 · Pair claims with guarded writes: acquiring a lock is insufficient
runtime.py · Atomic task claim
def claim(self, job, owner):
with transaction(self.db):
row, now = self.job(job), self.clock()
ready = (row['status'] in ('queued', 'retry_wait') and row['available_at'] <= now
or row['status'] == 'running' and row['lease_until'] <= now)
if not ready:
return None
if row['deadline'] <= now or row['attempts'] >= 3:
self.db.execute("UPDATE jobs SET status='failed',error='deadline_or_attempt_budget' WHERE id=?", (job,))
self.event(job, row['generation'], 'failed', 'deadline_or_attempt_budget')
return None
generation = row['generation'] + 1
self.db.execute("UPDATE jobs SET status='running',owner=?,generation=?,lease_until=?,"
'attempts=attempts+1 WHERE id=?', (owner, generation, now + self.lease_seconds, job))
self.event(job, generation, 'claimed', owner)
return generation
runtime.py · Validate the worker before writing
def guard(self, job, owner, generation):
row = self.job(job)
if row['status'] != 'running' or row['owner'] != owner or row['generation'] != generation or row['lease_until'] <= self.clock():
raise LostLease('stale worker or cancelled task')
return row
SQLite BEGIN IMMEDIATE starts a write transaction immediately. Each database still has only one writer, so competing writes may return SQLITE_BUSY. The claim transaction reads and updates one task and appends an event; it never waits for a model call. The connection waits up to five seconds for a lock, after which the caller handles failure. A transaction does not eliminate lock contention.
SQLite · Transactions documents write concurrency and IMMEDIATE behavior. (Source checked: 2026-09-25)
Consider this race: A claims generation=1, pauses until its lease expires, B claims generation=2, then A resumes and submits. The unique checkpoint slot may still be empty, so uniqueness alone does not stop A. guard checks running status, owner, generation, and lease together and rejects A's write. Cancellation also increments generation, removing the old worker's writeback authority.
This lab renews the lease at the start of each step and has no background heartbeat. A model call exceeding the 30-second lease can return successfully but still have its result rejected. Production needs periodic renewal and must stop further tool calls if renewal fails. External tools require their own idempotency or fencing contract: a local guard cannot recall a request already sent.
05 · Call tools outside transactions; commit results and events together
runtime.py · Commit results and completion status
def commit_step(self, job, owner, generation, step, payload):
with transaction(self.db):
row = self.guard(job, owner, generation)
self.validate_memory(row)
if row['deadline'] <= self.clock():
raise PermanentError('deadline exceeded')
self.db.execute('INSERT INTO checkpoints VALUES (?,?,?)', (job, step, encode(payload)))
self.event(job, generation, 'step_committed', step)
if step == 'publish':
self.db.execute("UPDATE jobs SET status='succeeded',owner=NULL,error=NULL WHERE id=?", (job,))
commit_step saves the result and step_committed event together. The final step also sets succeeded in that transaction, preventing a local success state without a durable publish checkpoint. Recovery reads existing checkpoints and executes only missing steps; it does not ask the model to infer the next step from logs.
| Crash point | State visible after restart | Next action |
|---|---|---|
| Before sending draft | collect committed; draft absent | Continue with draft |
| After the model responds but before committing draft | Only collect exists | The model may be called again, potentially charging again |
| After committing draft | The draft result exists | Reuse it without regeneration |
| After the final transaction commits | publish and succeeded both exist | A duplicate delivery returns succeeded |
LangGraph's persistence documentation describes threads and checkpoints. Whether using a framework or a custom runtime, verify durable storage, replayed steps, and deduplication of external effects. The implementation shown here was designed for this site; it is not LangGraph source code.
LangGraph · Persistence provides the reference for threads, checkpoints, and persistence. (Source checked: 2026-09-25)
06 · Failure experiment A: exit the process and resume after collect
Reproduction commands
python3 cli.py submit --db crash.sqlite
python3 cli.py run --db crash.sqlite --lease-seconds 2 --fault after_collect
# 预期退出码 75;独立执行上面这条命令
python3 cli.py inspect --db crash.sqlite
# 等待至少 2 秒,让旧租约到期
python3 cli.py run --db crash.sqlite
python3 evaluate.py --db crash.sqlite
The fault branch runs after the collect checkpoint commits. The CLI calls os._exit(75), terminating the process without finally cleanup or marking the task failed. The first inspect therefore shows running with collect committed, rather than the state produced by graceful exception handling.
| Check | After interruption | After recovery |
|---|---|---|
| Task status | running | succeeded |
| generation | 1 | 2 |
| Committed steps | collect | collect, draft, verify, publish |
| collect commit count | 1 | 1 |
If the second run still reports running, inspect lease_until: the previous lease may remain valid. Do not delete the task row to force recovery. Exceeding the default 600-second deadline terminates the task; resubmit with a new job ID. The three-claim limit includes initial execution, rather than allowing three additional retries.
07 · Failure experiment B: remote success without a local receipt
runtime.py · The independent service's idempotency contract
def publish(self, job, payload):
# Independent database simulates an external service's durable idempotency contract.
db = sqlite3.connect(self.path + '.publisher.sqlite', isolation_level=None, timeout=5)
try:
db.execute('CREATE TABLE IF NOT EXISTS receipts(op TEXT PRIMARY KEY, payload_hash TEXT NOT NULL, receipt TEXT NOT NULL)')
op, payload_hash = digest([job, 'publish/v1']), digest(payload)
with transaction(db):
old = db.execute('SELECT payload_hash,receipt FROM receipts WHERE op=?', (op,)).fetchone()
if old:
if old[0] != payload_hash:
raise PermanentError('idempotency key reused with different payload')
return {'receipt': old[1], 'deduplicated': True}
receipt = 'report-' + op[:12]
db.execute('INSERT INTO receipts VALUES (?,?,?)', (op, payload_hash, receipt))
return {'receipt': receipt, 'deduplicated': False}
finally:
db.close()
The operation key is SHA-256(job_id, publish/v1), unchanged across retries of the same logical publication. The service persists payload_hash and receipt. Matching key and content return the existing receipt; changed content under the same key fails. Including generation or retry count in the key would bypass deduplication on every takeover.
Reproduce the cross-system commit gap
python3 cli.py submit --db effect.sqlite
python3 cli.py run --db effect.sqlite --lease-seconds 2 --fault after_effect
# 等待至少 2 秒
python3 cli.py run --db effect.sqlite
python3 cli.py inspect --db effect.sqlite
Recorded result: At interruption, collect, draft, and verify were committed. Recovery adds the fourth checkpoint. The simulated publishing database retains one receipts row, and the recovered receipt has deduplicated=true. Inspect the original run snapshot.
This test demonstrates one failure against a service with a durable idempotency contract. It does not cover network partitions, remote disaster recovery, key expiry, or an incorrect provider implementation. If the target report API lacks idempotency keys, add operation_id status queries. If neither querying nor deduplication is available, preserve unknown and reconcile manually instead of retrying blindly.
08 · Encode retry, cancellation, and deadline rules
| Event | Lab behavior | Production requirement |
|---|---|---|
| HTTP 429 / 5xx / network error | retry_wait; wait 2 then 4 seconds by claim count; stop after the third failure | Scheduled redelivery, jitter, and Retry-After handling |
| 401 / 403 / invalid output / unknown source | failed with a recorded reason | Alerts or human correction by explicit error category |
| Cancellation | Set cancelled and increment generation | Compensate completed external effects separately; do not claim rollback |
| deadline exceeded | Stop at a claim or step boundary | Enforce an overall deadline on each external call |
| Unknown process error / forced exit | Retain running; allow takeover after lease expiry | Infrastructure alerts and limits on repeated crash loops |
The two- and four-second delays are reproducible lab settings, not production recommendations. Full jitter can sample from 0 to min(cap, base×2^attempt), subject to the business deadline. More retries can improve apparent completion rates while increasing costs or duplicate effects. Inspect failure reasons and final business records together.
The adapter's timeout=20 is a socket timeout, not a strict wall-clock limit for the whole call. Slow chunked responses may run longer and, without heartbeats, lose writeback authority when the lease expires. A real 60-second completion requirement needs a cancellable HTTP client, an overall deadline, and coordinated scheduling.
09 · Connect Qwen or a compatible API while retaining execution boundaries
First reproduce the default test-double path, then switch to provider=api. The model ID, region, workspace, and API key must match the provider console. The adapter reads LLM_API_KEY and sends sources and valid memory to your configured HTTPS endpoint. It neither persists credentials in the task database nor follows redirects.
Set LLM_API_KEY, then replace the placeholders and run
python3 cli.py submit --db api.sqlite --job api-001 \
--provider api --model YOUR_MODEL_ID \
--endpoint https://YOUR_TRUSTED_HOST/compatible-mode/v1/chat/completions
python3 cli.py run --db api.sqlite
python3 cli.py inspect --db api.sqlite
See provider.py for the implementation. Requests use max_tokens=800 and temperature=0 and parse plain JSON strictly. Markdown-fenced output fails explicitly rather than being repaired speculatively. Use failure fixtures to refine prompts or adopt supported structured output. temperature=0 provides no guarantee of identical output across service versions.
Alibaba Cloud Model Studio · Chat Completions documents request formats, regional endpoints, and workspace-specific domains; use the official documentation and console for your configuration. (Source checked: 2026-09-25)
The adapter was tested against simulated successful HTTP and 429 responses. No live model was called, and provider availability was not tested. A live integration still needs semantic review of summaries against sources; verify currently checks only shape and citation IDs.
10 · Integrate with Java: services and transaction boundaries
| Lab component | Java component | Constraint |
|---|---|---|
| cli.submit / jobs | TaskApplicationService + task table | Reuse request idempotency IDs; obtain tenant and user from authentication |
| claim / before_step | Scheduler + separate ClaimService | Claim in a short transaction; record owner, generation, and lease_until |
| provider.draft | ModelClient / HTTP client | Call outside transactions; start verification with read-only synthetic sources |
| commit_step | CheckpointService | Validate the worker and write checkpoint and event in one transaction |
| publish | ReportClient + existing business service | Persist operation keys at the provider; random replacement keys must not bypass deduplication |
| evaluate / tests | CI regression stage | Retain failure traces and receipts; limit effects on real services |
Java integration sequence
// 结构示意;不是可直接编译的完整 Java 项目
var claim = claimService.claim(taskId); // 独立短事务
if (claim == null) return;
var input = checkpointService.loadInput(claim);
var output = modelClient.generate(input); // 不持有数据库事务
checkpointService.commit(claim, output); // 校验 generation + lease
Do not wrap task lookup, a model call, and result storage in one large transaction-annotated method. Network waits hold connections and locks. Put short transactional methods in separate services and call through actual transaction boundaries; self-invocation may not trigger the framework proxy. Implement claim SQL for the chosen engine and test contention.
For PostgreSQL, evaluate FOR UPDATE SKIP LOCKED for queue claims. It skips locked rows so multiple consumers can claim pending tasks, but guarantees no strict completion order by creation time. Ordered work per account or project requires additional partitioned serialization; SKIP LOCKED is not FIFO.
PostgreSQL · SELECT describes SKIP LOCKED for multiple consumers of queue-like tables and its unsuitability for general consistent views. (Source checked: 2026-09-25)
11 · Troubleshooting and remaining production requirements
| Symptom | Inspect first | Cause and response |
|---|---|---|
| Task remains running | lease_until and the last step_started | A blocked call or exited worker; take over after expiry and check overall call timeouts |
| LostLease | owner, generation, and last renewal | A stale worker returned late; stop writeback without weakening the guard |
| database is locked | Transaction duration and write concurrency | Shorten transactions and reduce contention; migrate storage when scale requires it |
| Duplicate reports | Operation keys and remote receipts | New keys on retries, expired keys, or no durable provider deduplication |
| Old results after input changes | job_id and input_hash | Create a new task or migrate explicitly; never overwrite checkpoints manually |
| Failure after a memory update | memory_rev and the task snapshot | Invalidation correctly blocked recovery; create a new task instead of reusing the draft |
The code reproduces failures locally but supplies no HTTP service, authentication, production queue, database migrations, or monitoring deployment. Version 3 adds local approval experiments rather than a real approval backend. Before production, implement these components and test full disks, timeouts after remote commits, partitions, idempotency retention, database recovery, and permission changes. Eighteen passing base tests are not a production availability measurement.
CLI-generated scope trusts the local operator. Never map an open API's tenant_id directly into that trusted value. Actual writes need the existing business authorization service. Bind approvals to action type, resource ID, content digest, expiry, and requester. Changed content requires a new approval; an old approval cannot authorize a new action.
Start by reproducing the two failures, then run concurrent-claim and stale-generation tests, and finally replace the model and report adapters with read-only and sandbox integrations. Retain local database snapshots and remote receipts at each stage. Establish an explainable failure before increasing concurrency.
12 · Bind approval to an action, rather than just approved=true
Extend the report example so publish requires human approval. A minimal flow sets approved=true after a click and resumes. But the reviewed draft, actual payload, project, or target service may have changed. A boolean records no approved scope, leaving recovery unable to verify it.
| Bound field | Lab value or source | Failure prevented |
|---|---|---|
| job_id + input_hash | Task ID and frozen-input digest | Approval for an old task applied to new input |
| action + target | report.publish/v1 + local-report-store | Approval to save a report repurposed for another action |
| scope + requester | Task tenant, user, project, and requester | Cross-project reuse or self-approval |
| payload_hash | Canonical JSON digest of the committed draft | Silent content changes after approval |
| memory_rev / memory_expires | Memory snapshot revision and expiry | Reuse after correction or revocation |
| expires_at / decided_by | Request expiry and reviewer | Indefinite reuse or untraceable decisions |
approval_runtime.py · Fields included in the approval digest
def envelope(self, job, report):
row = self.job(job)
spec = json.loads(row['spec'])
return {'job_id': job, 'action': 'report.publish/v1',
'target': spec.get('report_target', 'local-report-store'),
'scope': row['scope'], 'requester': json.loads(row['scope'])[1],
'input_hash': row['input_hash'], 'payload_hash': digest(report),
'memory_rev': row['memory_rev'], 'memory_expires': row['memory_expires']}
Download full approval_runtime.py
binding_hash identifies the action description above. The review page must also show the concrete action, target, and complete draft. A hash alone provides no understandable review object. It is neither a signature nor a login credential; reviewer identity must come from trusted authentication.
OWASP · Transaction Authorization describes server-side authorization and binding reviewable transaction data to execution. This article's approval model is an independent teaching implementation. (Source checked: 2026-09-25)
Principal and --actor are local test fixtures, not real authentication. An HTTP interface must obtain roles and allowed_scopes from its authenticated session and authorization system rather than the browser. local-report-store names a simulated target, not a connected report service.
13 · Experiment D: save the draft, wait for approval, then resume
Version 3 adds a separate approval_cli.py entry point, reusing collection, drafting, verification, and checkpoints from runtime.py. The original cli.py behavior is retained for comparison. Download the version 3 lab package from the homepage, extract it, and run:
Create a task awaiting approval
python3 approval_cli.py submit
python3 approval_cli.py run
python3 approval_cli.py inspect
run reports waiting_approval. inspect shows committed collect, draft, and verify steps, with no publish checkpoint. Review checkpoints.draft and approval.action_json, then copy approval.binding_hash into the command below. Approval lasts 120 seconds and the task deadline is 600 seconds. These short lab settings make expiry observable; they are not production recommendations.
Confirm the reviewed action and resume
python3 approval_cli.py approve --actor reviewer --hash REVIEWED_BINDING_HASH
python3 approval_cli.py run
python3 approval_cli.py inspect
| Stage | Task status | Checkpoints and external effects |
|---|---|---|
| Initial run | waiting_approval | Three checkpoints; zero reports |
| approve | queued | Records only the decision; no publishing call |
| Second run | succeeded | Reuses the draft, adds publish, and retains one report |
No worker sleeps throughout the wait, and no database transaction spans a person's approval time. State is durable, so the CLI can exit. Write the decision in a short transaction and place the task in queued; a second run in this example, or a production scheduler, takes over. Waiting itself adds no claims; initial execution and recovery after approval each count as one claim.
LangGraph interrupt also supports a durable pause and resumption on the same thread. The interrupted node restarts from its beginning, so effects before interrupt still need idempotency. A pause mechanism does not define an application's approval scope or authorization policy.
LangGraph · Interrupts describes pauses, persistence, thread resumption, and node re-execution. (Source checked: 2026-09-25)
14 · Protect approval callbacks against unauthorized, duplicate, expired, or changed requests
| Test | Expected lab result | Retained evidence |
|---|---|---|
| approve --actor requester | Reject self-approval; remain pending | Reviewer and requester separation |
| approve --actor outsider / viewer | Reject an incorrect scope or role | Trusted role and allowed scope |
| Supply another binding_hash | Reject without changing the decision | Digest of the content actually reviewed |
| Same reviewer repeats the same decision | Return the existing decision while valid; no new authorization | One approval_approved event |
| Concurrent approve / reject | One succeeds; the other conflicts | One unmixed decision |
| Change draft or memory after approval | invalidated / failed; no dispatch | Current content digest and memory revision |
| Resume after expires_at | expired / failed; no dispatch | Time of the final authorization check |
Both approve and reject require the same reviewed digest. Rejection terminates the request; an old request cannot later approve it. Changed content needs a new task and another review. If the old task may already have dispatched, reconcile its outcome before creating a replacement operation.
Immediately before dispatch, recheck that the same action, content, and context are still allowed. Historical approval alone is insufficient. Validation and the dispatching marker commit in one short transaction. The next section explains why dispatching needs its own state.
approval_runtime.py · Revalidate and persist the dispatch boundary
def publish(self, job, payload):
if not self.execution or self.execution[0] != job:
raise PermissionError('publish requires the guarded execution path')
_, owner, generation = self.execution
with transaction(self.db):
row = self.guard(job, owner, generation)
request = self.approval(job)
if request is None or request['status'] != 'approved':
raise PermanentError('approval required')
self.validate_binding(job, request, payload)
if request['expires_at'] <= self.clock() or row['deadline'] <= self.clock():
raise PermanentError('approval_expired_before_dispatch')
# Authorization cutoff: after this durable marker, outcome may be unknown.
self.db.execute("UPDATE approvals SET status='dispatching' WHERE job_id=?", (job,))
self.event(job, generation, 'dispatch_started', request['operation_key'])
if self.fault == 'before_effect':
raise InjectedCrash('dispatch marker committed; no service call was made')
return super().publish(job, payload)
Download full approval_runtime.py
15 · A difficult window: dispatch succeeded, then approval expired
Consider approval → durable dispatching → remote save → local process exit → approval expiry. Marking the task unexecuted solely because approval expired hides a committed effect. Obtaining approval again and resending may duplicate it. Permission for a new write and the existence of a past write require separate decisions.
| Observable state | Permitted action | Unsupported inference |
|---|---|---|
| approved, before dispatching | Expiry or cancellation can block a new dispatch | Approval does not prove execution |
| dispatching, no local receipt | Query the remote receipt using the stable operation key | Missing receipts do not prove no request was sent |
| Receipt with matching payload_hash | Record the existing effect locally | Reconciliation neither grants new authorization nor reverses the effect |
| No receipt found | Remain reconciling; wait or investigate manually | Do not automatically change keys or tasks, or resend |
| Receipt payload digest differs | Report a conflict and stop | A mismatched receipt cannot establish success |
Experiment E: reconcile external effects after approval
python3 approval_cli.py submit --db approved-effect.sqlite
python3 approval_cli.py run --db approved-effect.sqlite
python3 approval_cli.py inspect --db approved-effect.sqlite
python3 approval_cli.py approve --db approved-effect.sqlite --hash REVIEWED_BINDING_HASH
python3 approval_cli.py run --db approved-effect.sqlite --lease-seconds 2 --fault after_effect
# 故障命令退出 75;等待至少 2 秒
python3 approval_cli.py reconcile --db approved-effect.sqlite
python3 approval_cli.py inspect --db approved-effect.sqlite
Interpretation: effect_confirmed records an established past write. publish.fresh_write_performed=false shows that recovery queried only and initiated no new publication. This differs from normal succeeded and does not mean a cancellation rolled back the report. Inspect the original record.
With before_effect, the process exits after committing dispatching but before calling the provider. The injector knows no request was sent; recovery knows only dispatching and no receipt, so it remains reconciling. This conservative choice limits automatic progress rather than treating uncertainty as permission to resend. A reliable provider idempotency contract may support a different recovery protocol; this extension implements read-only reconciliation.
Unresolved outcomes need manual reconciliation or remote terminal-state queries, rather than indefinite unattended waiting. This code has no background reconciler, alerts, or manual closure API. Production closure must retain evidence; operators must not simply change dispatching back to queued.
16 · Cancellation stops future work without inventing a rollback
| Cancellation point | Task state | External effect |
|---|---|---|
| Awaiting approval, or approved before dispatch | cancelled; approval invalidated | No report has been published in this lab |
| After dispatching | reconciling; record cancel_requested_after_dispatch | The request may still execute; absence of writes cannot be guaranteed |
| Reconciliation finds a receipt | effect_confirmed; retain cancellation intent | The report exists; reversal requires a separate business operation |
| Reconciliation finds no receipt | Remain reconciling | Continue investigation without claiming reversal |
The final authorization check before dispatch is the lab's boundary. Cancellation and the dispatching update compete in short database transactions. Cancellation committing first disqualifies the worker; dispatching committing first enters the unknown-outcome region. External requests are outside the transaction, so clicking cancel cannot guarantee zero business effects.
Compensation is not a generic inverse of recovery code. Draft deletion, publication withdrawal, refunds, and reversals have distinct rules. A Java integration needs explicit compensation commands, independent approval, and auditing from the business service. A model must not invent reverse SQL from a natural-language instruction.
17 · Java integration: separate approval, dispatch, and reconciliation
| Service or table | Transactional work | Work outside that transaction |
|---|---|---|
| ApprovalService / approvals | Validate reviewer, scope, digest, version, and state; record the decision | No model or publishing call |
| DispatchService | Validate active approval and worker lease; record dispatching | Call the business service with a stable operation_id |
| ReceiptReconciler | Record reconciliation and audit evidence; update terminal state as needed | Query by operation_id without creating a new effect |
| CheckpointService | Commit publish checkpoint, task state, and receipt together | Do not regenerate a draft to manufacture a missing receipt |
| PermissionGateway | Read trusted identity and server policy | Reauthorize each actual business access |
A suggested HTTP contract is POST /approvals/{id}/decision with decision, reviewed_hash, and a client request idempotency ID. Read actor from the session and compute permissions and tenant scope on the server. Suggested responses are 409 for version/decision conflicts, 403 for unauthorized approval, and 410 for expired requests. These codes are design suggestions; the lab implements no HTTP API.
This extension is still a local approval state machine with identity fixtures; it does not implement a login backend. During migration, first connect the existing login and role services to a read-only business sandbox. Then verify that expired approvals cannot dispatch actions, changed content requires a new review, cancellation reports business effects accurately, and an unknown receipt never causes a second operation to be created.
Sources and verification
Official sources establish the referenced mechanisms. The schema, program, and experiments were independently designed for this site.