Agent Application DevelopmentAccount
← Return to research directory

RUNTIME

Recovering long-running tasks: from a crash to a runnable harness

Follow checkpoints and leases through human approval, cancellation, and reconciliation of unknown outcomes, with runnable code, crash reproductions, and Java service boundaries.

01 · Why one retry can create two reports

Suppose you are building an internal technical research assistant. It reads two approved sources, generates a summary with source IDs, and saves it through a report service. A task may take tens of minutes; scheduling messages can arrive more than once, and model requests can time out. The business requirements are to retain completed work, explain unknown outcomes, and avoid saving the same report twice.

Time Worker A's action Established fact
T0 Reads sources and generates a summary The result exists only in memory and is lost if the process exits
T1 The report service saves successfully Report R1 exists remotely
T2 The process terminates before saving the receipt locally The local task still shows publish as incomplete
T3 Worker B takes over and retries publish A new operation key on each attempt may create R2

Even deterministic model output leaves a commit gap between T1 and T2. A local transaction controls only the local database; it cannot roll back a write already committed by another service. Recovery must answer two separate questions: where execution resumes, and how the business service handles duplicate write requests.

The approach: Checkpoints identify the next step, lease generations reject stale workers' writebacks, and stable operation keys with durable remote receipts handle retries. Each mechanism covers a different failure window.

02 · A verifiable example: four steps and two databases

The shared example is a fixed four-step research workflow. It validates the agent harness: scheduling, state, tool results, and recovery boundaries. By default, the draft step uses a deterministic stand-in for reproducibility; an HTTP model is optional. The model does not dynamically plan a tool route, so these results do not demonstrate an autonomous agent's planning ability.

Step Input → output Can it run again?
collect Two frozen sources + valid memory → context snapshot Yes; it reads only local data
draft Context snapshot → summary / citations / language Skip committed results; repeating an uncommitted call may incur another model charge
verify Draft + allowed source set → shape and citation-ID checks Yes; a failed check blocks publish
publish Committed draft → simulated report-service receipt Requests may be retried; one operation key retains one business record

The run creates lab.sqlite for tasks, steps, events, and memory, and lab.sqlite.publisher.sqlite for the simulated external report service. Using two independent databases deliberately prevents the remote write and local checkpoint from sharing a transaction. The webpage's recovery results were read from these databases; the source material is synthetic.

One example shared across all three articles

Python 3.10+ · Standard library · Offline by default · Source code, 36 tests, and run records included

Download the complete lab ZIP · Run instructions

  1. Create an empty directory, extract the package, and enter the directory containing cli.py.
  2. Run the commands below. The default path needs no pip install, account, or API key.
  3. Inspect the four checkpoints with inspect, then run evaluate.py for mechanical acceptance checks.

First run: the normal path

python3 cli.py memory-put
python3 cli.py submit
python3 cli.py run
python3 cli.py inspect
python3 evaluate.py
python3 -m unittest discover -s . -p test_lab.py -v

run should report succeeded. The evaluation JSON reports passed=true, while semantic_support and model_quality remain not_scored. These omissions are explicit: an existing citation ID establishes no semantic support, and an offline test double provides no measurement of live model quality.

03 · Database fields tied to specific failure modes

runtime.py · The schema used by the lab

SCHEMA = '''
CREATE TABLE IF NOT EXISTS jobs(
 id TEXT PRIMARY KEY, spec TEXT NOT NULL, input_hash TEXT NOT NULL,
 status TEXT NOT NULL, generation INTEGER NOT NULL DEFAULT 0, owner TEXT,
 lease_until REAL NOT NULL DEFAULT 0, attempts INTEGER NOT NULL DEFAULT 0,
 available_at REAL NOT NULL DEFAULT 0, scope TEXT NOT NULL,
 memory_rev INTEGER NOT NULL, memory_expires REAL NOT NULL,
 deadline REAL NOT NULL, error TEXT);
CREATE TABLE IF NOT EXISTS checkpoints(
 job_id TEXT NOT NULL REFERENCES jobs(id), step TEXT NOT NULL, payload TEXT NOT NULL,
 PRIMARY KEY(job_id,step));
CREATE TABLE IF NOT EXISTS events(
 seq INTEGER PRIMARY KEY AUTOINCREMENT, job_id TEXT NOT NULL REFERENCES jobs(id),
 at REAL NOT NULL, generation INTEGER NOT NULL, kind TEXT NOT NULL, detail TEXT NOT NULL);
'''

Download full runtime.py

Field or constraint Purpose Failure if omitted
jobs.id + input_hash Repeated submissions with one task ID must retain the same input and scope An old checkpoint is applied to a new question
generation + owner + lease_until Identify the current authorized worker and limit its execution window A stale process overwrites a newer worker's result
attempts + available_at + deadline Bound claims, backoff, and total execution time Retry storms or tasks that never terminate
checkpoints primary key (job_id,step) Retain one committed result per step Recovery cannot select the authoritative output
memory_rev + memory_expires Detect invalid context before recovery Revoked memory continues to affect output
events.seq + generation Retain ordered, correlated execution evidence Final status alone cannot explain failure

input_hash is computed from canonical JSON containing sources, provider, model, endpoint, workflow version, and scope. Resubmitting the same task ID with a different digest fails. Model credentials are excluded from spec. Recovery does not automatically load updated sources: create a new task ID explicitly. Changes to step semantics require a new report/v1 version and a new task under this lab's migration rule.

At larger scale, keep large documents outside jobs.spec. Store object-storage URIs, content hashes, and authorization snapshots in the task, and verify hashes when fetching content. This lab inlines two short synthetic sources to simplify reproduction.

04 · Pair claims with guarded writes: acquiring a lock is insufficient

runtime.py · Atomic task claim

    def claim(self, job, owner):
        with transaction(self.db):
            row, now = self.job(job), self.clock()
            ready = (row['status'] in ('queued', 'retry_wait') and row['available_at'] <= now
                     or row['status'] == 'running' and row['lease_until'] <= now)
            if not ready:
                return None
            if row['deadline'] <= now or row['attempts'] >= 3:
                self.db.execute("UPDATE jobs SET status='failed',error='deadline_or_attempt_budget' WHERE id=?", (job,))
                self.event(job, row['generation'], 'failed', 'deadline_or_attempt_budget')
                return None
            generation = row['generation'] + 1
            self.db.execute("UPDATE jobs SET status='running',owner=?,generation=?,lease_until=?,"
                'attempts=attempts+1 WHERE id=?', (owner, generation, now + self.lease_seconds, job))
            self.event(job, generation, 'claimed', owner)
            return generation

Download full runtime.py

runtime.py · Validate the worker before writing

    def guard(self, job, owner, generation):
        row = self.job(job)
        if row['status'] != 'running' or row['owner'] != owner or row['generation'] != generation or row['lease_until'] <= self.clock():
            raise LostLease('stale worker or cancelled task')
        return row

Download full runtime.py

SQLite BEGIN IMMEDIATE starts a write transaction immediately. Each database still has only one writer, so competing writes may return SQLITE_BUSY. The claim transaction reads and updates one task and appends an event; it never waits for a model call. The connection waits up to five seconds for a lock, after which the caller handles failure. A transaction does not eliminate lock contention.

SQLite · Transactions documents write concurrency and IMMEDIATE behavior. (Source checked: 2026-09-25)

Consider this race: A claims generation=1, pauses until its lease expires, B claims generation=2, then A resumes and submits. The unique checkpoint slot may still be empty, so uniqueness alone does not stop A. guard checks running status, owner, generation, and lease together and rejects A's write. Cancellation also increments generation, removing the old worker's writeback authority.

This lab renews the lease at the start of each step and has no background heartbeat. A model call exceeding the 30-second lease can return successfully but still have its result rejected. Production needs periodic renewal and must stop further tool calls if renewal fails. External tools require their own idempotency or fencing contract: a local guard cannot recall a request already sent.

05 · Call tools outside transactions; commit results and events together

runtime.py · Commit results and completion status

    def commit_step(self, job, owner, generation, step, payload):
        with transaction(self.db):
            row = self.guard(job, owner, generation)
            self.validate_memory(row)
            if row['deadline'] <= self.clock():
                raise PermanentError('deadline exceeded')
            self.db.execute('INSERT INTO checkpoints VALUES (?,?,?)', (job, step, encode(payload)))
            self.event(job, generation, 'step_committed', step)
            if step == 'publish':
                self.db.execute("UPDATE jobs SET status='succeeded',owner=NULL,error=NULL WHERE id=?", (job,))

Download full runtime.py

commit_step saves the result and step_committed event together. The final step also sets succeeded in that transaction, preventing a local success state without a durable publish checkpoint. Recovery reads existing checkpoints and executes only missing steps; it does not ask the model to infer the next step from logs.

Crash point State visible after restart Next action
Before sending draft collect committed; draft absent Continue with draft
After the model responds but before committing draft Only collect exists The model may be called again, potentially charging again
After committing draft The draft result exists Reuse it without regeneration
After the final transaction commits publish and succeeded both exist A duplicate delivery returns succeeded

LangGraph's persistence documentation describes threads and checkpoints. Whether using a framework or a custom runtime, verify durable storage, replayed steps, and deduplication of external effects. The implementation shown here was designed for this site; it is not LangGraph source code.

LangGraph · Persistence provides the reference for threads, checkpoints, and persistence. (Source checked: 2026-09-25)

06 · Failure experiment A: exit the process and resume after collect

Reproduction commands

python3 cli.py submit --db crash.sqlite
python3 cli.py run --db crash.sqlite --lease-seconds 2 --fault after_collect
# 预期退出码 75;独立执行上面这条命令
python3 cli.py inspect --db crash.sqlite
# 等待至少 2 秒,让旧租约到期
python3 cli.py run --db crash.sqlite
python3 evaluate.py --db crash.sqlite

The fault branch runs after the collect checkpoint commits. The CLI calls os._exit(75), terminating the process without finally cleanup or marking the task failed. The first inspect therefore shows running with collect committed, rather than the state produced by graceful exception handling.

Check After interruption After recovery
Task status running succeeded
generation 1 2
Committed steps collect collect, draft, verify, publish
collect commit count 1 1

If the second run still reports running, inspect lease_until: the previous lease may remain valid. Do not delete the task row to force recovery. Exceeding the default 600-second deadline terminates the task; resubmit with a new job ID. The three-claim limit includes initial execution, rather than allowing three additional retries.

07 · Failure experiment B: remote success without a local receipt

runtime.py · The independent service's idempotency contract

    def publish(self, job, payload):
        # Independent database simulates an external service's durable idempotency contract.
        db = sqlite3.connect(self.path + '.publisher.sqlite', isolation_level=None, timeout=5)
        try:
            db.execute('CREATE TABLE IF NOT EXISTS receipts(op TEXT PRIMARY KEY, payload_hash TEXT NOT NULL, receipt TEXT NOT NULL)')
            op, payload_hash = digest([job, 'publish/v1']), digest(payload)
            with transaction(db):
                old = db.execute('SELECT payload_hash,receipt FROM receipts WHERE op=?', (op,)).fetchone()
                if old:
                    if old[0] != payload_hash:
                        raise PermanentError('idempotency key reused with different payload')
                    return {'receipt': old[1], 'deduplicated': True}
                receipt = 'report-' + op[:12]
                db.execute('INSERT INTO receipts VALUES (?,?,?)', (op, payload_hash, receipt))
                return {'receipt': receipt, 'deduplicated': False}
        finally:
            db.close()

Download full runtime.py

The operation key is SHA-256(job_id, publish/v1), unchanged across retries of the same logical publication. The service persists payload_hash and receipt. Matching key and content return the existing receipt; changed content under the same key fails. Including generation or retry count in the key would bypass deduplication on every takeover.

Reproduce the cross-system commit gap

python3 cli.py submit --db effect.sqlite
python3 cli.py run --db effect.sqlite --lease-seconds 2 --fault after_effect
# 等待至少 2 秒
python3 cli.py run --db effect.sqlite
python3 cli.py inspect --db effect.sqlite

Recorded result: At interruption, collect, draft, and verify were committed. Recovery adds the fourth checkpoint. The simulated publishing database retains one receipts row, and the recovered receipt has deduplicated=true. Inspect the original run snapshot.

This test demonstrates one failure against a service with a durable idempotency contract. It does not cover network partitions, remote disaster recovery, key expiry, or an incorrect provider implementation. If the target report API lacks idempotency keys, add operation_id status queries. If neither querying nor deduplication is available, preserve unknown and reconcile manually instead of retrying blindly.

08 · Encode retry, cancellation, and deadline rules

Event Lab behavior Production requirement
HTTP 429 / 5xx / network error retry_wait; wait 2 then 4 seconds by claim count; stop after the third failure Scheduled redelivery, jitter, and Retry-After handling
401 / 403 / invalid output / unknown source failed with a recorded reason Alerts or human correction by explicit error category
Cancellation Set cancelled and increment generation Compensate completed external effects separately; do not claim rollback
deadline exceeded Stop at a claim or step boundary Enforce an overall deadline on each external call
Unknown process error / forced exit Retain running; allow takeover after lease expiry Infrastructure alerts and limits on repeated crash loops

The two- and four-second delays are reproducible lab settings, not production recommendations. Full jitter can sample from 0 to min(cap, base×2^attempt), subject to the business deadline. More retries can improve apparent completion rates while increasing costs or duplicate effects. Inspect failure reasons and final business records together.

The adapter's timeout=20 is a socket timeout, not a strict wall-clock limit for the whole call. Slow chunked responses may run longer and, without heartbeats, lose writeback authority when the lease expires. A real 60-second completion requirement needs a cancellable HTTP client, an overall deadline, and coordinated scheduling.

09 · Connect Qwen or a compatible API while retaining execution boundaries

First reproduce the default test-double path, then switch to provider=api. The model ID, region, workspace, and API key must match the provider console. The adapter reads LLM_API_KEY and sends sources and valid memory to your configured HTTPS endpoint. It neither persists credentials in the task database nor follows redirects.

Set LLM_API_KEY, then replace the placeholders and run

python3 cli.py submit --db api.sqlite --job api-001 \
  --provider api --model YOUR_MODEL_ID \
  --endpoint https://YOUR_TRUSTED_HOST/compatible-mode/v1/chat/completions
python3 cli.py run --db api.sqlite
python3 cli.py inspect --db api.sqlite

See provider.py for the implementation. Requests use max_tokens=800 and temperature=0 and parse plain JSON strictly. Markdown-fenced output fails explicitly rather than being repaired speculatively. Use failure fixtures to refine prompts or adopt supported structured output. temperature=0 provides no guarantee of identical output across service versions.

Alibaba Cloud Model Studio · Chat Completions documents request formats, regional endpoints, and workspace-specific domains; use the official documentation and console for your configuration. (Source checked: 2026-09-25)

The adapter was tested against simulated successful HTTP and 429 responses. No live model was called, and provider availability was not tested. A live integration still needs semantic review of summaries against sources; verify currently checks only shape and citation IDs.

10 · Integrate with Java: services and transaction boundaries

Lab component Java component Constraint
cli.submit / jobs TaskApplicationService + task table Reuse request idempotency IDs; obtain tenant and user from authentication
claim / before_step Scheduler + separate ClaimService Claim in a short transaction; record owner, generation, and lease_until
provider.draft ModelClient / HTTP client Call outside transactions; start verification with read-only synthetic sources
commit_step CheckpointService Validate the worker and write checkpoint and event in one transaction
publish ReportClient + existing business service Persist operation keys at the provider; random replacement keys must not bypass deduplication
evaluate / tests CI regression stage Retain failure traces and receipts; limit effects on real services

Java integration sequence

// 结构示意;不是可直接编译的完整 Java 项目
var claim = claimService.claim(taskId); // 独立短事务
if (claim == null) return;
var input = checkpointService.loadInput(claim);
var output = modelClient.generate(input); // 不持有数据库事务
checkpointService.commit(claim, output); // 校验 generation + lease

Do not wrap task lookup, a model call, and result storage in one large transaction-annotated method. Network waits hold connections and locks. Put short transactional methods in separate services and call through actual transaction boundaries; self-invocation may not trigger the framework proxy. Implement claim SQL for the chosen engine and test contention.

For PostgreSQL, evaluate FOR UPDATE SKIP LOCKED for queue claims. It skips locked rows so multiple consumers can claim pending tasks, but guarantees no strict completion order by creation time. Ordered work per account or project requires additional partitioned serialization; SKIP LOCKED is not FIFO.

PostgreSQL · SELECT describes SKIP LOCKED for multiple consumers of queue-like tables and its unsuitability for general consistent views. (Source checked: 2026-09-25)

11 · Troubleshooting and remaining production requirements

Symptom Inspect first Cause and response
Task remains running lease_until and the last step_started A blocked call or exited worker; take over after expiry and check overall call timeouts
LostLease owner, generation, and last renewal A stale worker returned late; stop writeback without weakening the guard
database is locked Transaction duration and write concurrency Shorten transactions and reduce contention; migrate storage when scale requires it
Duplicate reports Operation keys and remote receipts New keys on retries, expired keys, or no durable provider deduplication
Old results after input changes job_id and input_hash Create a new task or migrate explicitly; never overwrite checkpoints manually
Failure after a memory update memory_rev and the task snapshot Invalidation correctly blocked recovery; create a new task instead of reusing the draft

The code reproduces failures locally but supplies no HTTP service, authentication, production queue, database migrations, or monitoring deployment. Version 3 adds local approval experiments rather than a real approval backend. Before production, implement these components and test full disks, timeouts after remote commits, partitions, idempotency retention, database recovery, and permission changes. Eighteen passing base tests are not a production availability measurement.

CLI-generated scope trusts the local operator. Never map an open API's tenant_id directly into that trusted value. Actual writes need the existing business authorization service. Bind approvals to action type, resource ID, content digest, expiry, and requester. Changed content requires a new approval; an old approval cannot authorize a new action.

Start by reproducing the two failures, then run concurrent-claim and stale-generation tests, and finally replace the model and report adapters with read-only and sandbox integrations. Retain local database snapshots and remote receipts at each stage. Establish an explainable failure before increasing concurrency.

12 · Bind approval to an action, rather than just approved=true

Extend the report example so publish requires human approval. A minimal flow sets approved=true after a click and resumes. But the reviewed draft, actual payload, project, or target service may have changed. A boolean records no approved scope, leaving recovery unable to verify it.

Bound field Lab value or source Failure prevented
job_id + input_hash Task ID and frozen-input digest Approval for an old task applied to new input
action + target report.publish/v1 + local-report-store Approval to save a report repurposed for another action
scope + requester Task tenant, user, project, and requester Cross-project reuse or self-approval
payload_hash Canonical JSON digest of the committed draft Silent content changes after approval
memory_rev / memory_expires Memory snapshot revision and expiry Reuse after correction or revocation
expires_at / decided_by Request expiry and reviewer Indefinite reuse or untraceable decisions

approval_runtime.py · Fields included in the approval digest

    def envelope(self, job, report):
        row = self.job(job)
        spec = json.loads(row['spec'])
        return {'job_id': job, 'action': 'report.publish/v1',
                'target': spec.get('report_target', 'local-report-store'),
                'scope': row['scope'], 'requester': json.loads(row['scope'])[1],
                'input_hash': row['input_hash'], 'payload_hash': digest(report),
                'memory_rev': row['memory_rev'], 'memory_expires': row['memory_expires']}

Download full approval_runtime.py

binding_hash identifies the action description above. The review page must also show the concrete action, target, and complete draft. A hash alone provides no understandable review object. It is neither a signature nor a login credential; reviewer identity must come from trusted authentication.

OWASP · Transaction Authorization describes server-side authorization and binding reviewable transaction data to execution. This article's approval model is an independent teaching implementation. (Source checked: 2026-09-25)

Principal and --actor are local test fixtures, not real authentication. An HTTP interface must obtain roles and allowed_scopes from its authenticated session and authorization system rather than the browser. local-report-store names a simulated target, not a connected report service.

13 · Experiment D: save the draft, wait for approval, then resume

Version 3 adds a separate approval_cli.py entry point, reusing collection, drafting, verification, and checkpoints from runtime.py. The original cli.py behavior is retained for comparison. Download the version 3 lab package from the homepage, extract it, and run:

Create a task awaiting approval

python3 approval_cli.py submit
python3 approval_cli.py run
python3 approval_cli.py inspect

run reports waiting_approval. inspect shows committed collect, draft, and verify steps, with no publish checkpoint. Review checkpoints.draft and approval.action_json, then copy approval.binding_hash into the command below. Approval lasts 120 seconds and the task deadline is 600 seconds. These short lab settings make expiry observable; they are not production recommendations.

Confirm the reviewed action and resume

python3 approval_cli.py approve --actor reviewer --hash REVIEWED_BINDING_HASH
python3 approval_cli.py run
python3 approval_cli.py inspect
Stage Task status Checkpoints and external effects
Initial run waiting_approval Three checkpoints; zero reports
approve queued Records only the decision; no publishing call
Second run succeeded Reuses the draft, adds publish, and retains one report

No worker sleeps throughout the wait, and no database transaction spans a person's approval time. State is durable, so the CLI can exit. Write the decision in a short transaction and place the task in queued; a second run in this example, or a production scheduler, takes over. Waiting itself adds no claims; initial execution and recovery after approval each count as one claim.

LangGraph interrupt also supports a durable pause and resumption on the same thread. The interrupted node restarts from its beginning, so effects before interrupt still need idempotency. A pause mechanism does not define an application's approval scope or authorization policy.

LangGraph · Interrupts describes pauses, persistence, thread resumption, and node re-execution. (Source checked: 2026-09-25)

14 · Protect approval callbacks against unauthorized, duplicate, expired, or changed requests

Test Expected lab result Retained evidence
approve --actor requester Reject self-approval; remain pending Reviewer and requester separation
approve --actor outsider / viewer Reject an incorrect scope or role Trusted role and allowed scope
Supply another binding_hash Reject without changing the decision Digest of the content actually reviewed
Same reviewer repeats the same decision Return the existing decision while valid; no new authorization One approval_approved event
Concurrent approve / reject One succeeds; the other conflicts One unmixed decision
Change draft or memory after approval invalidated / failed; no dispatch Current content digest and memory revision
Resume after expires_at expired / failed; no dispatch Time of the final authorization check

Both approve and reject require the same reviewed digest. Rejection terminates the request; an old request cannot later approve it. Changed content needs a new task and another review. If the old task may already have dispatched, reconcile its outcome before creating a replacement operation.

Immediately before dispatch, recheck that the same action, content, and context are still allowed. Historical approval alone is insufficient. Validation and the dispatching marker commit in one short transaction. The next section explains why dispatching needs its own state.

approval_runtime.py · Revalidate and persist the dispatch boundary

    def publish(self, job, payload):
        if not self.execution or self.execution[0] != job:
            raise PermissionError('publish requires the guarded execution path')
        _, owner, generation = self.execution
        with transaction(self.db):
            row = self.guard(job, owner, generation)
            request = self.approval(job)
            if request is None or request['status'] != 'approved':
                raise PermanentError('approval required')
            self.validate_binding(job, request, payload)
            if request['expires_at'] <= self.clock() or row['deadline'] <= self.clock():
                raise PermanentError('approval_expired_before_dispatch')
            # Authorization cutoff: after this durable marker, outcome may be unknown.
            self.db.execute("UPDATE approvals SET status='dispatching' WHERE job_id=?", (job,))
            self.event(job, generation, 'dispatch_started', request['operation_key'])
        if self.fault == 'before_effect':
            raise InjectedCrash('dispatch marker committed; no service call was made')
        return super().publish(job, payload)

Download full approval_runtime.py

15 · A difficult window: dispatch succeeded, then approval expired

Consider approval → durable dispatching → remote save → local process exit → approval expiry. Marking the task unexecuted solely because approval expired hides a committed effect. Obtaining approval again and resending may duplicate it. Permission for a new write and the existence of a past write require separate decisions.

Observable state Permitted action Unsupported inference
approved, before dispatching Expiry or cancellation can block a new dispatch Approval does not prove execution
dispatching, no local receipt Query the remote receipt using the stable operation key Missing receipts do not prove no request was sent
Receipt with matching payload_hash Record the existing effect locally Reconciliation neither grants new authorization nor reverses the effect
No receipt found Remain reconciling; wait or investigate manually Do not automatically change keys or tasks, or resend
Receipt payload digest differs Report a conflict and stop A mismatched receipt cannot establish success

Experiment E: reconcile external effects after approval

python3 approval_cli.py submit --db approved-effect.sqlite
python3 approval_cli.py run --db approved-effect.sqlite
python3 approval_cli.py inspect --db approved-effect.sqlite
python3 approval_cli.py approve --db approved-effect.sqlite --hash REVIEWED_BINDING_HASH
python3 approval_cli.py run --db approved-effect.sqlite --lease-seconds 2 --fault after_effect
# 故障命令退出 75;等待至少 2 秒
python3 approval_cli.py reconcile --db approved-effect.sqlite
python3 approval_cli.py inspect --db approved-effect.sqlite

Interpretation: effect_confirmed records an established past write. publish.fresh_write_performed=false shows that recovery queried only and initiated no new publication. This differs from normal succeeded and does not mean a cancellation rolled back the report. Inspect the original record.

With before_effect, the process exits after committing dispatching but before calling the provider. The injector knows no request was sent; recovery knows only dispatching and no receipt, so it remains reconciling. This conservative choice limits automatic progress rather than treating uncertainty as permission to resend. A reliable provider idempotency contract may support a different recovery protocol; this extension implements read-only reconciliation.

Unresolved outcomes need manual reconciliation or remote terminal-state queries, rather than indefinite unattended waiting. This code has no background reconciler, alerts, or manual closure API. Production closure must retain evidence; operators must not simply change dispatching back to queued.

16 · Cancellation stops future work without inventing a rollback

Cancellation point Task state External effect
Awaiting approval, or approved before dispatch cancelled; approval invalidated No report has been published in this lab
After dispatching reconciling; record cancel_requested_after_dispatch The request may still execute; absence of writes cannot be guaranteed
Reconciliation finds a receipt effect_confirmed; retain cancellation intent The report exists; reversal requires a separate business operation
Reconciliation finds no receipt Remain reconciling Continue investigation without claiming reversal

The final authorization check before dispatch is the lab's boundary. Cancellation and the dispatching update compete in short database transactions. Cancellation committing first disqualifies the worker; dispatching committing first enters the unknown-outcome region. External requests are outside the transaction, so clicking cancel cannot guarantee zero business effects.

Compensation is not a generic inverse of recovery code. Draft deletion, publication withdrawal, refunds, and reversals have distinct rules. A Java integration needs explicit compensation commands, independent approval, and auditing from the business service. A model must not invent reverse SQL from a natural-language instruction.

17 · Java integration: separate approval, dispatch, and reconciliation

Service or table Transactional work Work outside that transaction
ApprovalService / approvals Validate reviewer, scope, digest, version, and state; record the decision No model or publishing call
DispatchService Validate active approval and worker lease; record dispatching Call the business service with a stable operation_id
ReceiptReconciler Record reconciliation and audit evidence; update terminal state as needed Query by operation_id without creating a new effect
CheckpointService Commit publish checkpoint, task state, and receipt together Do not regenerate a draft to manufacture a missing receipt
PermissionGateway Read trusted identity and server policy Reauthorize each actual business access

A suggested HTTP contract is POST /approvals/{id}/decision with decision, reviewed_hash, and a client request idempotency ID. Read actor from the session and compute permissions and tenant scope on the server. Suggested responses are 409 for version/decision conflicts, 403 for unauthorized approval, and 410 for expired requests. These codes are design suggestions; the lab implements no HTTP API.

This extension is still a local approval state machine with identity fixtures; it does not implement a login backend. During migration, first connect the existing login and role services to a read-only business sandbox. Then verify that expired approvals cannot dispatch actions, changed content requires a new review, cancellation reports business effects accurately, and an unknown receipt never causes a second operation to be created.

Sources and verification

Official sources establish the referenced mechanisms. The schema, program, and experiments were independently designed for this site.

Verification records · Approval recovery evidence