Agent Application DevelopmentAccount
Knowledge catalogChoose core direction and segmented content
knowledge unit 34AdvancedImplementationAbout 18 minutes

Understand → Implement → Debug → Design

Approval snapshots, durable waiting, and one logical resumption

Examine persistence waiting, one-time approval, expiration verification and concurrent recovery.

Human-in-the-loopApproval recoveryCompetition

Knowledge content check2026-10-03 · Check the source of the original question2026-10-02

Which step do you want to learn from this knowledge point?

Select the starting point based on the current basis, or you can go deeper one by one. When you encounter an unfamiliar concept, go back to the core principles first; use the knowledge exercises to check your understanding when you are finished.

Understand first

New to this knowledge point

Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.

Start with core principles →

Implement next

Prepare to write the principles into code

Understand implementation steps and boundaries, complete small tasks, and check results against acceptance requirements.

Reading implementation and trade-offs →

Debug failures

Need to handle failures and changes in conditions

Follow the continuous questioning to locate the failure premise, and then compare the migration cases to explain how the plan should be adjusted.

Continue to delve deeper into the problem →

Compare designs

Need to design or review plans

Combine engineering deductions and senior self-evaluation standards to explain the applicable conditions, costs and alternatives of the plan.

Analyze engineering scenarios →
Knowledge unit directory

LEARN · PRACTICE · REFLECT

Knowledge learning and personal records

My notes and review ↗

First read along the principles, Q&A and migration cases. When you need to check your understanding, switch to reinforcement exercises or start personal recording.

Answers and personal notes

Each modified commit will be kept as an independent history. Your level of mastery is up to you to evaluate yourself against the standards.

Core concept · Approval snapshots, durable waiting, and one logical resumption

Understand the core principles first

Preparatory concepts:Trusted Approval Identity, status condition update, Restoring reentrancy and idempotency

The authorization is a specific, reviewable snapshot of the action, and its content, permissions, and external conditions will change during the waiting period. Recovery requires verification that approval still applies, and merging repeated callbacks into the same logical action.

Waiting need not keep a process alive

Persist the run, pending action snapshot, recovery point, and approval state, then release the worker. A user callback is an event; the scheduler resumes from records. An in-memory approved flag or a three-day sleep cannot support restart or explain who approved which action.

Approval binds action semantics

Record action_hash, revision, approver, scope, expiry, and operation_id. Hash normalized complete business parameters, not a page title. Changes to recipients, cost, resource scope, or body can change the approved action. Authentication verifies identity; a model’s claim of agreement cannot create approval.

Validate twice during recovery

At callback time, check the approver and pending version. Before dispatch, recheck cancellation, unchanged content, current permissions, and business preconditions. Claim one logical resumption through a conditional transition; repeated callbacks return the same approval_id state. Recover failures with the stable operation key rather than deleting approval and inventing a run. Consuming approval and performing a remote action are separate commits; unknown outcomes require reconciliation.

An interrupt is not a program counter

LangGraph resumes an interrupted node from its beginning, repeating pre-interrupt code. Place effects in separately recoverable steps or protect them idempotently. Test double-clicks, duplicate callbacks, changed parameters, departed approvers, cancellation after approval, and remote success without local recording. Count business effects per logical approval, rather than HTTP responses.

Check understanding with a question

After waiting for three days for approval, the user clicks twice to agree. How can I safely restore the Agent?

Persist approval waits rather than keeping sleeping processes alive. Bind approval to run, action digest, version, approver, and expiry. Conditional transitions handle duplicate callbacks. Recheck access, business state, and content before resuming; changes or expiry require review. Protect effects in code that recovery may rerun.

Implementation and trade-offs

Waiting does not occupy the worker thread

Save waiting_approval, action digest, input version, and resume point, then release the worker. User callbacks only express approval events, and the scheduler is responsible for resuming tasks. Items to be approved can still be displayed from the database after restarting, and memory objects or WebSocket connections cannot be regarded as the only state. Tool versions, permissions, and business objects may change during long waits.

Approval and action binding

Approval contains approval_id, run_id, action_hash, revision, actor and expires_at. The server obtains the approver from a trusted identity and uses conditional writing to convert pending to approved or consumed; the same request returns to the processed status if it arrives repeatedly. The approval cannot be used to perform another action, nor can the model paraphrase "user said yes" instead of the actual record.

Re-authenticate on recovery

Check that the task is still waiting and has not been canceled, the action digest has not changed, the approval is valid and the approver still has permissions. Which code will be re-entered when the paused node is restored should be verified according to the framework version used, and side effects should be isolated in idempotent steps. Claiming tasks for recovery also requires concurrency control. Two Workers cannot execute them at the same time just because they both read approved.

How to wrap up when failure occurs

If the external action is successful but the approval-consumed state is not written, the result is determined by the operation’s idempotency key or receipt reconciliation instead of consuming again. Test double-click, repeated callbacks, revoking permissions, modifying parameters, cancellation after approval and worker crash. Each situation should have a clear business end state and explainable audit records.

Engineering deduction

scene
Interview hypothesis: The agree button was clicked repeatedly three days after the approval report was released.
design decisions
Persistently approve and verify by action summary, and receive a recovery with conditional update.
Verify target
A logical release is generated at most once, and repeated callbacks return the same state.
applicable boundary
Whether external publishing can guarantee a one-time effect still depends on the idempotency or receipt mechanism.

Continuous questions and answers

Continue reading along with the premises and constraints of the problem. Understand the reference answers first, then try to put away the answers and explain the cause and effect and trade-offs in your own words.

Draw inferences from one example: If the conditions change, how to deduce it?

First find out the conditions for change, and then determine which premises in the original plan still hold true. The following cases are teaching deductions to facilitate the transfer of principles to new problems.

Price is locked on approval and expires on restore

Changing conditions:The object is the same but the business conditions change

Extended question:Can I continue to pay 650 yuan for a hotel that was originally approved for 500 yuan after three days?

Derivation and reference solutions

If the action amount and terms have exceeded the snapshot, it will be re-approved or judged according to the upper limit rules specified by the user in advance. An approved status must not conceal changed prices; re-query the quotation and save the version before letting the user review it. If there are existing external orders, reconcile them first to avoid duplicate bookings caused by quotation changes.

The principles that remain unchanged:Approval is bound to specific conditions, and task waiting does not freeze the outside world.

Agency for approval

Changing conditions:The original approver is not online and another person approves on his behalf.

Extended question:Is it enough to hold an approval link?

Derivation and reference solutions

The link locates items to be reviewed, and the identity and agency qualifications need to be confirmed by the server. Record the actual actor, agency relationship and permission scope, and check the same action snapshot; repeated callbacks will still be subject to the same approval. Shared links should not be treated as unlimited bearer authorization, especially if there are external effects.

The principles that remain unchanged:Having access to the approval portal does not mean having the ability to approve. Trusted identity and action scope are indispensable.

Easy to make mistakes

  • Process sleep waits for approval
  • Only record approved=true
  • Double-click the callbacks to create and run them individually.

References

It is designed based on public technical information; the reference materials support the technical mechanism, and the scenarios and scoring standards are designed by this website and do not represent the original interview questions of a certain company. New Q&A and migration cases are added for principle explanation, and source verification and case operation verification are recorded separately.

Check how far you understand

After reading, you can explain the principles, boundaries, and trade-offs against these standards. It is up to you to evaluate your mastery; if further verification is needed, complete the small tasks below.

Basic standards met
Design wait states that persist across reboots and independently verify approval identities.
Intermediate and advanced signals
Ability to bind summaries, validity periods and handle repeat callbacks.
Senior criteria
Covers recovery reentrancy, concurrent collection, unknown side effects and permission changes.

Continue to do advanced research experiments

After the process really exits, how can it continue?

From the idempotent counterexample of two databases, continue to verify the lease, checkpoint and independent server receipt. Unzip the reliability experiment v3 and execute it in a separate directory.

Read full text and fault analysis → · Download Reliability Experiment v3 ↓

python3 cli.py submit --db crash.sqlite
python3 cli.py run --db crash.sqlite --lease-seconds 2 --fault after_collect
python3 cli.py inspect --db crash.sqlite
# 首次运行预期退出码 75;等待至少 2 秒后分别执行
python3 cli.py run --db crash.sqlite
python3 evaluate.py --db crash.sqlite

Keep evidence and check item by item

  • First inspect shows running, collect checkpoint saved.
  • After the lease expires, succeeded, generation=2, collect is only submitted once.
  • Press README and run after_effect to check that there is only one receipt in the independent publisher database.

Fixed collect → draft → verify → publish flow; verifying local persistence protocol with real process exit does not prove that any remote service executes exactly once.

Hands-on verificationComplete on demand · Suggestions15 minutes

Write conditional update pseudocode for the approval interface to simulate two consent requests arriving at the same time.

Expand acceptance requirements and checkpoints
  • No repeated consumption for the same approval
  • Action changes will refuse recovery
  • Canceled tasks are not restarted

Key inspections

  • Wait for persistence across processes
  • Approve binding actions and consume atomically
  • Verify changes and concurrency before recovery