Agent Application DevelopmentAccount
Knowledge catalogChoose core direction and segmented content
knowledge unit 32IntermediateImplementationAbout 12 minutes

Understand → Implement → Debug → Design

One state machine for retries, deadlines, cancellation, and budgets

Incorporate total deadlines, single timeouts, error classification, budget reservations, and collaborative cancellation into the same execution strategy.

Try againCanceltimeoutbudgetasyncio

Knowledge content check2026-10-03 · Check the source of the original question2026-10-02

Which step do you want to learn from this knowledge point?

Select the starting point based on the current basis, or you can go deeper one by one. When you encounter an unfamiliar concept, go back to the core principles first; use the knowledge exercises to check your understanding when you are finished.

Understand first

New to this knowledge point

Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.

Start with core principles →

Implement next

Prepare to write the principles into code

Understand implementation steps and boundaries, complete small tasks, and check results against acceptance requirements.

View the code example →

Debug failures

Need to handle failures and changes in conditions

Follow the continuous questioning to locate the failure premise, and then compare the migration cases to explain how the plan should be adjusted.

Continue to delve deeper into the problem →

Compare designs

Need to design or review plans

Combine engineering deductions and senior self-evaluation standards to explain the applicable conditions, costs and alternatives of the plan.

Analyze engineering scenarios →
Knowledge unit directory

LEARN · PRACTICE · REFLECT

Knowledge learning and personal records

My notes and review ↗

First read along the principles, Q&A and migration cases. When you need to check your understanding, switch to reinforcement exercises or start personal recording.

Answers and personal notes

Each modified commit will be kept as an independent history. Your level of mastery is up to you to evaluate yourself against the standards.

Core concept · One state machine for retries, deadlines, cancellation, and budgets

Understand the core principles first

Preparatory concepts:HTTP timeout semantics, Coroutine cancellation, Atomic Counting and Budgeting

A retry is another attempt at the same logical operation and must be subject to the overall deadline, budget, and current authorization. Timeout does not prove failure, cancellation does not undo actions that have taken place, and all boundaries must enter the state machine.

Four controls bound different resources

A run deadline bounds total duration; a call timeout bounds one wait; retry policy chooses retryable errors; a budget bounds consumption. Ten-second calls can still run indefinitely with unlimited retries. Backoff consumes the overall deadline. Validation and authorization failures usually need correction rather than retries.

Unknown outcomes constrain retries

A timed-out write may have succeeded. Query or retry idempotently using the same operation key. Backoff and jitter reduce retry bursts without deduplicating effects. Bound attempts and total waiting, use circuit breakers, and reserve budget before calls. Unknown usage cannot immediately release every reservation for other branches to spend.

Concurrency requires atomic reservations

Two branches may each see ten yuan available and each spend eight. An authoritative shared ledger must check and reserve atomically, recording reservation_id and settling against verifiable usage. Estimates are not upper bounds: strict budgets also limit output and tool allowances. Insufficient reservations require additional authorization or stopping, rather than only a later alert.

Cancellation must propagate through execution

asyncio cancellation takes effect at cooperative points; perform cleanup and re-raise CancelledError. Blocking SDK calls stall the loop, and moving work to a thread does not make that thread forcibly cancellable. Long remote calls need SDK cancellation, outcome queries, or process isolation. Temporal Activity cancellation relies on heartbeat delivery; verify the particular SDK behavior. Test cancellation before calls, after dispatch, after response loss, and before settlement, recording task state and actual effects separately.

Check understanding with a question

How to uniformly design Agent's retry, cancellation, timeout and cost budget?

Separate total deadlines from per-call timeouts. Retry transient errors within bounded backoff, using stable idempotency keys and counting waits against the deadline. Reserve budget atomically before calls and settle actual usage afterward. Propagate cancellation through tasks and tools, recording existing effects without claiming rollback. Clean up and re-raise CancelledError; heartbeat-dependent activities need timely heartbeat delivery.

Implementation and trade-offs

Four constraints should share the same execution context

The context contains at least deadline, cancel_requested, attempts_left, money_reserved and operation_id. The overall deadline covers queuing, backoff, tool execution, and result collection; a single timeout limits only one attempt. Setting three retries but allowing ten minutes each time will make users wait far longer than expected. Cancellation is a control signal, not a reason to retry due to tool exception. The task status can pass cancel_requested and then enter canceled. During this period, necessary cleanup and reconciliation are allowed, but continued planning of new operations is prohibited.

Retry with error semantics instead of catching all exceptions

If the network is temporarily unavailable, rate limiting, etc., you can have limited backoff and respect the server’s retry-delay guidance; authentication failure, parameter errors, and business rejections should directly fail or request supplementary information. A timeout on a write operation may mean that the response was lost when it was completed and must be queried first or retried using idempotency. Temporal documentation distinguishes between Activity retries and Workflow retries, and the default behavior does not replace the application's overall deadline and error classification. Record retry counts, total waiting time, and tool charges in the trace to make it easy to locate cost and service issues. Adding jitter when concurrency is high can reduce the impact of collective retries at the same time.

The budget must be reserved first and then settled. Cancellation cannot be automatically refunded.

If both branches start calling after reading the remaining budget, they will both exceed the budget. Before calling, deduct the available quota and increase the reserved quota based on atomic conditions; release the difference after getting the real billing data. Calls already dispatched may still be billed after cancellation, and the entire reserved quota cannot be refunded before allowing other branches to consume. If the bill is temporarily unknown, keep a conservative reservation and check it asynchronously. Model token, external search, storage or code sandbox costs may be measured separately. The allowable estimation error and hard limit policy should be clearly written. If real-time accurate billing is not possible, the absolute limit cannot be promised.

Test cancellation propagation with injected failures

The cancellation of Python asyncio is triggered at the next responsive opportunity of the coroutine, and cleanup should be placed in finally; swallowing CancelledError may destroy structured concurrency. Synchronous blocking functions, thread tasks, and remote services do not necessarily stop when the coroutine is canceled. Temporal long-running activities get cancellation notifications through heartbeats, which may also have propagation delays. Acceptance includes cancellation during backoff, mid-call cancellation, result return coincident with cancellation, budget just exhausted, and failure response unknown. The following code only limits the number of standard library coroutine attempts and the total time; it does not simulate real billing, cross-process atomic budgeting, or remote cancellation.

code example

Limit the number of retries and the total deadline at the same time

Requires Python 3.11+. Only retries for explicit temporary exceptions; single timeout fails directly by default, and cancellation will not be swallowed. Remote side effects and billing budgets are implemented separately.

import asyncio
class TransientError(Exception): pass
async def run(tool, max_attempts=3, total_seconds=1):
    async with asyncio.timeout(total_seconds):
        for attempt in range(1, max_attempts + 1):
            try:
                async with asyncio.timeout(0.2):
                    return await tool(attempt)
            except TransientError:
                if attempt == max_attempts: raise
                await asyncio.sleep(min(0.01 * 2**(attempt-1), 0.05))
    raise RuntimeError('unreachable')
async def tool(attempt):
    if attempt == 1: raise TransientError('temporary outage')
    return 'ok after 2 attempts'
print(asyncio.run(run(tool)))

expected output

ok after 2 attempts

Engineering deduction

scene
Hypothetical engineering scenario: The research agent encounters current limit when calling the search API, and the user subsequently cancels.
design decisions
Sharing deadlines and cancellation signals, backoffs can be canceled, requests issued to hold billing reservations and verified.
Verify target
The acceptance goal is to no longer initiate new searches after cancellation, and the task retains evidence of completion and the number of real attempts.
applicable boundary
When the server does not support revocation, searches that have been issued may still be completed and billed.

Continuous questions and answers

Continue reading along with the premises and constraints of the problem. Understand the reference answers first, then try to put away the answers and explain the cause and effect and trade-offs in your own words.

Draw inferences from one example: If the conditions change, how to deduce it?

First find out the conditions for change, and then determine which premises in the original plan still hold true. The following cases are teaching deductions to facilitate the transfer of principles to new problems.

Read-only expensive search

Changing conditions:No side effects written, but billed per request

Extended question:Do we still need idempotency and budget when retrying after timeout?

Derivation and reference solutions

Read requests usually do not repeat business writes, but will repeatedly consume fees, rate-limit quota, and time. The budget is reserved on a per-try basis, and whether caching or request deduplication saves money depends on the service contract. The total deadline and backoff should still be uniformly limited; infinite retries should not be caused by read-only.

The principles that remain unchanged:Retryability and resource affordability are different conditions, and all attempts are still subject to the overall constraint.

As soon as two branches win, the other one is cancelled.

Changing conditions:Cancellation is used to reduce delays and costs

Extended question:Are reservations for canceled branches released immediately?

Derivation and reference solutions

If it has not been sent and the stop has been confirmed, it can be released; if it has been sent but the usage has not been confirmed, it will be retained and reconciled to avoid repeated allocation of the same quota. Successful external effects are recorded and cannot be ignored on the basis that another path has won. Settlement is idempotent to avoid repeated budget returns for canceled callbacks and successful callbacks.

The principles that remain unchanged:Just because the task no longer requires results does not mean that costs and side effects have not occurred.

Easy to make mistakes

  • Always retry permission errors and cancellations
  • Reset the total deadline for each retry
  • Immediately after cancellation, it is determined that the remote operation did not occur
  • The budget is checked first and then called, without atomicity reserved.

References

It is designed based on public technical information; the reference materials support the technical mechanism, and the scenarios and scoring standards are designed by this website and do not represent the original interview questions of a certain company. New Q&A and migration cases are added for principle explanation, and source verification and case operation verification are recorded separately.

Check how far you understand

After reading, you can explain the principles, boundaries, and trade-offs against these standards. It is up to you to evaluate your mastery; if further verification is needed, complete the small tasks below.

Basic standards met
Can distinguish retryable, permanent failure and user cancellation by error type.
Intermediate and advanced signals
Unify deadlines, backoffs, frequency and cost budgets.
Senior criteria
Handle concurrency budget reservation, cross-layer retry amplification and external action verification.

View verification records for independent examples

Continue to do advanced research experiments

After the process really exits, how can it continue?

From the idempotent counterexample of two databases, continue to verify the lease, checkpoint and independent server receipt. Unzip the reliability experiment v3 and execute it in a separate directory.

Read full text and fault analysis → · Download Reliability Experiment v3 ↓

python3 cli.py submit --db crash.sqlite
python3 cli.py run --db crash.sqlite --lease-seconds 2 --fault after_collect
python3 cli.py inspect --db crash.sqlite
# 首次运行预期退出码 75;等待至少 2 秒后分别执行
python3 cli.py run --db crash.sqlite
python3 evaluate.py --db crash.sqlite

Keep evidence and check item by item

  • First inspect shows running, collect checkpoint saved.
  • After the lease expires, succeeded, generation=2, collect is only submitted once.
  • Press README and run after_effect to check that there is only one receipt in the independent publisher database.

Fixed collect → draft → verify → publish flow; verifying local persistence protocol with real process exit does not prove that any remote service executes exactly once.

Hands-on verificationComplete on demand · Suggestions15 minutes

Write processing matrices for model throttling, permission denial, and tool timeouts.

Expand acceptance requirements and checkpoints
  • Do not blindly retry permission denials
  • Retries do not reset total deadline
  • No new action will be started after cancellation

Key inspections

  • Able to distinguish temporary errors, permanent errors and unknown results
  • Ability to limit both single execution and complete task time
  • Can explain the propagation of cancellations and the processing of operations that have occurred
  • Can explain why budget reservation for concurrent calls requires atomicity