Understand first
New to this knowledge point
Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.
Start with core principles →Understand → Implement → Debug → Design
Incorporate total deadlines, single timeouts, error classification, budget reservations, and collaborative cancellation into the same execution strategy.
Knowledge content check2026-10-03 · Check the source of the original question2026-10-02
It is recommended to understand first:
Tool calls: structure, authorization, and business contracts →Select the starting point based on the current basis, or you can go deeper one by one. When you encounter an unfamiliar concept, go back to the core principles first; use the knowledge exercises to check your understanding when you are finished.
New to this knowledge point
Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.
Start with core principles →Prepare to write the principles into code
Understand implementation steps and boundaries, complete small tasks, and check results against acceptance requirements.
View the code example →Need to handle failures and changes in conditions
Follow the continuous questioning to locate the failure premise, and then compare the migration cases to explain how the plan should be adjusted.
Continue to delve deeper into the problem →Need to design or review plans
Combine engineering deductions and senior self-evaluation standards to explain the applicable conditions, costs and alternatives of the plan.
Analyze engineering scenarios →LEARN · PRACTICE · REFLECT
First read along the principles, Q&A and migration cases. When you need to check your understanding, switch to reinforcement exercises or start personal recording.
Can be practiced directly. After logging in, answers, favorites, and notes will be saved to your account.
Log in and saveEach modified commit will be kept as an independent history. Your level of mastery is up to you to evaluate yourself against the standards.
Core concept · One state machine for retries, deadlines, cancellation, and budgets
Preparatory concepts:HTTP timeout semantics, Coroutine cancellation, Atomic Counting and Budgeting
A retry is another attempt at the same logical operation and must be subject to the overall deadline, budget, and current authorization. Timeout does not prove failure, cancellation does not undo actions that have taken place, and all boundaries must enter the state machine.
A run deadline bounds total duration; a call timeout bounds one wait; retry policy chooses retryable errors; a budget bounds consumption. Ten-second calls can still run indefinitely with unlimited retries. Backoff consumes the overall deadline. Validation and authorization failures usually need correction rather than retries.
A timed-out write may have succeeded. Query or retry idempotently using the same operation key. Backoff and jitter reduce retry bursts without deduplicating effects. Bound attempts and total waiting, use circuit breakers, and reserve budget before calls. Unknown usage cannot immediately release every reservation for other branches to spend.
Two branches may each see ten yuan available and each spend eight. An authoritative shared ledger must check and reserve atomically, recording reservation_id and settling against verifiable usage. Estimates are not upper bounds: strict budgets also limit output and tool allowances. Insufficient reservations require additional authorization or stopping, rather than only a later alert.
asyncio cancellation takes effect at cooperative points; perform cleanup and re-raise CancelledError. Blocking SDK calls stall the loop, and moving work to a thread does not make that thread forcibly cancellable. Long remote calls need SDK cancellation, outcome queries, or process isolation. Temporal Activity cancellation relies on heartbeat delivery; verify the particular SDK behavior. Test cancellation before calls, after dispatch, after response loss, and before settlement, recording task state and actual effects separately.
Separate total deadlines from per-call timeouts. Retry transient errors within bounded backoff, using stable idempotency keys and counting waits against the deadline. Reserve budget atomically before calls and settle actual usage afterward. Propagate cancellation through tasks and tools, recording existing effects without claiming rollback. Clean up and re-raise CancelledError; heartbeat-dependent activities need timely heartbeat delivery.
The context contains at least deadline, cancel_requested, attempts_left, money_reserved and operation_id. The overall deadline covers queuing, backoff, tool execution, and result collection; a single timeout limits only one attempt. Setting three retries but allowing ten minutes each time will make users wait far longer than expected. Cancellation is a control signal, not a reason to retry due to tool exception. The task status can pass cancel_requested and then enter canceled. During this period, necessary cleanup and reconciliation are allowed, but continued planning of new operations is prohibited.
If the network is temporarily unavailable, rate limiting, etc., you can have limited backoff and respect the server’s retry-delay guidance; authentication failure, parameter errors, and business rejections should directly fail or request supplementary information. A timeout on a write operation may mean that the response was lost when it was completed and must be queried first or retried using idempotency. Temporal documentation distinguishes between Activity retries and Workflow retries, and the default behavior does not replace the application's overall deadline and error classification. Record retry counts, total waiting time, and tool charges in the trace to make it easy to locate cost and service issues. Adding jitter when concurrency is high can reduce the impact of collective retries at the same time.
If both branches start calling after reading the remaining budget, they will both exceed the budget. Before calling, deduct the available quota and increase the reserved quota based on atomic conditions; release the difference after getting the real billing data. Calls already dispatched may still be billed after cancellation, and the entire reserved quota cannot be refunded before allowing other branches to consume. If the bill is temporarily unknown, keep a conservative reservation and check it asynchronously. Model token, external search, storage or code sandbox costs may be measured separately. The allowable estimation error and hard limit policy should be clearly written. If real-time accurate billing is not possible, the absolute limit cannot be promised.
The cancellation of Python asyncio is triggered at the next responsive opportunity of the coroutine, and cleanup should be placed in finally; swallowing CancelledError may destroy structured concurrency. Synchronous blocking functions, thread tasks, and remote services do not necessarily stop when the coroutine is canceled. Temporal long-running activities get cancellation notifications through heartbeats, which may also have propagation delays. Acceptance includes cancellation during backoff, mid-call cancellation, result return coincident with cancellation, budget just exhausted, and failure response unknown. The following code only limits the number of standard library coroutine attempts and the total time; it does not simulate real billing, cross-process atomic budgeting, or remote cancellation.
Requires Python 3.11+. Only retries for explicit temporary exceptions; single timeout fails directly by default, and cancellation will not be swallowed. Remote side effects and billing budgets are implemented separately.
import asyncio
class TransientError(Exception): pass
async def run(tool, max_attempts=3, total_seconds=1):
async with asyncio.timeout(total_seconds):
for attempt in range(1, max_attempts + 1):
try:
async with asyncio.timeout(0.2):
return await tool(attempt)
except TransientError:
if attempt == max_attempts: raise
await asyncio.sleep(min(0.01 * 2**(attempt-1), 0.05))
raise RuntimeError('unreachable')
async def tool(attempt):
if attempt == 1: raise TransientError('temporary outage')
return 'ok after 2 attempts'
print(asyncio.run(run(tool)))
expected output
ok after 2 attemptsContinue reading along with the premises and constraints of the problem. Understand the reference answers first, then try to put away the answers and explain the cause and effect and trade-offs in your own words.
Level 1Why can't HTTP timeout directly determine that the tool failed?
Drill down from unified strategy to retry the most critical outcome uncertainty.
Timeout only means that the caller did not get confirmation within the time limit, and the remote end may not have started, has succeeded, or is executing. Read operations can be retried according to the policy; write operations keep the operation_id and query receipts or retry idempotently, and new keys cannot be generated. If the target cannot be verified, mark unknown and pause instead of recording request exceptions as business failures.
Level 1How do two Agent branches reserve remaining budgets atomically?
Multiple branches turn budget read-after-write into a double-spend risk.
Use conditional updates or short transactions to check available >= reservation in the shared ledger, deduct available and create unique reservation records. Concurrent branches compete for the same authoritative balance, and will be called only by those who successfully reserve; retry to inherit or explicitly add reservations, and the same reservation cannot be deducted repeatedly. The difference will be settled after the actual usage is returned, and the amount for unknown usage will be retained until reconciliation.
Follow this answer further
Level 2If five yuan is reserved but seven yuan is actually used, is the atomic reservation still considered effective control?
After the father asked to solve the concurrent reservation, the cost estimation error became a new constraint.
It prevents concurrent duplicate allocation, but it does not constitute a strict upper limit if the reservation is only an estimate. Limit the maximum token or tool resources before calling to form a conservative upper bound, or explicitly allow limited overage and design additional reservations. When the actual settlement exceeds the reservation, new calls will be blocked and an alarm will be issued. You cannot use subsequent negative deductions to claim that the budget has never been exceeded.
Follow this answer further
Level 3The response is lost and the actual usage cannot be found. Can the reservation remain stuck?
Beyond the upper cost bound, incomplete settlement information requires the definition of an observable final state.
Set unknown usage status, reconciliation period and manual processing path; use it conservatively before confirmation under strict budget to avoid spending again. If the business allows risk reserves, clarify the upper bound and write-off strategy, but it cannot be silently reset to zero. Persistent unknowns require alerting or stopping new calls to the provider, rather than relying on retries while inflating uncertainty charges.
Level 1How can the synchronization SDK ensure that it can be canceled when it blocks the event loop?
There are different boundaries between canceling a coroutine and canceling the underlying operation.
Prioritize the use of asynchronous cancelable SDK; I/O blocking can use asyncio.to_thread to keep the event loop responsive, but canceling the wait will not automatically terminate the in-thread function. Set the actual request timeout and cooperative cancellation signal for the SDK, and use an isolation process if necessary; the remote end still needs to reconcile after it has been submitted. You cannot use wait_for to prove that the underlying action has stopped.
First find out the conditions for change, and then determine which premises in the original plan still hold true. The following cases are teaching deductions to facilitate the transfer of principles to new problems.
Changing conditions:No side effects written, but billed per request
Extended question:Do we still need idempotency and budget when retrying after timeout?
Read requests usually do not repeat business writes, but will repeatedly consume fees, rate-limit quota, and time. The budget is reserved on a per-try basis, and whether caching or request deduplication saves money depends on the service contract. The total deadline and backoff should still be uniformly limited; infinite retries should not be caused by read-only.
The principles that remain unchanged:Retryability and resource affordability are different conditions, and all attempts are still subject to the overall constraint.
Changing conditions:Cancellation is used to reduce delays and costs
Extended question:Are reservations for canceled branches released immediately?
If it has not been sent and the stop has been confirmed, it can be released; if it has been sent but the usage has not been confirmed, it will be retained and reconciled to avoid repeated allocation of the same quota. Successful external effects are recorded and cannot be ignored on the basis that another path has won. Settlement is idempotent to avoid repeated budget returns for canceled callbacks and successful callbacks.
The principles that remain unchanged:Just because the task no longer requires results does not mean that costs and side effects have not occurred.
It is designed based on public technical information; the reference materials support the technical mechanism, and the scenarios and scoring standards are designed by this website and do not represent the original interview questions of a certain company. New Q&A and migration cases are added for principle explanation, and source verification and case operation verification are recorded separately.
After reading, you can explain the principles, boundaries, and trade-offs against these standards. It is up to you to evaluate your mastery; if further verification is needed, complete the small tasks below.
View verification records for independent examples
From the idempotent counterexample of two databases, continue to verify the lease, checkpoint and independent server receipt. Unzip the reliability experiment v3 and execute it in a separate directory.
Read full text and fault analysis → · Download Reliability Experiment v3 ↓
python3 cli.py submit --db crash.sqlite
python3 cli.py run --db crash.sqlite --lease-seconds 2 --fault after_collect
python3 cli.py inspect --db crash.sqlite
# 首次运行预期退出码 75;等待至少 2 秒后分别执行
python3 cli.py run --db crash.sqlite
python3 evaluate.py --db crash.sqliteFixed collect → draft → verify → publish flow; verifying local persistence protocol with real process exit does not prove that any remote service executes exactly once.
Write processing matrices for model throttling, permission denial, and tool timeouts.