Understand first
New to this knowledge point
Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.
Start with core principles →Understand → Implement → Debug → Design
Examine asynchronous concurrency, deadlines, cancellation propagation, and partial results.
Knowledge content check2026-10-03 · Check the source of the original question2026-10-02
It is recommended to understand first:
Control and completion criteria in agent loops →Tool calls: structure, authorization, and business contracts →Select the starting point based on the current basis, or you can go deeper one by one. When you encounter an unfamiliar concept, go back to the core principles first; use the knowledge exercises to check your understanding when you are finished.
New to this knowledge point
Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.
Start with core principles →Prepare to write the principles into code
Understand implementation steps and boundaries, complete small tasks, and check results against acceptance requirements.
View the code example →Need to handle failures and changes in conditions
Follow the continuous questioning to locate the failure premise, and then compare the migration cases to explain how the plan should be adjusted.
Continue to delve deeper into the problem →Need to design or review plans
Combine engineering deductions and senior self-evaluation standards to explain the applicable conditions, costs and alternatives of the plan.
Analyze engineering scenarios →LEARN · PRACTICE · REFLECT
First read along the principles, Q&A and migration cases. When you need to check your understanding, switch to reinforcement exercises or start personal recording.
Can be practiced directly. After logging in, answers, favorites, and notes will be saved to your account.
Log in and saveEach modified commit will be kept as an independent history. Your level of mastery is up to you to evaluate yourself against the standards.
Core concept · Delivery conditions and shared deadlines for parallel results
Preparatory concepts:event loop, Structured concurrency, Request timeout
Parallelism only shortens independent waiting and does not change task dependencies and success conditions. Global deadlines constrain all attempts and backoffs; the cancellation of a local task does not prove that the remote side effects have stopped, and the results must be retained and classified item by item.
News searches may permit a missing region; pre-publication checks may all be mandatory. Specify required results at task creation. Declaring failed checks optional afterward invents success. Calls depending on another tool’s output cannot start simultaneously.
Declaring async does not make blocking HTTP nonblocking. Use asynchronous I/O or controlled threads/processes with underlying timeouts. Blocking work can delay all tools, timers, and cancellation. Cancelling a coroutine may not stop its thread operation.
By default, Python gather propagates the first exception while other tasks continue. TaskGroup cancels remaining tasks after failure and waits for cleanup. Choose according to all-required or partial-result policy. Use a monotonic clock for a shared deadline; retries, queueing, and backoff consume its remaining duration.
Use dependencies to choose parallel calls and define all-required versus partial success upfront. Share an overall deadline and allocate remaining time and concurrency. Preserve valid independent read results; coordinate dependent calls and writes more strictly. Local cancellation does not establish that remote work stopped. Report success, failure, unknown timeout, and cancellation separately.
If the three tools obtain news from three regions respectively, partial results with gaps can be accepted; if they together constitute a pre-publication verification, any key verification failure should prevent publication. Put the required tag into the task definition so that it does not temporarily determine whether it is successful after an exception occurs. Calls that rely on the output of another tool must wait until the input has been validated and cannot be hard-parallelized for speed.
The task entry calculates the absolute deadline, uses the remaining time for each step, and does not reset the complete timeout for each retry. Use the concurrency upper limit to limit the connections and quotas occupied at the same time, and the timeout wait is also included in the budget. The asynchronous library's task group, aggregate waiting and cancellation behaviors are different, and the selected semantics should be clearly defined; synchronous blocking functions may block the event loop even if they are written in async functions.
The results are encapsulated into tool ID, status, data, time consumption, error type and whether external actions need to be checked. It is not necessary to discard all successful results when a single independent read task fails; unstarted actions are canceled when a critical task fails. The write request that has been sent will enter the receipt check, and the CancelledError cannot be directly interpreted as the business cancellation was successful. Cleaning up connections and releasing concurrency slots are placed on the cleanup path that can be executed.
Simulation A succeeds in 100ms, B fails immediately, C waits forever, and the global deadline is 500ms. The verifier exits bounded, valid results from A are retained, errors from B are visible, and C consumes no legacy resources. If all requirements are successful, the final must not be marked as succeeded. Further let C complete the remote write after cancellation and check whether the candidate distinguishes between local coroutine state and external facts.
Minimal demonstration of a read-only asynchronous task; remote write requests are not guaranteed to be canceled with local cancellation.
import asyncio
async def ok():
return "result"
async def fail():
raise RuntimeError("unavailable")
async def slow():
await asyncio.Event().wait()
async def main():
tasks = {"a": asyncio.create_task(ok()), "b": asyncio.create_task(fail()),
"c": asyncio.create_task(slow())}
done, pending = await asyncio.wait(tasks.values(), timeout=0.02)
for task in pending:
task.cancel()
await asyncio.gather(*pending, return_exceptions=True)
for name, task in tasks.items():
if task in pending:
status = "timeout"
elif task.exception() is not None:
status = "failed"
else:
status = "succeeded"
print(name, status)
asyncio.run(main())
expected output
a succeeded
b failed
c timeoutContinue reading along with the premises and constraints of the problem. Understand the reference answers first, then try to put away the answers and explain the cause and effect and trade-offs in your own words.
Level 1What's the problem with calling the synchronous HTTP library in the async function?
Parallelism requirements yield execution rights, which is not guaranteed by functional syntax.
The synchronization library occupies the event loop thread, and other coroutines cannot be scheduled, and even the timeout may be delayed. Use an asynchronous client instead, or put blocking calls into a capacity-limited thread/process and set connect, read, and overall timeouts; just the async wrapper won't work. Thread cancellation usually does not terminate its network request, and in-flight writes still need to be checked.
Level 1Are other tasks still running after aggregation wait encounters an exception?
Error handling changes the concurrent task life cycle and requires checking the real library semantics.
Look at the primitives chosen. gather will throw the first exception by default, but other awaitables will not be automatically canceled; return_exceptions=True will collect itemized exceptions. When a TaskGroup encounters a non-cancellation exception, it will cancel the remaining tasks and wait. In either case, it is necessary to clearly understand the result collection, cleaning and in-transit side effects. It cannot be assumed that throwing an exception equals a complete stop.
Follow this answer further
Level 2I want to retain the successful result of A and cancel C when B fails. How to achieve this?
After understanding the primitives, you need to translate the library behavior into an application delivery strategy.
Save the structured results of each task and decide when to cancel according to required rules. If the B key fails, record the completed data of A for diagnosis, cancel C and wait for limited cleanup; if B is optional, continue to wait for C until the global deadline. Exception encapsulation must retain the failure category, and cannot misjudge all successes after treating it as ordinary data.
Follow this answer further
Level 3C swallows the cancellation exception and continues to run. Can TaskGroup guarantee timely return?
Controlling a task group does not mean controlling all subcode and remote side effects.
No. Cancellation is a cooperative mechanism. Subtasks swallowing signals or executing blocking code will delay the completion. Corrected subtask cancellation processing and propagation of exceptions after cleanup; use process isolation and external timeout for untrusted or uninterruptible tasks. If a remote write request has been issued, reconciliation is required even if the local process is terminated.
Level 1Why does the tool retry reusing the remaining time?
Local retries must obey the global delivery budget.
Resetting the full timeout each time will cause the total time to balloon with the number of attempts. Subtract the current monotonic clock from the entry deadline to get the remaining amount, including backoff, queuing and connection establishment; when the remaining amount is insufficient, stop retrying and return a clear partial status. When the amount or remote effect is unknown, it is marked separately and cannot be displayed as unexecuted.
First find out the conditions for change, and then determine which premises in the original plan still hold true. The following cases are teaching deductions to facilitate the transfer of principles to new problems.
Changing conditions:Independent data reading becomes release threshold
Extended question:Can slow check timeout be released first?
No. Failure to pass any required check will prevent release, and the success evidence and failure reasons will be retained for repair; a new round of rechecking may be invalid versions. Performance optimizations can adjust concurrency and check ranges, but cannot sneakily modify success definitions after timeouts.
The principles that remain unchanged:Parallelism does not change dependencies and acceptance.
Changing conditions:All read-only becomes read-write mixed
Extended question:Will canceling all tasks after expiration end it?
Stop the unstarted operation, cancel to cancel the waiting, and submit the write to enter the result verification. Successful read results can be saved, but a summary that relies on the write results cannot generate a definite conclusion; when the receipt is late, the ledger is updated according to the operation ID. Mixed reading and writing requires separation of business state and coroutine state.
The principles that remain unchanged:Local concurrency state is not external business state.
It is designed based on public technical information; the reference materials support the technical mechanism, and the scenarios and scoring standards are designed by this website and do not represent the original interview questions of a certain company. New Q&A and migration cases are added for principle explanation, and source verification and case operation verification are recorded separately.
After reading, you can explain the principles, boundaries, and trade-offs against these standards. It is up to you to evaluate your mastery; if further verification is needed, complete the small tasks below.
View verification records for independent examples
Write pseudocode to call three read-only tools in parallel and return item-by-item status at the 500ms deadline.