Understand first
New to this knowledge point
Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.
Start with core principles →Understand → Implement → Debug → Design
Starting from the tool invocation loop, design termination conditions, budgets, cancellations, and verifiable completion statuses.
Knowledge content check2026-10-03 · Check the source of the original question2026-10-02
It is recommended to understand first:
Your first model call and response contract →Select the starting point based on the current basis, or you can go deeper one by one. When you encounter an unfamiliar concept, go back to the core principles first; use the knowledge exercises to check your understanding when you are finished.
New to this knowledge point
Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.
Start with core principles →Prepare to write the principles into code
Understand implementation steps and boundaries, complete small tasks, and check results against acceptance requirements.
Reading implementation and trade-offs →Need to handle failures and changes in conditions
Follow the continuous questioning to locate the failure premise, and then compare the migration cases to explain how the plan should be adjusted.
Continue to delve deeper into the problem →Need to design or review plans
Combine engineering deductions and senior self-evaluation standards to explain the applicable conditions, costs and alternatives of the plan.
Analyze engineering scenarios →LEARN · PRACTICE · REFLECT
First read along the principles, Q&A and migration cases. When you need to check your understanding, switch to reinforcement exercises or start personal recording.
Can be practiced directly. After logging in, answers, favorites, and notes will be saved to your account.
Log in and saveEach modified commit will be kept as an independent history. Your level of mastery is up to you to evaluate yourself against the standards.
Core concept · Control and completion criteria in agent loops
Preparatory concepts:Request life cycle, Asynchronous cancellation, Business acceptance
Model output proposes the next step; the runtime owns execution and terminal state. Stopping only establishes that the loop ended. Success additionally requires evidence that the agreed business conditions hold. Budget exhaustion, cancellation, and unknown outcomes therefore need distinct labels rather than one generic 'completed' state.
A conventional backend follows a defined program after receiving a request. An agent adds model-proposed next actions. The decision source changes, but server responsibilities for permissions, cost, and effects remain. The loop generates candidate actions, checks them at the gateway, executes them, and updates recorded facts. Final prose is a candidate artifact.
One round can make several calls, each with retries. Limiting rounds alone resembles limiting transactions without limiting their SQL operations. Bound model rounds, actual call attempts, concurrency, and absolute deadlines separately; account for actual usage. OpenAI Runner’s maximum turns protects the loop, while the application must enforce business budgets.
“Fixed” does not establish a passing test; “cancelled” does not erase a submitted release. Separate task state from operation states so a stopped task can retain completed effects and unresolved outcomes. A test model that declares completion early or repeatedly requests tools can expose control failures more reliably than one successful live-model run.
The model proposes actions; runtime executes permitted tools, returns results to context, and continues the loop. Runtime independently enforces permissions, rounds, time, and cost. Business evidence determines success. Budget exhaustion and tool failure are distinct terminal outcomes. Record the stopping reason and verified artifacts so users can see completed and outstanding work.
A model response may contain the final answer, or it may contain one or more tool calls. The runtime is responsible for dispatching calls, collecting results, and deciding whether to continue. OpenAI Agents Runner appends the tool results and calls the model again, and provides a max_turns limit; this shows that "letting the model stop itself" is not a complete control strategy. Multiple tool executions within a model round need to be counted independently, and retries within the tool cannot be hidden outside the budget. Model input should include goals, allowed actions and existing evidence. Authorization identities, keys and real budgets are saved on the server side.
I would differentiate between succeeded, failed, canceled, waiting_approval and budget_exhausted. Success needs to meet acceptance requirements, for example, the target file exists, the content format is legal, and necessary checks pass; the model says "completed" and only generates candidates to be accepted. The budget must include at least the number of model rounds, tool calls, total deadline, and cost cap. Estimated costs are reserved before each call, and settled according to actual usage after the call; concurrent requests share the same budget ledger to prevent each branch from thinking that there is still a balance. After the limit is exceeded, the completed artifacts and remaining tasks are saved, and retry cannot continue into an infinite loop.
Cancel first blocks new actions and then propagates the signal to running cancelable calls. For external write operations that have been issued, the results need to be verified, and the remote end cannot be declared not executed due to local timeout. Tool returns are classified by recoverable errors, permission denial, and permanent business errors; only operations that are truly retryable and idempotent are retried. Consecutive failures with the same tool and normalization parameters should be escalated to diagnostics or exit rather than feeding the same error to the model dozens of times. Event records run_id, turn, tool_call_id, time consumption and stop reason, and sensitive content is redacted.
Simulate the model with a fixed response sequence: keep asking the same tool, make multiple calls at once, claim success early, cancel during execution. After the assertion exceeds the limit, the tool will no longer be called. Failure will not generate succeeded. There will be no new side effects after cancellation. If the artifact acceptance fails, you can go back to repair or clear failure. Validate these deterministic controls before using real models to evaluate task completion rates. Giving such counterexamples in interviews can better prove implementation capabilities than reciting "Plan-Action-Reflection".
Continue reading along with the premises and constraints of the problem. Understand the reference answers first, then try to put away the answers and explain the cause and effect and trade-offs in your own words.
Level 1One model round generates five tool calls, how to budget?
One round is not equal to one action, and the abstract cycle needs to be dropped into the resource ledger.
Count one model turn and five candidate tool requests separately. Count each execution attempt that actually starts, including retries inside tools. Atomically reserve shared quota before starting; if it is insufficient, reject the batch or execute the permitted subset according to the task contract. Five concurrent branches must not each spend the same remaining balance. Return an explicit reason for unexecuted calls so the model does not mistake them for completed actions.
Follow this answer further
Level 2There are only two call quotas left, and there is a key verification among the five candidates. How should I choose?
Shared counting solves the problem of excess. Next, we need to consider how to maintain the meaning of acceptance when the quota is insufficient.
Determined by task dependencies and priorities configured at runtime, the available budget for critical verifications is guaranteed first; the model is allowed to recommend ordering, but it cannot reduce verifications to optional on its own. If the full cost of a critical step is no longer affordable, it is more honest to save the current artifact and end up underbudgeted than to claim success after performing two irrelevant steps.
Follow this answer further
Level 3A quota has been reserved for critical verification, but a started tool has not been returned. Can the budget be released?
Reservation policies encounter unknown results and need to differentiate between local capacity and external facts.
Potential costs and side effects cannot be immediately dismissed as disappearing due to local wait timeouts. Settle confirmable resources separately, keep unknown operation records, and stop subsequent dependencies. The fee retention time and reconciliation method are determined by the supplier's behavior; the capacity will be released after the cancellation of cancellable pure calculations, and the receipt of the write request that has been issued will be checked first.
Level 1When the user cancels, the remote write operation has been successful. How do you end it?
The cancellation signal acts on the execution process, and the business facts that have occurred need to be ended separately.
Disable new actions first, and then verify submitted requests. If the remote end has succeeded, it will record a receipt, telling the user that the subsequent steps have been cancelled, but the action has taken effect; if the business supports revocation, an independent operation is initiated based on the revocation permission, and whether it is successful is recorded. The original success record cannot be deleted, nor can the issuance of a compensation request be interpreted as restoration to the original status.
Level 1Is a correct answer successful if the model did not run the required tests?
Controlling the final state must come back to the acceptance criteria, not the credibility of the text.
If the task contract requires running the test, it will not be judged as successful. Correct answers or code are only part of the conditions. If there is a lack of test evidence, it will be marked as to be verified. If there is still budget, make up the test. If there is no environment, the delivery scope and unverified items will be clarified. Only after the user explicitly adjusts the acceptance contract can it be completed according to the new scope, and the model cannot unilaterally delete the requirements.
First find out the conditions for change, and then determine which premises in the original plan still hold true. The following cases are teaching deductions to facilitate the transfer of principles to new problems.
Changing conditions:Acceptance changes from automatic testing to evidence coverage and manual judgment
Extended question:How to judge the completion of data research?
Start by defining the scope of coverage, the types of sources that must be consulted, how references can be read back, and how open questions are presented. References and required fields can be automatically checked during runtime, and disputed conclusions are submitted to manual or specialized review. Reaching the scope and truthfully reporting unknowns counts as completion. The model cannot be required to guarantee that nothing is missing in the world, nor can the number of searches be used as a substitute for evidence coverage.
The principles that remain unchanged:Completion always relies on predefined observable conditions, it's just that the validator changes from testing to evidence review.
Changing conditions:One task becomes multiple runs per day
Extended question:After today's failure, can yesterday's success status be automatically used tomorrow?
Each run binds a snapshot of the input with proof of acceptance. The facts published yesterday can be reused, today's data refresh, permissions and budgets are rechecked; the same business actions continue to use the same operation keys, and new date artifacts use new business identifiers. One success cannot overwrite a failure under new input, and the system should display the results of each run separately.
The principles that remain unchanged:Task lifecycle and business operation lifecycle must be separated.
It is designed based on public technical information; the reference materials support the technical mechanism, and the scenarios and scoring standards are designed by this website and do not represent the original interview questions of a certain company. New Q&A and migration cases are added for principle explanation, and source verification and case operation verification are recorded separately.
After reading, you can explain the principles, boundaries, and trade-offs against these standards. It is up to you to evaluate your mastery; if further verification is needed, complete the small tasks below.
Draw a state machine for a tool loop and simulate the model requiring the same failure action three times in a row.