Agent Application DevelopmentAccount
Knowledge catalogChoose core direction and segmented content
knowledge unit 01FundamentalsConceptsAbout 12 minutes

Understand → Implement → Debug → Design

Control and completion criteria in agent loops

Starting from the tool invocation loop, design termination conditions, budgets, cancellations, and verifiable completion statuses.

Agent LoopbudgetTermination conditionCancel

Knowledge content check2026-10-03 · Check the source of the original question2026-10-02

Which step do you want to learn from this knowledge point?

Select the starting point based on the current basis, or you can go deeper one by one. When you encounter an unfamiliar concept, go back to the core principles first; use the knowledge exercises to check your understanding when you are finished.

Understand first

New to this knowledge point

Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.

Start with core principles →

Realize again

Prepare to write the principles into code

Understand implementation steps and boundaries, complete small tasks, and check results against acceptance requirements.

Reading implementation and trade-offs →

Will troubleshoot

Need to handle failures and changes in conditions

Follow the continuous questioning to locate the failure premise, and then compare the migration cases to explain how the plan should be adjusted.

Continue to delve deeper into the problem →

Able to choose

Need to design or review plans

Combine engineering deductions and senior self-evaluation standards to explain the applicable conditions, costs and alternatives of the plan.

Analyze engineering scenarios →
Knowledge unit directory

LEARN · PRACTICE · REFLECT

Knowledge learning and personal records

My notes and review ↗

First read along the principles, Q&A and migration cases. When you need to check your understanding, switch to reinforcement exercises or start personal recording.

Answers and personal notes

Each modified commit will be kept as an independent history. Your level of mastery is up to you to evaluate yourself against the standards.

Core concept · Control and completion criteria in agent loops

Understand the core principles first

Preparatory concepts:Request life cycle, Asynchronous cancellation, Business acceptance

Model output proposes the next step; the runtime owns execution and terminal state. Stopping only establishes that the loop ended. Success additionally requires evidence that the agreed business conditions hold. Budget exhaustion, cancellation, and unknown outcomes therefore need distinct labels rather than one generic 'completed' state.

View the loop from an HTTP request

A conventional backend follows a defined program after receiving a request. An agent adds model-proposed next actions. The decision source changes, but server responsibilities for permissions, cost, and effects remain. The loop generates candidate actions, checks them at the gateway, executes them, and updates recorded facts. Final prose is a candidate artifact.

A round limit is insufficient

One round can make several calls, each with retries. Limiting rounds alone resembles limiting transactions without limiting their SQL operations. Bound model rounds, actual call attempts, concurrency, and absolute deadlines separately; account for actual usage. OpenAI Runner’s maximum turns protects the loop, while the application must enforce business budgets.

Terminal states require evidence

“Fixed” does not establish a passing test; “cancelled” does not erase a submitted release. Separate task state from operation states so a stopped task can retain completed effects and unresolved outcomes. A test model that declares completion early or repeatedly requests tools can expose control failures more reliably than one successful live-model run.

Check understanding with a question

How does the Agent loop stop? How to draw the boundary between model decision-making and runtime control?

The model proposes actions; runtime executes permitted tools, returns results to context, and continues the loop. Runtime independently enforces permissions, rounds, time, and cost. Business evidence determines success. Budget exhaustion and tool failure are distinct terminal outcomes. Record the stopping reason and verified artifacts so users can see completed and outstanding work.

Realization and trade-offs

Actual control of the loop

A model response may contain the final answer, or it may contain one or more tool calls. The runtime is responsible for dispatching calls, collecting results, and deciding whether to continue. OpenAI Agents Runner appends the tool results and calls the model again, and provides a max_turns limit; this shows that "letting the model stop itself" is not a complete control strategy. Multiple tool executions within a model round need to be counted independently, and retries within the tool cannot be hidden outside the budget. Model input should include goals, allowed actions and existing evidence. Authorization identities, keys and real budgets are saved on the server side.

How to design the final state and budget

I would differentiate between succeeded, failed, canceled, waiting_approval and budget_exhausted. Success needs to meet acceptance requirements, for example, the target file exists, the content format is legal, and necessary checks pass; the model says "completed" and only generates candidates to be accepted. The budget must include at least the number of model rounds, tool calls, total deadline, and cost cap. Estimated costs are reserved before each call, and settled according to actual usage after the call; concurrent requests share the same budget ledger to prevent each branch from thinking that there is still a balance. After the limit is exceeded, the completed artifacts and remaining tasks are saved, and retry cannot continue into an infinite loop.

Cancellation and failure cannot only change one flag

Cancel first blocks new actions and then propagates the signal to running cancelable calls. For external write operations that have been issued, the results need to be verified, and the remote end cannot be declared not executed due to local timeout. Tool returns are classified by recoverable errors, permission denial, and permanent business errors; only operations that are truly retryable and idempotent are retried. Consecutive failures with the same tool and normalization parameters should be escalated to diagnostics or exit rather than feeding the same error to the model dozens of times. Event records run_id, turn, tool_call_id, time consumption and stop reason, and sensitive content is redacted.

How to verify control boundaries

Simulate the model with a fixed response sequence: keep asking the same tool, make multiple calls at once, claim success early, cancel during execution. After the assertion exceeds the limit, the tool will no longer be called. Failure will not generate succeeded. There will be no new side effects after cancellation. If the artifact acceptance fails, you can go back to repair or clear failure. Validate these deterministic controls before using real models to evaluate task completion rates. Giving such counterexamples in interviews can better prove implementation capabilities than reciting "Plan-Action-Reflection".

Engineering deduction

scene
Hypothetical engineering scenario: After the R&D Agent repairs the interface, it repeatedly runs the same failed test, and the task never ends.
design decisions
Set up shared budgets and repeated action detection, and convert failure logs into structured diagnostics; success must be accepted through the specified interface.
Verify target
Expected behavior: Repeated failures trigger a restricted end, retaining modification and error evidence; only runs that pass acceptance will enter the successful final state.
applicable boundary
This is a design example and does not claim real team performance improvement metrics; budget thresholds need to be calibrated based on specific task distribution.

Continuous questions and answers

Continue reading along with the premises and constraints of the problem. Understand the reference answers first, then try to put away the answers and explain the cause and effect and trade-offs in your own words.

Draw inferences from one example: If the conditions change, how to deduce it?

First find out the conditions for change, and then determine which premises in the original plan still hold true. The following cases are teaching deductions to facilitate the transfer of principles to new problems.

Research assignment has no definite answer

Changing conditions:Acceptance changes from automatic testing to evidence coverage and manual judgment

Extended question:How to judge the completion of data research?

Derivation and reference solutions

Start by defining the scope of coverage, the types of sources that must be consulted, how references can be read back, and how open questions are presented. References and required fields can be automatically checked during runtime, and disputed conclusions are submitted to manual or specialized review. Reaching the scope and truthfully reporting unknowns counts as completion. The model cannot be required to guarantee that nothing is missing in the world, nor can the number of searches be used as a substitute for evidence coverage.

The principles that remain unchanged:Completion always relies on predefined observable conditions, it's just that the validator changes from testing to evidence review.

Scheduled tasks run repeatedly

Changing conditions:One task becomes multiple runs per day

Extended question:After today's failure, can yesterday's success status be automatically used tomorrow?

Derivation and reference solutions

Each run binds a snapshot of the input with proof of acceptance. The facts published yesterday can be reused, today's data refresh, permissions and budgets are rechecked; the same business actions continue to use the same operation keys, and new date artifacts use new business identifiers. One success cannot overwrite a failure under new input, and the system should display the results of each run separately.

The principles that remain unchanged:Task lifecycle and business operation lifecycle must be separated.

Easy to make mistakes

  • Treat max_turns as the total cost or total number of tools limit
  • Directly map the final text of the model to business success
  • Tools that retry unconditionally after timeout have side effects

References

It is designed based on public technical information; the reference materials support the technical mechanism, and the scenarios and scoring standards are designed by this website and do not represent the original interview questions of a certain company. New Q&A and migration cases are added for principle explanation, and source verification and case operation verification are recorded separately.

Check how far you understand

After reading, you can explain the principles, boundaries, and trade-offs against these standards. It is up to you to evaluate your mastery; if further verification is needed, complete the small tasks below.

Basic standards met
It can distinguish between model proposal completion and business acceptance, and provide bounded stopping conditions.
Intermediate and advanced signals
Implement rounds, tool counts, fees, cancellations, and error classifications into runtime.
Senior Signal
Can explain the unknown state of shared budget reservation and external actions, and verify it with counterexamples.

Hands-on verificationComplete on demand · Suggestions15 minutes

Draw a state machine for a tool loop and simulate the model requiring the same failure action three times in a row.

Expand acceptance requirements and checkpoints
  • Repeated failures will result in bounded termination.
  • The model is claimed to be complete but still needs to pass acceptance
  • Stop reason traceable

Key inspections

  • Can explain the difference between model call rounds and tool execution times
  • Distinguish final text from business acceptance
  • Provide budget, cancellation, exception final status and recording method