Agent Application DevelopmentAccount
Knowledge catalogChoose core direction and segmented content
knowledge unit 61FundamentalsConceptsAbout 12 minutes

Understand → Implement → Debug → Design

Token accounting, context capacity, and generation headroom

Use a concrete budget example to distinguish input capacity, generation limits, and next-turn headroom, then decide when to split, filter, or compress.

Tokencontext windowOutput budgetTruncateevidence use

Knowledge content check2026-10-03 · Check the source of the original question2026-10-03

Which step do you want to learn from this knowledge point?

Select the starting point based on the current basis, or you can go deeper one by one. When you encounter an unfamiliar concept, go back to the core principles first; use the knowledge exercises to check your understanding when you are finished.

Understand first

New to this knowledge point

Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.

Start with core principles →

Implement next

Prepare to write the principles into code

Understand implementation steps and boundaries, complete small tasks, and check results against acceptance requirements.

Reading implementation and trade-offs →

Debug failures

Need to handle failures and changes in conditions

Follow the continuous questioning to locate the failure premise, and then compare the migration cases to explain how the plan should be adjusted.

Continue to delve deeper into the problem →

Compare designs

Need to design or review plans

Combine engineering deductions and senior self-evaluation standards to explain the applicable conditions, costs and alternatives of the plan.

Analyze engineering scenarios →
Knowledge unit directory

LEARN · PRACTICE · REFLECT

Knowledge learning and personal records

My notes and review ↗

First read along the principles, Q&A and migration cases. When you need to check your understanding, switch to reinforcement exercises or start personal recording.

Answers and personal notes

Each modified commit will be kept as an independent history. Your level of mastery is up to you to evaluate yourself against the standards.

Core concept · Token accounting, context capacity, and generation headroom

Understand the core principles first

Preparatory concepts:Request and response life cycle, Result type of model interface

The context window limits how much information one request can hold; the generation limit caps how much that request may generate. Neither proves that the information is used correctly. Count the complete request according to the target interface. Newly added results consume input in the next turn, so the previous turn's remaining headroom does not guarantee that later requests fit.

Capacity differs from comprehension

A file fitting in memory does not establish correct parsing. Similarly, an accepted model request satisfies capacity rules without demonstrating that the model identified interacting conditions or cited the right number. Original research on positional sensitivity establishes observations for its tested models and tasks, not a universal degradation rate for newer models.

Output reservation is a policy

An output cap sets a maximum; generation may finish early or exhaust the budget. Where reasoning counts toward generated usage, even short visible text can exhaust G. Choose reservations from representative usage distributions, permitted cost, and required delivery length. Monitor limit-trigger rates rather than reserving the same amount for every task.

Recalculate each round

In this exercise, the first round has room for another 2,000 tokens, but a 4,000-token tool result requires changing the next request. One valid request does not establish that the whole agent run fits. Tools can return summaries and citations before selected passages are read. If exact wording determines the conclusion, retrieve original text in stages or choose sufficient capacity rather than dropping conditions. This unit addresses request capacity; the context compaction unit covers long-task ledgers, authorization, and summary drift.

Check understanding with a question

What are the constraints on token, context window and output upper limit respectively? How do you leave room for the next round?

Tokens measure model input and usage, not fixed Chinese-character counts. Inputs include instructions, tools, history, and evidence; context and output limits differ. Reasoning may consume generation budget without visible text. Count under the target API, reserve generation and safety margins, and recalculate after tool results. Handle input overflow separately from output truncation. Larger windows do not guarantee correct evidence use.

Implementation and trade-offs

First distinguish measurement and capacity

Token is the unit for model encoding and processing content, and is not equal to a Chinese character or English word. The same text may have different amounts in different encodings, and pure text will miss role boundaries, tool schemas, and message structures. When developing, select the counting method of the target model, count complete requests as much as possible, and then use the actual usage returned by the interface to compare. Token pricing, context capacity and request rate limiting are also different constraints. Cheap price does not mean a larger window.

What restrictions do input and output share?

The complete input is denoted by I, the planned total generation budget is denoted by G, and the model context capacity is denoted by C. For interfaces that use input and generation to share a context, first check that I+G does not exceed C, and then check that G does not exceed the independent maximum output limit of the model; the specific counting and reservation rules are subject to the deployment interface. If the amount of reasoning is included in the generation budget, G must also accommodate reasoning and visible answers, and cannot be set only according to the word count of the final article. Retaining safety margin S is an engineering choice to cover input-estimation error or expected growth and is not an automatically given space by the supplier.

A budget calculation example with given assumptions

All numbers below are teaching assumptions and do not represent any current model: C = 32000, independent maximum output = 8000, current full input I = 23000, planned G = 6000, safety margin S = 1000. Using the local rule of I+G+S≤C, 2000 input tokens can currently be added. If the 4000 token tool result is added in the next round and the input becomes 27000, the same generated budget will require 34000, and the total will exceed the limit by 2000. At this time, you should filter out irrelevant input, read in steps, or reduce the generation range when delivery allows; just increasing the output parameters will make it more crowded.

Handle different failures separately

Input that is too large may be rejected by the interface, or may trigger explicitly configured cropping. Automatic cropping cannot be performed by default while still retaining key materials. Check the completion status and reasons when the generation reaches the upper limit. Partial JSON must not be used as executable parameters. Network flow interruption and token budget exhaustion are not the same problem and need to be diagnosed according to the actual status. Plain text can be generated in chapters and accepted paragraph by paragraph. Tool calls that rely on complete parameters must first produce valid, complete objects. Compression, retrieval and window enlargement each have trade-offs: abstracts are lossy, retrieval may miss recall, long input is costly and evidence utilization is not guaranteed, and after selection, it is necessary to check whether the answers and references are supported.

Engineering deduction

scene
Teaching scenario: The input is 23,000 tokens, and the tool will add 4,000 in the next round. Assume that the window is 32,000, the generation budget is 6,000, and the safety margin is 1,000.
design decisions
Calculate the next request before adding and find that the demand is 34,000. Choose to remove at least 2,000 irrelevant tokens or read them step by step according to the problem; the key basis is to retain the original text and version.
Verify target
The arithmetic is checkable: initial headroom is 2,000 tokens, and the shortfall after adding results is 2,000. Acceptance requires a valid budget, retained coverage of necessary evidence, and complete output. No real model was called.
applicable boundary
Counting and window sharing rules are given assumptions; safety margins cannot prove semantic completeness, and conditions may still be missed after segmentation or summarization.

Continuous questions and answers

Continue reading along with the premises and constraints of the problem. Understand the reference answers first, then try to put away the answers and explain the cause and effect and trade-offs in your own words.

Draw inferences from one example: If the conditions change, how to deduce it?

First find out the conditions for change, and then determine which premises in the original plan still hold true. The following cases are teaching deductions to facilitate the transfer of principles to new problems.

Full contract terms determine exceptions

Changing conditions:Ordinary summarization tasks become dependent on precise explanations of definitions, negatives, and exceptions.

Extended question:When the budget is insufficient, can you summarize all the terms and explain them?

Derivation and reference solutions

The abstract can locate relevant sections, but read the full definitions, terms, and exceptions of the applicable version before key conclusions, and retain the citation. If the combination still exceeds the window, split it into acceptable sub-tasks according to the problem, save the intermediate claims and basis and then synthesize them; if the necessary context cannot be retained, narrow the scope of the conclusion. Changing tokens cannot offset errors caused by missing conditions.

The principles that remain unchanged:Capacity adjustments must retain evidence that determines the conclusion, and independent acceptance of semantic adequacy is required.

Fixed format report output is very long

Changing conditions:The input is short, but requires many chapters to be delivered at once and the output cap is easily reached.

Extended question:Is changing to a larger input window the first choice?

Derivation and reference solutions

First check the independent output upper limit and actual generation usage. Large input space does not mean that longer output is allowed. Generate it in chapters according to the stable directory, save and verify it section by section, and then check the overall references and duplicate omissions; if the complete object must be returned at once, reduce the scope of delivery or choose an interface with sufficient output capabilities. Don't splice guess fields at the end of truncated JSON.

The principles that remain unchanged:Input capacity and output capacity constrain delivery separately, with complete results coming from explicit acceptance.

Easy to make mistakes

  • Apply rough conversion ratios from English to Chinese, code or multi-modal complete requests.
  • It is believed that setting max_output_tokens will generate the specified length, or that this parameter only counts visible text.
  • Inputs fill the window before responses and next round of tool results are considered.
  • Think of increasing the window as a guarantee of search quality, citation support, and factual correctness.
  • Based on the truncated JSON, add a few parentheses and then perform an external write operation.

References

Based on official interface documents and original paper design; budget figures are clear teaching assumptions, real models have not been run, and do not represent a company's original topic or current product quota. New Q&A and migration cases are added for principle explanation, and source verification and case operation verification are recorded separately.

Check how far you understand

After reading, you can explain the principles, boundaries, and trade-offs against these standards. It is up to you to evaluate your mastery; if further verification is needed, complete the small tasks below.

Basic standards met
Make it clear that token is not the number of words, and distinguish between complete input, context capacity and maximum output.
Intermediate and advanced signals
Correctly calculate a given budget study, know that inference may account for the generated budget, and diagnose over-windowing and truncation separately.
Senior criteria
Set capacity rules for the next round of increments, weigh segmentation, retrieval vs. summarization, and verify that necessary evidence is used instead of just verifying that the request is accepted.

Hands-on verificationComplete on demand · Suggestions20 minutes

The teaching assumptions for this question are used: window 32000, independent maximum output 8000, complete input 23000, generation budget 6000, and safety margin 1000. First calculate the amount of new inputs, then add the 4000 token tool results, give two legal adjustment options, and explain how to check the necessary evidence and complete output for each option.

Expand acceptance requirements and checkpoints
  • Calculate 2,000 tokens of initial input headroom, 34,000 tokens of total demand after adding results, and a 2,000-token shortfall.
  • At least give two options of reducing input while retaining the generation budget, and shrinking the generation scope when the task allows. Do not treat increasing output as a universal fix.
  • The illustrated numbers are teaching assumptions, and real requests are counted by target interface.
  • The parameters are not executed when the output status is incomplete; the necessary basis is verified one by one after adding or deleting evidence.

Key inspections

  • Can distinguish between token, complete input, context window, maximum output and cost budget.
  • Input and generation margins are calculated using the given assumptions to account for the impact of the next round of additional tool results.
  • Able to distinguish between inputs that are too large, generation budget exhaustion, and network outages, and do not execute incomplete parameters.
  • Knowing the long window capacity and evidence usage effect need to be verified separately.