Review the prerequisites
Suitable for: Can write basic programs and handle model context for the first time.
- Token
- The unit of model encoding and measurement content is not equal to a fixed number of Chinese characters or words; the experiment directly gives the count and does not implement tokenization.
- context
- The information available to the model in this request, including instructions, history, tool definitions, and source material.
- Generate budget
- The upper limit of content that is allowed to be generated; some interfaces also include inference usage, please check the target interface for specific rules.
- safety margin
- Headroom actively set aside by the application for counting errors or projected increments; not a substitute for true counts.
How does the mechanism work?
- Assembly request
Count instructions, history, tool definitions, and evidence respectively.
- Check constraints
Also check the total capacity and independent output upper limit.
- Call and check status
Distinguish between completion, budget exhaustion, rejection, and transmission interruption.
- Prepare for the next round
Recount after adding tool results to check whether necessary evidence is retained.
Foundation · Understand the concepts
Why can a short document still exceed the request limit?
Objectives of this level: Can indicate which parts of the complete input are included and distinguish between insufficient capacity and insufficient generation.
Count the whole request
A program reading a contract sends more than the contract itself: instructions, conversation history, tool definitions, and retrieved evidence also occupy input space. A few pages of text do not establish that the complete request is small. Tokens are an accounting unit; 2,000 characters cannot be assumed to equal 2,000 tokens.
Distinguish context capacity from the output limit
Context capacity limits the information a request can accommodate. A separate output limit caps generation. For interfaces where input and generation share capacity, a permitted output setting may still overflow the context when combined with the input. Increasing the output setting does not expand the context window.
Work through an explicit assumption
The lab assumes total capacity of 32,000, input of 23,000, a generation budget of 6,000, and a reserve of 1,000. Its local rule leaves 2,000. Adding a 4,000-token tool result on the next round exceeds capacity by 2,000. This is a request-capacity problem. If input fits but generation stops at its output limit, that is a different failure. A network disconnection alone establishes neither.
Run experiments and observe counterexamples
Offline arithmetic experiment for a given number of tokens; no model calls, no verification of tokenizers, streaming, or actual vendor limitations.
Python 3.10+ · Runs by default using only the standard library · Runs on your computer
- Calculate by hand first, then run the script to verify
- Change tool result increment and recalculate
- Deliberately exceed the independent output limit, check the wrong branch
python3 context_budget.pyView the entry-point script
"""Given token counts: no tokenizer, model request or provider-specific limit."""
import json
def remaining(context, output_limit, output, reserve, components):
values = [context, output_limit, output, reserve, *components.values()]
if any(type(v) is not int or v < 0 for v in values):
raise ValueError("counts must be non-negative integers")
if output > output_limit:
raise ValueError("output limit exceeded")
return context - sum(components.values()) - output - reserve
def demo():
components = dict(instructions=2000, history=10000, tools=3000, evidence=8000)
before = remaining(32000, 8000, 6000, 1000, components)
after = remaining(32000, 8000, 6000, 1000, {**components, "result": 4000})
assert before == 2000 and after == -2000
return dict(input_tokens=sum(components.values()), current_room=before,
next_turn_room=after, next_turn_fits=after >= 0)
if __name__ == "__main__":
print(json.dumps(demo(), sort_keys=True))
Expected output when running locally
{"current_room": 2000, "input_tokens": 23000, "next_turn_fits": false, "next_turn_room": -2000}- Complete input items are consistent
- Next round of capacity recalculation
- Passing the experiment does not mean passing the answer quality
Acceptance task for this level
Calculate the current request and the margin after adding the tool results by hand, and write down three failure phenomena.
Check each item after completion
- Write 23000+6000+1000=30000
- Calculate the next round margin to be −2000
- Input overwindows, generation truncation, and network interruptions are explained separately.
Save your own processes, code and results. Acceptance requirements are provided here, and course mastery status will not be automatically graded or saved at this time.
Hide the answer and check your understanding
The upper limit of output is 8000. Why does this example allow 6000 to be generated but it cannot be loaded in the next round?
Expand reference derivation
The independent output cap is just a condition. The next round inputs 27,000 plus the generated 6,000 and the margin 1,000, for a total of 34,000, exceeding the teaching capacity of 32,000.
Further explanations and practice
When encountering unfamiliar principles, first read the implementation, continuous questioning and migration cases, and then independently explain the premise and boundaries. Answers and notes are saved to the original account record.
Sources and verification scope
The principles are based on public information; the numbers, cases and tasks are the teaching design of this website. Offline experiments verify the range noted on this page, and the learning effect still needs to be judged through independent tasks and feedback.