Review the prerequisites
Suitable for: We have understood the premise that input and generation jointly occupy capacity.
- Token
- The unit of model encoding and measurement content is not equal to a fixed number of Chinese characters or words; the experiment directly gives the count and does not implement tokenization.
- context
- The information available to the model in this request, including instructions, history, tool definitions, and source material.
- Generate budget
- The upper limit of content that is allowed to be generated; some interfaces also include inference usage, please check the target interface for specific rules.
- safety margin
- Headroom actively set aside by the application for counting errors or projected increments; not a substitute for true counts.
How does the mechanism work?
- Assembly request
Count instructions, history, tool definitions, and evidence respectively.
- Check constraints
Also check the total capacity and independent output upper limit.
- Call and check status
Distinguish between completion, budget exhaustion, rejection, and transmission interruption.
- Prepare for the next round
Recount after adding tool results to check whether necessary evidence is retained.
Implementation · Build it
Turn capacity rules into a testable function
Objectives of this level: Ability to run budget experiments and include the next round of increments in the review.
Count components before summing them
The offline lab's components count instructions, history, tool definitions, and evidence separately. remaining validates count types and the independent output limit before calculating capacity. It rejects booleans because Python bool values can behave as integers, while the contract requires explicit counts.
Check again on every round
A request that fits on round one may overflow on round two. Tool results, model responses, and new questions all add input, and real interfaces can charge for message structure too. This lab accepts given numbers. A production adapter needs the target model's counting method and actual usage records to calibrate estimates.
Passing the budget check is only one condition
The budget function establishes that local capacity rules permit the request. It checks no relevance, citation completeness, or tool authorization. A partial JSON object left after generation reaches its limit cannot be executed because an earlier capacity check passed. Complete-response status, parsing, and business checks remain necessary.
Run experiments and observe counterexamples
Offline arithmetic experiment for a given number of tokens; no model calls, no verification of tokenizers, streaming, or actual vendor limitations.
Python 3.10+ · Runs by default using only the standard library · Runs on your computer
- Calculate by hand first, then run the script to verify
- Change tool result increment and recalculate
- Deliberately exceed the independent output limit, check the wrong branch
python3 context_budget.pyView the entry-point script
"""Given token counts: no tokenizer, model request or provider-specific limit."""
import json
def remaining(context, output_limit, output, reserve, components):
values = [context, output_limit, output, reserve, *components.values()]
if any(type(v) is not int or v < 0 for v in values):
raise ValueError("counts must be non-negative integers")
if output > output_limit:
raise ValueError("output limit exceeded")
return context - sum(components.values()) - output - reserve
def demo():
components = dict(instructions=2000, history=10000, tools=3000, evidence=8000)
before = remaining(32000, 8000, 6000, 1000, components)
after = remaining(32000, 8000, 6000, 1000, {**components, "result": 4000})
assert before == 2000 and after == -2000
return dict(input_tokens=sum(components.values()), current_room=before,
next_turn_room=after, next_turn_fits=after >= 0)
if __name__ == "__main__":
print(json.dumps(demo(), sort_keys=True))
Expected output when running locally
{"current_room": 2000, "input_tokens": 23000, "next_turn_fits": false, "next_turn_room": -2000}- Complete input items are consistent
- Next round of capacity recalculation
- Passing the experiment does not mean passing the answer quality
Acceptance task for this level
Download context_budget.py, and after running it, change result to 1000, and then change output to 9000.
Check each item after completion
- Original experimental output current_room=2000, next_turn_room=−2000
- After the result increment is changed to 1000, the next round margin is 1000
- Standalone output cap error triggered when generation budget is changed to 9000
Save your own processes, code and results. Acceptance requirements are provided here, and course mastery status will not be automatically graded or saved at this time.
Hide the answer and check your understanding
The function returns a positive number, does it prove that the answer must be correct?
Expand reference derivation
It can only be proven that a given number complies with the capacity rules. Evidence quality, generation status, factual support and business results must also be verified separately.
Further explanations and practice
When encountering unfamiliar principles, first read the implementation, continuous questioning and migration cases, and then independently explain the premise and boundaries. Answers and notes are saved to the original account record.
Sources and verification scope
The principles are based on public information; the numbers, cases and tasks are the teaching design of this website. Offline experiments verify the range noted on this page, and the learning effect still needs to be judged through independent tasks and feedback.