Review the prerequisites
Suitable for: A long conversation needs to be located or a tool loop fails.
- Token
- The unit of model encoding and measurement content is not equal to a fixed number of Chinese characters or words; the experiment directly gives the count and does not implement tokenization.
- context
- The information available to the model in this request, including instructions, history, tool definitions, and source material.
- Generate budget
- The upper limit of content that is allowed to be generated; some interfaces also include inference usage, please check the target interface for specific rules.
- safety margin
- Headroom actively set aside by the application for counting errors or projected increments; not a substitute for true counts.
How does the mechanism work?
- Assembly request
Count instructions, history, tool definitions, and evidence respectively.
- Check constraints
Also check the total capacity and independent output upper limit.
- Call and check status
Distinguish between completion, budget exhaustion, rejection, and transmission interruption.
- Prepare for the next round
Recount after adding tool results to check whether necessary evidence is retained.
Debugging · Diagnose failures
Use response states to diagnose overflow, truncation, and lost evidence
Objectives of this level: Failure stages can be located based on request, response and evidence snapshots.
Preserve facts that distinguish causes
Record input counts by component, generation settings, terminal response status, tool-result size, and evidence IDs. Debug logs must respect the evidence's access rules; troubleshooting does not justify retaining all user text. A single request-failed message cannot distinguish capacity rejection, network failure, and incomplete generation.
Construct three separate failures
Keep the generation budget fixed and enlarge tool results until the input rule fails. Next, keep input within capacity but lower the output allowance and inspect the incomplete artifact. Finally, interrupt transport. The numeric lab covers no transport interruption, so verify that case with a real interface or a reliable test double. Increasing tokens is not a common remedy for all three.
Check what compression removed
A summary can shorten evidence while dropping exception clauses. Label the evidence each fixed task must retain, then compare coverage and answers before and after compression. A request that now fits demonstrates a capacity improvement. If it silently omits necessary conditions, the repair is incomplete.
Run experiments and observe counterexamples
Offline arithmetic experiment for a given number of tokens; no model calls, no verification of tokenizers, streaming, or actual vendor limitations.
Python 3.10+ · Runs by default using only the standard library · Runs on your computer
- Calculate by hand first, then run the script to verify
- Change tool result increment and recalculate
- Deliberately exceed the independent output limit, check the wrong branch
python3 context_budget.pyView the entry-point script
"""Given token counts: no tokenizer, model request or provider-specific limit."""
import json
def remaining(context, output_limit, output, reserve, components):
values = [context, output_limit, output, reserve, *components.values()]
if any(type(v) is not int or v < 0 for v in values):
raise ValueError("counts must be non-negative integers")
if output > output_limit:
raise ValueError("output limit exceeded")
return context - sum(components.values()) - output - reserve
def demo():
components = dict(instructions=2000, history=10000, tools=3000, evidence=8000)
before = remaining(32000, 8000, 6000, 1000, components)
after = remaining(32000, 8000, 6000, 1000, {**components, "result": 4000})
assert before == 2000 and after == -2000
return dict(input_tokens=sum(components.values()), current_room=before,
next_turn_room=after, next_turn_fits=after >= 0)
if __name__ == "__main__":
print(json.dumps(demo(), sort_keys=True))
Expected output when running locally
{"current_room": 2000, "input_tokens": 23000, "next_turn_fits": false, "next_turn_room": -2000}- Complete input items are consistent
- Next round of capacity recalculation
- Passing the experiment does not mean passing the answer quality
Acceptance task for this level
Write a checklist to list observation signals and processing methods for capacity rejection, generation truncation, and streaming interruption.
Check each item after completion
- Capacity error associates full input with generated budget
- For truncation, check the terminal response status and artifact completeness
- Review the necessary evidence after compression rather than just seeing whether the request was successful.
Save your own processes, code and results. Acceptance requirements are provided here, and course mastery status will not be automatically graded or saved at this time.
Hide the answer and check your understanding
The same request can run after being compressed. Can you directly announce that the problem has been fixed?
Expand reference derivation
No. Also check that evidence of mission constraints and exceptions is preserved, and that the final conclusion is still supported by the material.
Further explanations and practice
When encountering unfamiliar principles, first read the implementation, continuous questioning and migration cases, and then independently explain the premise and boundaries. Answers and notes are saved to the original account record.
Sources and verification scope
The principles are based on public information; the numbers, cases and tasks are the teaching design of this website. Offline experiments verify the range noted on this page, and the learning effect still needs to be judged through independent tasks and feedback.