Review the prerequisites
Suitable for: Information flow and budgeting strategies for long tasks need to be designed.
- Token
- The unit of model encoding and measurement content is not equal to a fixed number of Chinese characters or words; the experiment directly gives the count and does not implement tokenization.
- context
- The information available to the model in this request, including instructions, history, tool definitions, and source material.
- Generate budget
- The upper limit of content that is allowed to be generated; some interfaces also include inference usage, please check the target interface for specific rules.
- safety margin
- Headroom actively set aside by the application for counting errors or projected increments; not a substitute for true counts.
How does the mechanism work?
- Assembly request
Count instructions, history, tool definitions, and evidence respectively.
- Check constraints
Also check the total capacity and independent output upper limit.
- Call and check status
Distinguish between completion, budget exhaustion, rejection, and transmission interruption.
- Prepare for the next round
Recount after adding tool results to check whether necessary evidence is retained.
Design · Explain the trade-offs
Choose between retrieval, chunking, and summarization
Objectives of this level: Able to select information delivery methods based on task objectives and clarify quality and cost acceptance.
Start from the task's information needs
Contract review needs item-level evidence; code changes need relevant files and constraints; conversation summaries need decisions and unresolved questions. These structures differ. Dropping the oldest messages may remove restrictions that still apply. Manage task constraints, evidence, and recomputable material separately to decide what can be compressed.
Record the cost of each choice
On-demand retrieval reduces input but risks missing evidence. Staged generation reduces output pressure but requires consistency across sections. Summaries reduce history through a lossy transformation. Larger windows admit more material without guaranteeing that the model uses it correctly. Compare evidence coverage, result quality, latency, and usage together.
Design for the next round
Reserve space for likely tool results and subsequent generation. Read smaller batches when capacity is tight. Store large artifacts behind stable references and keep key constraints in a structure that can be read again. Derive reserves from task distributions and verified interface accounting rather than a universal percentage.
Run experiments and observe counterexamples
Offline arithmetic experiment for a given number of tokens; no model calls, no verification of tokenizers, streaming, or actual vendor limitations.
Python 3.10+ · Runs by default using only the standard library · Runs on your computer
- Calculate by hand first, then run the script to verify
- Change tool result increment and recalculate
- Deliberately exceed the independent output limit, check the wrong branch
python3 context_budget.pyView the entry-point script
"""Given token counts: no tokenizer, model request or provider-specific limit."""
import json
def remaining(context, output_limit, output, reserve, components):
values = [context, output_limit, output, reserve, *components.values()]
if any(type(v) is not int or v < 0 for v in values):
raise ValueError("counts must be non-negative integers")
if output > output_limit:
raise ValueError("output limit exceeded")
return context - sum(components.values()) - output - reserve
def demo():
components = dict(instructions=2000, history=10000, tools=3000, evidence=8000)
before = remaining(32000, 8000, 6000, 1000, components)
after = remaining(32000, 8000, 6000, 1000, {**components, "result": 4000})
assert before == 2000 and after == -2000
return dict(input_tokens=sum(components.values()), current_room=before,
next_turn_room=after, next_turn_fits=after >= 0)
if __name__ == "__main__":
print(json.dumps(demo(), sort_keys=True))
Expected output when running locally
{"current_room": 2000, "input_tokens": 23000, "next_turn_fits": false, "next_turn_room": -2000}- Complete input items are consistent
- Next round of capacity recalculation
- Passing the experiment does not mean passing the answer quality
Acceptance task for this level
Design a capacity policy for "Read a batch of contracts and generate a risk report" that explains when to segment, retrieve, or digest.
Check each item after completion
- List the terms and task constraints that must not be lost
- Set explicit capacity checks for next round of tool results
- Compare quality, usage, and latency and design regression examples
Save your own processes, code and results. Acceptance requirements are provided here, and course mastery status will not be automatically graded or saved at this time.
Hide the answer and check your understanding
By enlarging the window tenfold, would it be possible to remove the evidence coverage check?
Expand reference derivation
No. The capacity to accommodate information and the correct use of evidence are different abilities. The expansion of the window still requires verification of whether the necessary evidence and conditions are adopted.
Further explanations and practice
When encountering unfamiliar principles, first read the implementation, continuous questioning and migration cases, and then independently explain the premise and boundaries. Answers and notes are saved to the original account record.
Sources and verification scope
The principles are based on public information; the numbers, cases and tasks are the teaching design of this website. Offline experiments verify the range noted on this page, and the learning effect still needs to be judged through independent tasks and feedback.