Agent Application DevelopmentAccount
Knowledge catalogChoose core direction and segmented content

SYSTEMATIC LEARNING / FOUR-LEVEL COURSE

Context and generation budgets

Calculate the capacity of a complete request and connect input size, output budget, growth in the next turn, and evidence retention.

Learning objectives: Calculate capacity for a request and the next turn of tool results, and distinguish context overflow, output truncation, and network interruption.

Content checked: 2026-10-04 · Each level has independent explanations, tasks and inspections

Choose a starting point based on your familiarity with this topic. Current level: Design · Explain the trade-offs. After completing the task, continue to the next level. Reading and self-checks alone do not establish mastery.

On this level

Review the prerequisites

Suitable for: Information flow and budgeting strategies for long tasks need to be designed.

Token
The unit of model encoding and measurement content is not equal to a fixed number of Chinese characters or words; the experiment directly gives the count and does not implement tokenization.
context
The information available to the model in this request, including instructions, history, tool definitions, and source material.
Generate budget
The upper limit of content that is allowed to be generated; some interfaces also include inference usage, please check the target interface for specific rules.
safety margin
Headroom actively set aside by the application for counting errors or projected increments; not a substitute for true counts.

How does the mechanism work?

  1. Assembly request

    Count instructions, history, tool definitions, and evidence respectively.

  2. Check constraints

    Also check the total capacity and independent output upper limit.

  3. Call and check status

    Distinguish between completion, budget exhaustion, rejection, and transmission interruption.

  4. Prepare for the next round

    Recount after adding tool results to check whether necessary evidence is retained.

Design · Explain the trade-offs

Choose between retrieval, chunking, and summarization

Objectives of this level: Able to select information delivery methods based on task objectives and clarify quality and cost acceptance.

Start from the task's information needs

Contract review needs item-level evidence; code changes need relevant files and constraints; conversation summaries need decisions and unresolved questions. These structures differ. Dropping the oldest messages may remove restrictions that still apply. Manage task constraints, evidence, and recomputable material separately to decide what can be compressed.

Record the cost of each choice

On-demand retrieval reduces input but risks missing evidence. Staged generation reduces output pressure but requires consistency across sections. Summaries reduce history through a lossy transformation. Larger windows admit more material without guaranteeing that the model uses it correctly. Compare evidence coverage, result quality, latency, and usage together.

Design for the next round

Reserve space for likely tool results and subsequent generation. Read smaller batches when capacity is tight. Store large artifacts behind stable references and keep key constraints in a structure that can be read again. Derive reserves from task distributions and verified interface accounting rather than a universal percentage.

Run experiments and observe counterexamples

Offline arithmetic experiment for a given number of tokens; no model calls, no verification of tokenizers, streaming, or actual vendor limitations.

Python 3.10+ · Runs by default using only the standard library · Runs on your computer

  1. Calculate by hand first, then run the script to verify
  2. Change tool result increment and recalculate
  3. Deliberately exceed the independent output limit, check the wrong branch
Downloadcontext_budget.py ↓
python3 context_budget.py
View the entry-point script
"""Given token counts: no tokenizer, model request or provider-specific limit."""
import json


def remaining(context, output_limit, output, reserve, components):
    values = [context, output_limit, output, reserve, *components.values()]
    if any(type(v) is not int or v < 0 for v in values):
        raise ValueError("counts must be non-negative integers")
    if output > output_limit:
        raise ValueError("output limit exceeded")
    return context - sum(components.values()) - output - reserve


def demo():
    components = dict(instructions=2000, history=10000, tools=3000, evidence=8000)
    before = remaining(32000, 8000, 6000, 1000, components)
    after = remaining(32000, 8000, 6000, 1000, {**components, "result": 4000})
    assert before == 2000 and after == -2000
    return dict(input_tokens=sum(components.values()), current_room=before,
                next_turn_room=after, next_turn_fits=after >= 0)


if __name__ == "__main__":
    print(json.dumps(demo(), sort_keys=True))

Expected output when running locally

{"current_room": 2000, "input_tokens": 23000, "next_turn_fits": false, "next_turn_room": -2000}
  • Complete input items are consistent
  • Next round of capacity recalculation
  • Passing the experiment does not mean passing the answer quality
View the running environment, output and verification records →

Acceptance task for this level

Design a capacity policy for "Read a batch of contracts and generate a risk report" that explains when to segment, retrieve, or digest.

Check each item after completion

  • List the terms and task constraints that must not be lost
  • Set explicit capacity checks for next round of tool results
  • Compare quality, usage, and latency and design regression examples

Save your own processes, code and results. Acceptance requirements are provided here, and course mastery status will not be automatically graded or saved at this time.

Hide the answer and check your understanding

By enlarging the window tenfold, would it be possible to remove the evidence coverage check?

Further explanations and practice

When encountering unfamiliar principles, first read the implementation, continuous questioning and migration cases, and then independently explain the premise and boundaries. Answers and notes are saved to the original account record.

All linked explanations and exercises (3 )

Sources and verification scope

The principles are based on public information; the numbers, cases and tasks are the teaching design of this website. Offline experiments verify the range noted on this page, and the learning effect still needs to be judged through independent tasks and feedback.