Agent Application DevelopmentAccount
Knowledge catalogChoose core direction and segmented content
knowledge unit 19IntermediateImplementationAbout 12 minutes

Understand → Implement → Debug → Design

Semantic completeness and evidence location in chunking

Examine document structure, search units, contextual completion, and reviewable chunking experiments.

ChunkingDocument parsingevidence

Knowledge content check2026-10-03 · Check the source of the original question2026-10-02

Which step do you want to learn from this knowledge point?

Select the starting point based on the current basis, or you can go deeper one by one. When you encounter an unfamiliar concept, go back to the core principles first; use the knowledge exercises to check your understanding when you are finished.

Understand first

New to this knowledge point

Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.

Start with core principles →

Implement next

Prepare to write the principles into code

Understand implementation steps and boundaries, complete small tasks, and check results against acceptance requirements.

Reading implementation and trade-offs →

Debug failures

Need to handle failures and changes in conditions

Follow the continuous questioning to locate the failure premise, and then compare the migration cases to explain how the plan should be adjusted.

Continue to delve deeper into the problem →

Compare designs

Need to design or review plans

Combine engineering deductions and senior self-evaluation standards to explain the applicable conditions, costs and alternatives of the plan.

Analyze engineering scenarios →
Knowledge unit directory

LEARN · PRACTICE · REFLECT

Knowledge learning and personal records

My notes and review ↗

First read along the principles, Q&A and migration cases. When you need to check your understanding, switch to reinforcement exercises or start personal recording.

Answers and personal notes

Each modified commit will be kept as an independent history. Your level of mastery is up to you to evaluate yourself against the standards.

Core concept · Semantic completeness and evidence location in chunking

Understand the core principles first

Preparatory concepts:Document parsing, Token budget, Source and permission metadata

Chunking trades retrieval granularity against evidence completeness. Each chunk must preserve the conditions and relationships needed to understand it. High vector similarity cannot restore missing relationships.

Text length does not define an evidence boundary

Fixed-length chunks offer a cheap baseline but can separate rules from exceptions or values from units. “500” alone may denote yuan, ten thousand yuan, milliseconds, or a version. Chunking should produce interpretable, citable evidence rather than evenly sized text.

Preserve clause hierarchy and definitions in contracts, headers and units in tables, and symbols, versions, dependencies, and call boundaries in code. Parsing errors precede chunking; overlap cannot restore a missing header. Retain pages, ranges, and versions for source verification.

Expand context under explicit conditions

Small chunks target local questions; parent or neighboring sections add context. Expansion from a public paragraph must not expose private attachments. Repeated expansion of one parent can also waste tokens. Check authorization and deduplicate expanded evidence.

Evaluate with boundary-sensitive questions

Compare strategies on main rules, exceptions, table relationships, and cross-function behavior. Measure candidate coverage, complete evidence, and citation accuracy. Overlap costs index space and duplicate candidates without guaranteeing preserved relationships. Choose for the task distribution, with separate handling for unusual structures.

Check understanding with a question

How should contracts, forms, and code be cut into pieces? What’s wrong with sticking to 500 words?

Choose chunks around questions and document structure. Preserve contract hierarchy and exceptions, table headers and units, and code symbols and dependencies. Retrieve small chunks and expand to parent or neighboring context within access and token budgets, controlling duplicates. Use fixed length as a baseline, then evaluate annotated evidence coverage, answer completeness, and citation positions.

Implementation and trade-offs

First define the problem and evidence unit

Asking "What are the liquidated damages?" requires the amount, triggering conditions and exceptions; asking about the trend of the table requires the header, unit and time columns. Splitting solely by character count can separate negations or units from their statements. Each block retains at least document_id, version, chapter path, page number or line range, and the numbers in the extracted text can be traced back to the source for verification. Parsing failures must be visible, and empty text cannot be stored as a success.

Split according to structure and then control the length

First use title, paragraph, table and function boundaries to establish candidate blocks, and then recursively split them when the budget is exceeded. When the table is divided into batches, the necessary headers are repeated and indicate whether there are continuation rows; the code retains function signatures, file paths and related type references. Overlap is only used to reduce boundary loss. Excessive overlap will squeeze multiple approximate blocks into the top few, reducing evidence diversity.

Small block retrieval and parent block expansion

Exact chunks are suitable for matching questions where the answer may require context and parent segments or adjacency chunks can be read after recall. Extensions must be made under the same document version and current permissions, and the budget calculated again. Just because the parent document is visible does not mean that all associated documents are visible by default. Answer citations should still be directed to specific passages that support the conclusion, not just a link to the entire manual.

Select parameters through experimentation

Prepare a set of questions including cross-segment exceptions, table units, and code call relationships, and mark the necessary evidence for each question. Hold retrieval and generation configurations fixed, change only the chunking strategy, and compare necessary evidence coverage, final answer correctness, repetition rate, and cost. Larger chunks may improve recall but reduce answer focus, so average similarity cannot be used to select the best solution.

Engineering deduction

scene
Interview Assumption: Contract answers only cite the main clause and leave out applicable exceptions in the next paragraph.
design decisions
Chapter relationships are retained, and adjacent exceptions are read after small block recall.
Verify target
The answer explains conditions and exceptions and points to the original text in the relevant version.
applicable boundary
Slice length is an experimental parameter and there is no fixed value that will fit all documents.

Continuous questions and answers

Continue reading along with the premises and constraints of the problem. Understand the reference answers first, then try to put away the answers and explain the cause and effect and trade-offs in your own words.

Draw inferences from one example: If the conditions change, how to deduce it?

First find out the conditions for change, and then determine which premises in the original plan still hold true. The following cases are teaching deductions to facilitate the transfer of principles to new problems.

A code base instead of a regular article

Changing conditions:The evidence involves functions and external calls

Extended question:The function exceeds the length budget, is it enough to cut it into several pieces?

Derivation and reference solutions

Preserve symbols, signatures, control flow fragments, and file versions first, then create navigable references to dependencies. When the user asks for timeout processing, read the relevant exception branches and called functions as needed; do not insert the entire repository for the sake of completeness of a single block. Code explanations are required to indicate which behavior depends on external configuration, and necessary evidence is retrieved individually across files.

The principles that remain unchanged:To preserve the semantic relationships required by the problem, text length cannot be substituted for behavioral boundaries.

Exceptions to the main terms and appendices of the contract

Changing conditions:Answer relies on non-adjacent paragraphs

Extended question:Can only expanding adjacent blocks avoid missing disclaimers?

Derivation and reference solutions

No. Build links by clause reference and definition relationships, and when recalling a main clause check the appendices, exceptions and applicable versions it references. Leave evidence gaps for appendices that are necessary but do not have the right to read, and refuse to give a complete conclusion. Such non-adjacent evidence problems are added to the annotation set to verify the expansion strategy.

The principles that remain unchanged:Contextual integrity is determined by dependencies, physical proximity is just one clue.

Easy to make mistakes

  • Same character length for all documents
  • Only text without location and version
  • Replacing final correctness with similarity

References

It is designed based on public technical information; the reference materials support the technical mechanism, and the scenarios and scoring standards are designed by this website and do not represent the original interview questions of a certain company. New Q&A and migration cases are added for principle explanation, and source verification and case operation verification are recorded separately.

Check how far you understand

After reading, you can explain the principles, boundaries, and trade-offs against these standards. It is up to you to evaluate your mastery; if further verification is needed, complete the small tasks below.

Basic standards met
Can point out segmentation differences in contracts, forms, and code.
Intermediate and advanced signals
Propose parent-child blocks, deduplication and necessary evidence annotation.
Senior criteria
Ability to isolate variable evaluation and handle cross-version, permissions and parsing failures.

Continue to do advanced research experiments

Why can’t the old context continue to be used after the policy changes?

Transferring pre-retrieval filtered ideas to memory versions and recovery. Observe how existing drafts become invalid after cancellation.

Read full text and fault analysis → · Download Reliability Experiment v3 ↓

python3 cli.py memory-put --db memory.sqlite
python3 cli.py submit --db memory.sqlite
python3 cli.py run --db memory.sqlite --lease-seconds 2 --fault after_draft
python3 cli.py memory-forget --db memory.sqlite
# 等待至少 2 秒后分别执行
python3 cli.py run --db memory.sqlite
python3 cli.py inspect --db memory.sqlite

Keep evidence and check item by item

  • The draft checkpoint already exists after the first exit.
  • Recovery after memory-forget gets failed with memory_changed_or_expired.
  • No publish checkpoint; explain the difference between fail blocking and complete deletion.

Verify local scope, version, and undo blocking; old checkpoints remain, no complete deletion of logs, backups, or checkpoints is provided.

Hands-on verificationComplete on demand · Suggestions15 minutes

Design block structure and reference fields for a document containing main clauses, exceptions, and tables.

Expand acceptance requirements and checkpoints
  • Exceptions are not disconnected from the main clause
  • Table units reserved
  • Each block can locate the original text

Key inspections

  • Preserve semantic boundaries according to document structure
  • Evidence location includes version and scope
  • Cover assessment slices with necessary evidence