Understand first
New to this knowledge point
Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.
Start with core principles →Understand → Implement → Debug → Design
Establish staged evidence and indicators to avoid all errors being solved by changing models.
Knowledge content check2026-10-03 · Check the source of the original question2026-10-02
It is recommended to understand first:
Prompts, structured output, and iterative acceptance checks →Context and generation budgets →Tool calls: structure, authorization, and business contracts →Select the starting point based on the current basis, or you can go deeper one by one. When you encounter an unfamiliar concept, go back to the core principles first; use the knowledge exercises to check your understanding when you are finished.
New to this knowledge point
Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.
Start with core principles →Prepare to write the principles into code
Understand implementation steps and boundaries, complete small tasks, and check results against acceptance requirements.
View the code example →Need to handle failures and changes in conditions
Follow the continuous questioning to locate the failure premise, and then compare the migration cases to explain how the plan should be adjusted.
Continue to delve deeper into the problem →Need to design or review plans
Combine engineering deductions and senior self-evaluation standards to explain the applicable conditions, costs and alternatives of the plan.
Analyze engineering scenarios →LEARN · PRACTICE · REFLECT
First read along the principles, Q&A and migration cases. When you need to check your understanding, switch to reinforcement exercises or start personal recording.
Can be practiced directly. After logging in, answers, favorites, and notes will be saved to your account.
Log in and saveEach modified commit will be kept as an independent history. Your level of mastery is up to you to evaluate yourself against the standards.
Core concept · Evidence flow and stage-by-stage RAG diagnosis
Preparatory concepts:information retrieval ranking, Pipeline intermediate results, Labeled datasets and controlled comparisons
The evidence that answers the dependency must first exist, be recalled, retained, and fed into the model in its entirety. Find the earliest loss point along the evidence flow to know which stage of changes can affect the error.
An upstream field can disappear during parsing, filtering, ranking, or serialization. Likewise, an answer existing in a document, candidate set, context, and final response are different facts. Diagnose from the source rather than assuming every error is poor retrieval.
Annotate minimum evidence, such as an error code, trigger condition, and remedy. Fix index, authorization, and query versions; record each stage’s inputs, outputs, and exclusion reasons by chunk ID. Evidence ranked eighteenth may disappear when only five chunks are assembled. A wrong answer may require better ranking, assembly, or combining evidence.
Recall@k measures candidate coverage; ranking metrics measure position; context inspection checks complete text and conditions; claim evaluation checks support. Vector and RRF scores are ranking signals, not factual probabilities. Final accuracy can conceal compensating stage errors; recall alone can reward excessive noise.
Keep the generator fixed and replace candidates with curated evidence. Remaining errors point toward assembly or generation. Supplying complete evidence directly helps isolate instructions, reasoning, and verification. A stronger model answering from learned knowledge while citing incorrectly does not establish repaired retrieval. Include unanswerable, permission-restricted, and version-specific questions so increased answer volume cannot masquerade as reliability.
Fix the question, index version, and access scope. Save rewritten queries, candidates, rankings, actual model context, and citations. Missing candidates indicate data or retrieval problems; dropped candidates indicate ranking problems; truncated selected evidence indicates assembly problems. Wrong answers with complete evidence point toward generation. Evaluate annotated questions by stage rather than fluency alone.
The original question, rewritten question, user permissions, knowledge version, chunking strategy and search parameters are retained. First confirm whether the original data contains answers and whether tables, titles or key conditions are missing in the database analysis. When a document does not exist or is out of date, no matter how powerful the model is, it can only guess. Document_id, chunk_id, version, retrieval channel, and ranking are logged for each candidate; these are used for diagnostics, but the logs themselves must also adhere to permissions and anonymization requirements. Without a snapshot of the actual context, looking only at the final answer cannot reliably locate the error.
Label the test questions with pieces of evidence that support the answer. The candidate stage checks Recall@k; the sorting stage looks at the location of the correct evidence, using MRR or nDCG; the context stage checks whether key paragraphs are complete, whether conflicting versions have been mixed, and whether the reference identification still corresponds to the original text; the generation stage checks whether each conclusion is supported by evidence and whether applicable conditions are missing. Just because the evidence is in the candidate list does not mean that the model has seen it, and the fact that the evidence is in does not mean that the answer actually refers to it. Rejected questions should be evaluated separately to avoid counting only samples with answers.
Terminology, error codes, and precise identification often require keyword retrieval, and expression changes can be supplemented by vector recall. Hybrid retrieval requires merging different rankings and cannot directly add similarities and BM25 scores with different dimensions. Azure AI Search uses RRF to fuse by rank, semantic reranking is a subsequent stage, and the scores have other meanings. The following code only demonstrates the fusion mechanism. If the recall is insufficient, then check the rewriting, chunking, filtering and number of candidates; if the sorting is not good, adjust the reranking; if the context is insufficient, reduce the repetition and retain the conditional paragraphs; only adjust the instructions and output verification when it is unfaithful.
Pinned a batch of questions covering precise terminology, vague expressions, cross-document reasoning, no answers, and permission restrictions. Only change one type of configuration at a time, and compare stage indicators, final evidence support rate, latency and cost. Manually review failed samples and mark root causes. You cannot use a score from the model's self-evaluation to replace all factual judgments. Data, index, reranker and prompt versions are recorded together for playback. Getting a certain problem better does not mean that the whole problem is getting better. In particular, expanding top-k may bring more noise and context costs.
Python 3 standard library demo. The inputs are two hypothetical ranked lists, without real retrieval, vector model or semantic reranking; rank_constant is a different parameter than retrieval top-k.
def rrf(rankings, rank_constant=60):
if rank_constant <= 0:
raise ValueError("rank_constant must be positive")
scores = {}
for ranking in rankings:
seen = set()
for rank, doc_id in enumerate(ranking, start=1):
if doc_id in seen:
continue
seen.add(doc_id)
scores[doc_id] = scores.get(doc_id, 0.0) + 1 / (rank_constant + rank)
return sorted(scores.items(), key=lambda item: (-item[1], item[0]))
keyword = ["exact-error-code", "overview", "faq"]
vector = ["overview", "exact-error-code", "runbook"]
for doc_id, score in rrf([keyword, vector]):
print(doc_id, f"{score:.6f}")
expected output
exact-error-code 0.032522
overview 0.032522
faq 0.015873
runbook 0.015873Continue reading along with the premises and constraints of the problem. Understand the reference answers first, then try to put away the answers and explain the cause and effect and trade-offs in your own words.
Level 1The correct evidence is in the top-20, but it has not entered the final prompt. How can I check it?
Moving from the main question to a specific boundary where evidence has been recalled but lost along the way.
Compare reranked output and final context snapshot item by item, check rank truncation, duplicate deduplication, token budget, parent block expansion and format conversion. Document clear reasons for discarded blocks. If the necessary evidence is in the candidate but has not been sent, sort or assemble it first; do not directly upgrade the generation model. Also check the actual text sent to the API, the candidate list displayed on the page cannot replace it.
Follow this answer further
Level 2The block is ranked in the top five, but the conditional sentence has not yet entered the prompt. Where to check next?
The parent question eliminates candidate deletions and further eliminates text transformation distortions exposed after sorting.
Check text truncation and snippet conversion, not just rankings anymore. Save the character or paragraph mapping of the original block to the actual context, especially checking headers, negation, and "only in certain versions" conditions. The same block ID does not prove that the content is the same; if the assembler only retains the first segment, it needs to reorganize according to necessary evidence instead of mechanically retaining the first N blocks.
Follow this answer further
Level 3When compressing evidence to save tokens, how can you check that the summary preserves the conclusion?
After budget repair introduces compression, the new failure is that semantic conditions are lost and fidelity needs to be checked instead of length.
List the key claims, conditions, values and source positioning as fidelity check items, and check the abstract with the original text. Compression can delete duplicate descriptions, but you cannot change "only for new indexes" to "for all indexes". If the random inspection fails, keep the original evidence or read it in batches; the summary is used for navigation, and the key judgments still refer to the original paragraphs that can be located.
Level 1Why can't we just add the BM25 score and the vector score directly?
The ranking problem further requires understanding the semantics of fractions.
The amplitude of BM25 is affected by word frequency, corpus statistics, etc. The vector score is determined by distance metric and engine transformation. The same value does not mean the same meaning. Direct addition may cause one dimension to overwhelm another. RRF can be done based on rankings, or annotated data can be used for calibration and learning fusion; the calibration results cannot be unchanged across corpus and query distribution by default.
Level 1The answer is worse after expanding top-k, what could be the reason?
After changing the number of candidates, examine how downstream constraints affect the answer inversely.
More candidates may bring in duplicate blocks, old versions, and conflicting conditions, and limited budgets crowd out key evidence; longer context does not guarantee a more faithful model. Only by measuring candidate evidence coverage and final context coverage separately can we distinguish between "the recall becomes better but the assembly becomes worse" and "the new additions are all noise." You should adjust deduplication, evidence selection, and budget instead of assuming that bigger top-k is better.
First find out the conditions for change, and then determine which premises in the original plan still hold true. The following cases are teaching deductions to facilitate the transfer of principles to new problems.
Changing conditions:From single block question and answer to multi-evidence combination
Extended question:Every document is recalled, but the answer is missing restrictions. How to locate it?
Mark the chunks that support the main conclusion and constraints into a necessary set of evidence, and check that both come together into the final context. Ordinary "hit any block" indicator will report success, so add an evidence-set completeness metric. If the complete feed is still missed, check whether the generator retains the limit as claimed instead of continuing to expand the recall.
The principles that remain unchanged:The necessary collection of evidence must be present at all stages, and single-block correlation is not a substitute for completeness of the answer.
Changing conditions:The query is the same, but the authorization set is small
Extended question:Can I temporarily remove permission filtering to confirm if there is an answer?
Production requests cannot be unfiltered. Use an offline annotation set with legal testing permissions to compare the recall behavior of engine pre-filtering and post-filtering, and record whether the authorized evidence is truncated by the candidate. After confirming the filter expression and identity version, adjust the search mode or candidate size within the authorization scope. Empty results may also indicate insufficient accessible data.
The principles that remain unchanged:Diagnostics do not extend data access boundaries, and the correct answer must be based on evidence that allows reading.
It is designed based on public technical information; the reference materials support the technical mechanism, and the scenarios and scoring standards are designed by this website and do not represent the original interview questions of a certain company. New Q&A and migration cases are added for principle explanation, and source verification and case operation verification are recorded separately.
After reading, you can explain the principles, boundaries, and trade-offs against these standards. It is up to you to evaluate your mastery; if further verification is needed, complete the small tasks below.
View verification records for independent examples
Transferring pre-retrieval filtered ideas to memory versions and recovery. Observe how existing drafts become invalid after cancellation.
Read full text and fault analysis → · Download Reliability Experiment v3 ↓
python3 cli.py memory-put --db memory.sqlite
python3 cli.py submit --db memory.sqlite
python3 cli.py run --db memory.sqlite --lease-seconds 2 --fault after_draft
python3 cli.py memory-forget --db memory.sqlite
# 等待至少 2 秒后分别执行
python3 cli.py run --db memory.sqlite
python3 cli.py inspect --db memory.sqliteVerify local scope, version, and undo blocking; old checkpoints remain, no complete deletion of logs, backups, or checkpoints is provided.
Give correct evidence for the 30th candidate, and only take the top 5 diagnoses for context.