Agent Application DevelopmentAccount
Knowledge catalogChoose core direction and segmented content
knowledge unit 17IntermediateSystem designAbout 15 minutes

Understand → Implement → Debug → Design

Evidence flow and stage-by-stage RAG diagnosis

Establish staged evidence and indicators to avoid all errors being solved by changing models.

RAGrecallrearrangeReview

Knowledge content check2026-10-03 · Check the source of the original question2026-10-02

Which step do you want to learn from this knowledge point?

Select the starting point based on the current basis, or you can go deeper one by one. When you encounter an unfamiliar concept, go back to the core principles first; use the knowledge exercises to check your understanding when you are finished.

Understand first

New to this knowledge point

Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.

Start with core principles →

Implement next

Prepare to write the principles into code

Understand implementation steps and boundaries, complete small tasks, and check results against acceptance requirements.

View the code example →

Debug failures

Need to handle failures and changes in conditions

Follow the continuous questioning to locate the failure premise, and then compare the migration cases to explain how the plan should be adjusted.

Continue to delve deeper into the problem →

Compare designs

Need to design or review plans

Combine engineering deductions and senior self-evaluation standards to explain the applicable conditions, costs and alternatives of the plan.

Analyze engineering scenarios →
Knowledge unit directory

LEARN · PRACTICE · REFLECT

Knowledge learning and personal records

My notes and review ↗

First read along the principles, Q&A and migration cases. When you need to check your understanding, switch to reinforcement exercises or start personal recording.

Answers and personal notes

Each modified commit will be kept as an independent history. Your level of mastery is up to you to evaluate yourself against the standards.

Core concept · Evidence flow and stage-by-stage RAG diagnosis

Understand the core principles first

Preparatory concepts:information retrieval ranking, Pipeline intermediate results, Labeled datasets and controlled comparisons

The evidence that answers the dependency must first exist, be recalled, retained, and fed into the model in its entirety. Find the earliest loss point along the evidence flow to know which stage of changes can affect the error.

Inspect RAG as a pipeline that can lose information

An upstream field can disappear during parsing, filtering, ranking, or serialization. Likewise, an answer existing in a document, candidate set, context, and final response are different facts. Diagnose from the source rather than assuming every error is poor retrieval.

Annotate minimum evidence, such as an error code, trigger condition, and remedy. Fix index, authorization, and query versions; record each stage’s inputs, outputs, and exclusion reasons by chunk ID. Evidence ranked eighteenth may disappear when only five chunks are assembled. A wrong answer may require better ranking, assembly, or combining evidence.

Metrics describe different stages

Recall@k measures candidate coverage; ranking metrics measure position; context inspection checks complete text and conditions; claim evaluation checks support. Vector and RRF scores are ranking signals, not factual probabilities. Final accuracy can conceal compensating stage errors; recall alone can reward excessive noise.

Substitute evidence to identify causes

Keep the generator fixed and replace candidates with curated evidence. Remaining errors point toward assembly or generation. Supplying complete evidence directly helps isolate instructions, reasoning, and verification. A stronger model answering from learned knowledge while citing incorrectly does not establish repaired retrieval. Include unanswerable, permission-restricted, and version-specific questions so increased answer volume cannot masquerade as reliability.

Check understanding with a question

When RAG gives an incorrect answer, how to determine whether it is a recall, reranking, context, or generated question?

Fix the question, index version, and access scope. Save rewritten queries, candidates, rankings, actual model context, and citations. Missing candidates indicate data or retrieval problems; dropped candidates indicate ranking problems; truncated selected evidence indicates assembly problems. Wrong answers with complete evidence point toward generation. Evaluate annotated questions by stage rather than fluency alone.

Implementation and trade-offs

Fixed reproducible diagnostic input

The original question, rewritten question, user permissions, knowledge version, chunking strategy and search parameters are retained. First confirm whether the original data contains answers and whether tables, titles or key conditions are missing in the database analysis. When a document does not exist or is out of date, no matter how powerful the model is, it can only guess. Document_id, chunk_id, version, retrieval channel, and ranking are logged for each candidate; these are used for diagnostics, but the logs themselves must also adhere to permissions and anonymization requirements. Without a snapshot of the actual context, looking only at the final answer cannot reliably locate the error.

Compare whether the evidence is still there stage by stage

Label the test questions with pieces of evidence that support the answer. The candidate stage checks Recall@k; the sorting stage looks at the location of the correct evidence, using MRR or nDCG; the context stage checks whether key paragraphs are complete, whether conflicting versions have been mixed, and whether the reference identification still corresponds to the original text; the generation stage checks whether each conclusion is supported by evidence and whether applicable conditions are missing. Just because the evidence is in the candidate list does not mean that the model has seen it, and the fact that the evidence is in does not mean that the answer actually refers to it. Rejected questions should be evaluated separately to avoid counting only samples with answers.

Optimize to match failure reasons

Terminology, error codes, and precise identification often require keyword retrieval, and expression changes can be supplemented by vector recall. Hybrid retrieval requires merging different rankings and cannot directly add similarities and BM25 scores with different dimensions. Azure AI Search uses RRF to fuse by rank, semantic reranking is a subsequent stage, and the scores have other meanings. The following code only demonstrates the fusion mechanism. If the recall is insufficient, then check the rewriting, chunking, filtering and number of candidates; if the sorting is not good, adjust the reranking; if the context is insufficient, reduce the repetition and retain the conditional paragraphs; only adjust the instructions and output verification when it is unfaithful.

How to prove that the changes are effective

Pinned a batch of questions covering precise terminology, vague expressions, cross-document reasoning, no answers, and permission restrictions. Only change one type of configuration at a time, and compare stage indicators, final evidence support rate, latency and cost. Manually review failed samples and mark root causes. You cannot use a score from the model's self-evaluation to replace all factual judgments. Data, index, reranker and prompt versions are recorded together for playback. Getting a certain problem better does not mean that the whole problem is getting better. In particular, expanding top-k may bring more noise and context costs.

code example

Implementing a minimal RRF fusion using ranking only

Python 3 standard library demo. The inputs are two hypothetical ranked lists, without real retrieval, vector model or semantic reranking; rank_constant is a different parameter than retrieval top-k.

def rrf(rankings, rank_constant=60):
    if rank_constant <= 0:
        raise ValueError("rank_constant must be positive")
    scores = {}
    for ranking in rankings:
        seen = set()
        for rank, doc_id in enumerate(ranking, start=1):
            if doc_id in seen:
                continue
            seen.add(doc_id)
            scores[doc_id] = scores.get(doc_id, 0.0) + 1 / (rank_constant + rank)
    return sorted(scores.items(), key=lambda item: (-item[1], item[0]))

keyword = ["exact-error-code", "overview", "faq"]
vector = ["overview", "exact-error-code", "runbook"]
for doc_id, score in rrf([keyword, vector]):
    print(doc_id, f"{score:.6f}")

expected output

exact-error-code 0.032522
overview 0.032522
faq 0.015873
runbook 0.015873

Engineering deduction

scene
Hypothetical engineering scenario: The user asks for an interface error code, and the vector search returns a concept introduction, but misses the operation and maintenance document containing the precise error code.
design decisions
Add keyword recall and use ranking fusion; record candidates and final context, review error codes and applicable versions.
Verify target
Expected behavior: After the exact document is entered as a candidate, it can be checked whether it is retained by the reranking, and whether the answer reference corresponds to it is tracked.
applicable boundary
There is no promised metric improvement for the examples; the fusion parameters, number of candidates, and reranking selection must be verified on the local annotation set.

Continuous questions and answers

Continue reading along with the premises and constraints of the problem. Understand the reference answers first, then try to put away the answers and explain the cause and effect and trade-offs in your own words.

Draw inferences from one example: If the conditions change, how to deduce it?

First find out the conditions for change, and then determine which premises in the original plan still hold true. The following cases are teaching deductions to facilitate the transfer of principles to new problems.

The answer needs to span two documents

Changing conditions:From single block question and answer to multi-evidence combination

Extended question:Every document is recalled, but the answer is missing restrictions. How to locate it?

Derivation and reference solutions

Mark the chunks that support the main conclusion and constraints into a necessary set of evidence, and check that both come together into the final context. Ordinary "hit any block" indicator will report success, so add an evidence-set completeness metric. If the complete feed is still missed, check whether the generator retains the limit as claimed instead of continuing to expand the recall.

The principles that remain unchanged:The necessary collection of evidence must be present at all stages, and single-block correlation is not a substitute for completeness of the answer.

Suddenly the result is empty after the permissions are narrowed.

Changing conditions:The query is the same, but the authorization set is small

Extended question:Can I temporarily remove permission filtering to confirm if there is an answer?

Derivation and reference solutions

Production requests cannot be unfiltered. Use an offline annotation set with legal testing permissions to compare the recall behavior of engine pre-filtering and post-filtering, and record whether the authorized evidence is truncated by the candidate. After confirming the filter expression and identity version, adjust the search mode or candidate size within the authorization scope. Empty results may also indicate insufficient accessible data.

The principles that remain unchanged:Diagnostics do not extend data access boundaries, and the correct answer must be based on evidence that allows reading.

Easy to make mistakes

  • Only upgrade the generative model when no answer can be retrieved
  • Only the final answer is recorded, candidates and context are not recorded
  • Treat the RRF score or similarity as the probability of answer authenticity

References

It is designed based on public technical information; the reference materials support the technical mechanism, and the scenarios and scoring standards are designed by this website and do not represent the original interview questions of a certain company. New Q&A and migration cases are added for principle explanation, and source verification and case operation verification are recorded separately.

Check how far you understand

After reading, you can explain the principles, boundaries, and trade-offs against these standards. It is up to you to evaluate your mastery; if further verification is needed, complete the small tasks below.

Basic standards met
Can split parsing, recall, reranking, context and generation.
Intermediate and advanced signals
Locate missing points using intermediate candidates, necessary evidence, and a fixed query set.
Senior criteria
Isolate the source of the problem through ablation and combine authority, timeliness and final reference judgment.

View verification records for independent examples

Continue to do advanced research experiments

Why can’t the old context continue to be used after the policy changes?

Transferring pre-retrieval filtered ideas to memory versions and recovery. Observe how existing drafts become invalid after cancellation.

Read full text and fault analysis → · Download Reliability Experiment v3 ↓

python3 cli.py memory-put --db memory.sqlite
python3 cli.py submit --db memory.sqlite
python3 cli.py run --db memory.sqlite --lease-seconds 2 --fault after_draft
python3 cli.py memory-forget --db memory.sqlite
# 等待至少 2 秒后分别执行
python3 cli.py run --db memory.sqlite
python3 cli.py inspect --db memory.sqlite

Keep evidence and check item by item

  • The draft checkpoint already exists after the first exit.
  • Recovery after memory-forget gets failed with memory_changed_or_expired.
  • No publish checkpoint; explain the difference between fail blocking and complete deletion.

Verify local scope, version, and undo blocking; old checkpoints remain, no complete deletion of logs, backups, or checkpoints is provided.

Hands-on verificationComplete on demand · Suggestions15 minutes

Give correct evidence for the 30th candidate, and only take the top 5 diagnoses for context.

Expand acceptance requirements and checkpoints
  • Not directly attributable to model capabilities
  • Able to distinguish between recall and reranking
  • Review final answer after fix

Key inspections

  • Locate specific failure stages through evidence trails
  • Understand that vector scores, BM25 and RRF scores are not on the same scale
  • Able to design labeled and reproducible stage evaluations