Agent Application DevelopmentAccount
Knowledge catalogChoose core direction and segmented content

SYSTEMATIC LEARNING / FOUR-LEVEL COURSE

RAG evidence flow and failure diagnosis

Follow the same evidence through retrieval, ranking, context assembly, and the answer to see where the information needed for a correct answer can be lost.

Learning objectives: Use required evidence to locate the first failing stage, and compare the conditions and costs of different fixes.

Content checked: 2026-10-04 · Each level has independent explanations, tasks and inspections

Choose a starting point based on your familiarity with this topic. Current level: Foundation · Understand the concepts. After completing the task, continue to the next level. Reading and self-checks alone do not establish mastery.

On this level

Review the prerequisites

Suitable for: First Contact Retrieval Enhanced Answers.

recall
Finding potentially relevant candidates from the database does not mean that they will eventually enter the model input.
rearrange
Re-prioritize candidates, affecting which material is placed first in a limited context.
necessary evidence
A collection of materials sufficient to support the conclusion of this question may require both text and exceptions.
evidence support
Whether the conclusion is supported by actual materials and applicable conditions; there may or may not be support with reference links.

How does the mechanism work?

  1. Confirm information

    The original material exists, is valid, and has the right to be used.

  2. Recall and sort

    Record candidate identification and ranking.

  3. Assembly context

    Maintain identification of necessary materials, conditions and sources.

  4. Generate and verify

    Check whether each conclusion and rejection is consistent with the evidence.

Foundation · Understand the concepts

Why is the answer wrong even though the document was retrieved?

Objectives of this level: It can be explained that retrieval, feeding into the model and supporting the conclusion are different stages.

Think of RAG as an open-book answer

RAG retrieves material and puts it into model context before generating an answer. It changes the available evidence without guaranteeing every conclusion. Missing answers, expired documents, and unauthorized sources cannot be repaired simply by using a more capable model.

One conclusion can require multiple passages

In this teaching scenario, ordinary refunds have a seven-day window, while presale items follow another condition. Supplying only the ordinary rule can produce a plausible but incorrect answer. Label the full set of evidence needed for the conclusion rather than one relevant document.

Evidence can disappear at every stage

The correct passage may never be retrieved, rank too low, or be cut from context after correct ranking. Even complete context can be misinterpreted during generation. Diagnose the earliest gap to choose between repairing retrieval, ranking, assembly, and generation.

Run experiments and observe counterexamples

Pipeline diagnostics for fixed evidence identification lists; does not perform vector retrieval, ground truth reranking, language modeling, or semantic refereeing.

Python 3.10+ · Runs by default using only the standard library · Runs on your computer

  1. Check the two pieces of necessary evidence
  2. Running context missing counterexample
  3. Change candidate and answer support conditions separately
Downloadrag_evidence.py ↓
python3 rag_evidence.py
View the entry-point script
"""Trace evidence survival through fixed lists; no search or LLM is executed."""
import json


def diagnose(required, candidates, ranked, context, supported_claims):
    if not required.issubset(set(candidates)):
        return "retrieval"
    if not required.issubset(set(ranked)):
        return "ranking"
    if not required.issubset(set(context)):
        return "context"
    if not supported_claims:
        return "generation"
    return "supported"


def demo():
    required = {"policy-current", "policy-exception"}
    candidates = ["old-policy", "policy-current", "policy-exception"]
    ranked = ["policy-current", "policy-exception", "old-policy"]
    stage = diagnose(required, candidates, ranked, ranked[:1], True)
    fixed = diagnose(required, candidates, ranked, ranked[:2], True)
    assert stage == "context" and fixed == "supported"
    return dict(first_failure=stage, fixed_evidence_check=fixed,
                generation_still_needs_review=True)


if __name__ == "__main__":
    print(json.dumps(demo(), sort_keys=True))

Expected output when running locally

{"first_failure": "context", "fixed_evidence_check": "supported", "generation_still_needs_review": true}
  • Locating the earliest evidence gaps
  • Collection checking separate from semantic checking
  • Final repair needs to verify the true conclusion
View the running environment, output and verification records →

Acceptance task for this level

Draw the five stages of information, candidate, sorting, context and answer, and mark at which step general terms and exceptions are lost.

Check each item after completion

  • Distinguish between candidates and actual input
  • Necessary evidence includes both rules and exceptions
  • Don’t attribute all failures to the model not being strong enough

Save your own processes, code and results. Acceptance requirements are provided here, and course mastery status will not be automatically graded or saved at this time.

Hide the answer and check your understanding

If the correct document appears in the retrieval log, does it prove that the model has seen the correct terms?

Further explanations and practice

When encountering unfamiliar principles, first read the implementation, continuous questioning and migration cases, and then independently explain the premise and boundaries. Answers and notes are saved to the original account record.

All linked explanations and exercises (4 )

Sources and verification scope

The principles are based on public information; the numbers, cases and tasks are the teaching design of this website. Offline experiments verify the range noted on this page, and the learning effect still needs to be judged through independent tasks and feedback.