Agent Application DevelopmentAccount
Knowledge catalogChoose core direction and segmented content

SYSTEMATIC LEARNING / FOUR-LEVEL COURSE

RAG evidence flow and failure diagnosis

Follow the same evidence through retrieval, ranking, context assembly, and the answer to see where the information needed for a correct answer can be lost.

Learning objectives: Use required evidence to locate the first failing stage, and compare the conditions and costs of different fixes.

Content checked: 2026-10-04 · Each level has independent explanations, tasks and inspections

Choose a starting point based on your familiarity with this topic. Current level: Design · Explain the trade-offs. After completing the task, continue to the next level. Reading and self-checks alone do not establish mastery.

On this level

Review the prerequisites

Suitable for: It is necessary to design a quality system for corporate knowledge Q&A.

recall
Finding potentially relevant candidates from the database does not mean that they will eventually enter the model input.
rearrange
Re-prioritize candidates, affecting which material is placed first in a limited context.
necessary evidence
A collection of materials sufficient to support the conclusion of this question may require both text and exceptions.
evidence support
Whether the conclusion is supported by actual materials and applicable conditions; there may or may not be support with reference links.

How does the mechanism work?

  1. Confirm information

    The original material exists, is valid, and has the right to be used.

  2. Recall and sort

    Record candidate identification and ranking.

  3. Assembly context

    Maintain identification of necessary materials, conditions and sources.

  4. Generate and verify

    Check whether each conclusion and rejection is consistent with the evidence.

Design · Explain the trade-offs

Balance sufficient evidence, access permissions, and refusal

Objectives of this level: Able to define evidence contracts and design phased indicators and rejection criteria.

Work backward from the answer contract

Define required sources, versions, and conditions for each conclusion before choosing parsing, chunking, retrieval, and assembly. Contract exceptions, table headings, and code-call relationships have different evidence structures. A uniform character limit cannot define every useful chunk. The goal is complete, traceable support.

Enforce permissions along the evidence flow

Unauthorized text must not enter candidates, caches, or source-readback paths. Relevance ranking performs no authentication. Source access, derived summaries, and caches must follow permission and version changes. Allow partial answers or abstention when evidence is insufficient, and identify what is missing.

Compare designs with layered evaluation

Evaluate corpus coverage, required-evidence recall, ranking, assembly, claim support, and abstention separately. Open-ended text may use human review or calibrated graders; permissions remain independent hard checks. Thresholds depend on business risk and samples. This lesson supplies no universal accuracy target for release.

Run experiments and observe counterexamples

Pipeline diagnostics for fixed evidence identification lists; does not perform vector retrieval, ground truth reranking, language modeling, or semantic refereeing.

Python 3.10+ · Runs by default using only the standard library · Runs on your computer

  1. Check the two pieces of necessary evidence
  2. Running context missing counterexample
  3. Change candidate and answer support conditions separately
Downloadrag_evidence.py ↓
python3 rag_evidence.py
View the entry-point script
"""Trace evidence survival through fixed lists; no search or LLM is executed."""
import json


def diagnose(required, candidates, ranked, context, supported_claims):
    if not required.issubset(set(candidates)):
        return "retrieval"
    if not required.issubset(set(ranked)):
        return "ranking"
    if not required.issubset(set(context)):
        return "context"
    if not supported_claims:
        return "generation"
    return "supported"


def demo():
    required = {"policy-current", "policy-exception"}
    candidates = ["old-policy", "policy-current", "policy-exception"]
    ranked = ["policy-current", "policy-exception", "old-policy"]
    stage = diagnose(required, candidates, ranked, ranked[:1], True)
    fixed = diagnose(required, candidates, ranked, ranked[:2], True)
    assert stage == "context" and fixed == "supported"
    return dict(first_failure=stage, fixed_evidence_check=fixed,
                generation_still_needs_review=True)


if __name__ == "__main__":
    print(json.dumps(demo(), sort_keys=True))

Expected output when running locally

{"first_failure": "context", "fixed_evidence_check": "supported", "generation_still_needs_review": true}
  • Locating the earliest evidence gaps
  • Collection checking separate from semantic checking
  • Final repair needs to verify the true conclusion
View the running environment, output and verification records →

Acceptance task for this level

Define evidence and rejection contracts for enterprise refund Q&A, and add three samples: old policy, pre-sale exception, and no permission.

Check each item after completion

  • Conclusion binding source, version and applicable conditions
  • Unauthorized material does not flow into models or references
  • Distinguish between partial answers, refusals, and well-documented complete answers

Save your own processes, code and results. Acceptance requirements are provided here, and course mastery status will not be automatically graded or saved at this time.

Hide the answer and check your understanding

A valid link is referenced, why should it still be rejected?

Further explanations and practice

When encountering unfamiliar principles, first read the implementation, continuous questioning and migration cases, and then independently explain the premise and boundaries. Answers and notes are saved to the original account record.

All linked explanations and exercises (4 )

Sources and verification scope

The principles are based on public information; the numbers, cases and tasks are the teaching design of this website. Offline experiments verify the range noted on this page, and the learning effect still needs to be judged through independent tasks and feedback.