Review the prerequisites
Suitable for: A RAG has been built and quality degradation needs to be located.
- recall
- Finding potentially relevant candidates from the database does not mean that they will eventually enter the model input.
- rearrange
- Re-prioritize candidates, affecting which material is placed first in a limited context.
- necessary evidence
- A collection of materials sufficient to support the conclusion of this question may require both text and exceptions.
- evidence support
- Whether the conclusion is supported by actual materials and applicable conditions; there may or may not be support with reference links.
How does the mechanism work?
- Confirm information
The original material exists, is valid, and has the right to be used.
- Recall and sort
Record candidate identification and ranking.
- Assembly context
Maintain identification of necessary materials, conditions and sources.
- Generate and verify
Check whether each conclusion and rejection is consistent with the evidence.
Debugging · Diagnose failures
Turn an incorrect answer into a staged experiment
Objectives of this level: Ability to preserve reproduction conditions, locate evidence loss points, and verify repairs.
Freeze the input conditions
Record the question, rewrite, permissions, knowledge version, chunks, retrieval channels, and candidate count. Label the required evidence for failures. Without these conditions, changing documents can explain inconsistent answers, making parameter attribution unreliable.
Repair the first gap
Missing exact product identifiers may suggest keyword retrieval. Correct candidates ranked too low suggest reranking. A well-ranked exception cut from context requires an assembly fix. Hybrid retrieval combines rankings or explicitly calibrated signals; unrelated score scales cannot simply be added. Azure's RRF is one concrete rank-fusion mechanism rather than a rule for every system.
Measure the new tradeoffs
Larger top-k can improve coverage while adding noise, cost, and input length. Fix answerable, unanswerable, outdated, and restricted cases and compare stage results and final support. One improved case establishes only that case's improvement.
Run experiments and observe counterexamples
Pipeline diagnostics for fixed evidence identification lists; does not perform vector retrieval, ground truth reranking, language modeling, or semantic refereeing.
Python 3.10+ · Runs by default using only the standard library · Runs on your computer
- Check the two pieces of necessary evidence
- Running context missing counterexample
- Change candidate and answer support conditions separately
python3 rag_evidence.pyView the entry-point script
"""Trace evidence survival through fixed lists; no search or LLM is executed."""
import json
def diagnose(required, candidates, ranked, context, supported_claims):
if not required.issubset(set(candidates)):
return "retrieval"
if not required.issubset(set(ranked)):
return "ranking"
if not required.issubset(set(context)):
return "context"
if not supported_claims:
return "generation"
return "supported"
def demo():
required = {"policy-current", "policy-exception"}
candidates = ["old-policy", "policy-current", "policy-exception"]
ranked = ["policy-current", "policy-exception", "old-policy"]
stage = diagnose(required, candidates, ranked, ranked[:1], True)
fixed = diagnose(required, candidates, ranked, ranked[:2], True)
assert stage == "context" and fixed == "supported"
return dict(first_failure=stage, fixed_evidence_check=fixed,
generation_still_needs_review=True)
if __name__ == "__main__":
print(json.dumps(demo(), sort_keys=True))
Expected output when running locally
{"first_failure": "context", "fixed_evidence_check": "supported", "generation_still_needs_review": true}- Locating the earliest evidence gaps
- Collection checking separate from semantic checking
- Final repair needs to verify the true conclusion
Acceptance task for this level
Write diagnostic and control experiments for "Candidate No. 30 had correct evidence, but ended up using only 5 blocks."
Check each item after completion
- First confirm whether the first 5 blocks lack necessary evidence
- Change ordering and context scale separately
- Also review unanswered samples, missing conditions, and input costs
Save your own processes, code and results. Acceptance requirements are provided here, and course mastery status will not be automatically graded or saved at this time.
Hide the answer and check your understanding
Increasing top-k for all problems, what possible degradations are there?
Expand reference derivation
Noise and duplicate materials may crowd out key conditions and increase capacity and delay costs; the final conclusion may still be unfaithful to the evidence, so fixed samples are needed for comparison.
Further explanations and practice
When encountering unfamiliar principles, first read the implementation, continuous questioning and migration cases, and then independently explain the premise and boundaries. Answers and notes are saved to the original account record.
All linked explanations and exercises (4 )
- Evidence flow and stage-by-stage RAG diagnosis · answer independently
- Semantic completeness and evidence location in chunking · answer independently
- Complementary retrieval and rank fusion in hybrid search · answer independently
- Evidence support, partial answers, and calibrated abstention · answer independently
Sources and verification scope
The principles are based on public information; the numbers, cases and tasks are the teaching design of this website. Offline experiments verify the range noted on this page, and the learning effect still needs to be judged through independent tasks and feedback.