Review the prerequisites
Suitable for: Prepare to write the RAG principle into a reproducible project.
- Warehouse
- Parse documents, determine tiles and metadata, and record versions and summaries for content and sources.
- Embedding
- Text is mapped to vectors by the specified model; documents and queries need to use compatible models and processing.
- Retrieval candidates
- Vector or keyword ranking provides potentially relevant blocks, filtered by current permissions and versions.
- RRF
- The results are fused according to the ranking of each ranking; the fusion constant in this experiment is 60, and the final number of contexts is controlled by top_k.
- Necessary evidence recall
- How many of the required human-annotated sources enter the context, calculated per answerable question.
- vector snapshot
- The real local model inference artifacts of this batch of texts, with models, dimensions, corpus summaries and model file summaries, support offline playback without model download.
How does the mechanism work?
- Parse and preserve sources
Five copies of fictional English Markdown generate seven independent blocks, retaining tenants, versions, chapters, body, and abstracts.
- Generate compatibility vector
The authors actually ran bge-small-en-v1.5, recording 384-dimensional vectors of seven document blocks and three queries.
- Filter first then rank
SQLite parameterizes filtering of the current tenant and valid versions, and then performs keyword and cosine searches.
- Fusion and assembly
BM25 is fused with vector ranking using RRF, intercepting the final context and writing the request with the source ID.
- Acceptance Candidate
Author candidates are checked for citation membership, permissions, version and limited policy tags, and free text semantics are individually marked.
Implementation · Build it
Run parsing, retrieval, context assembly, and acceptance checks step by step
Objectives of this level: Run real document parsing and retrieval algorithms, view vector artifacts, and generate evidence-based request and acceptance reports.
Run the complete offline pipeline
Download and extract this page's lab, then run:
python3 rag_pipeline.py
python3 application_project.py > application-report.json
python3 -m unittest test_application -v
The first command reports seven chunks, a 384-dimensional model, actual contextIds for three queries, contract checks, and required-evidence recall. The second saves a combined model-response, prompt, and RAG project artifact with each question's request, candidate, and failure categories.
Read application_data.py for document parsing, embeddings.json for recorded inference vectors, and rag_pipeline.py for the in-memory SQLite index, filtering, cosine/BM25 ranking, and RRF fusion. prompt_iteration.py builds requests and runs the grader. Expected labels enter evaluation only; model input contains the question and selected text.
Retain retrieval evidence
Each chunk has a stable ID, tenant, current flag, text, and digest. Parameterized SQL selects the allowed set. Both vector and keyword ranking run only within that set, excluding the other tenant's VIP policy and old versions. The lab assumes trusted ingestion metadata. Production permissions must come from real sessions and authorization data.
dense and keyword retain both rankings. RRF sums 1/(60+rank) and selects the final top_k. This is a fusion baseline; compare a model reranker as an additional stage. In-memory tables and exact cosine scans make the mechanism inspectable. Larger corpora can compare inverted indexes, ANN, and filtering strategies.
Rebuild embeddings yourself
Default replay reads recorded vectors. To change documents or queries, install the verified version and rebuild in your own environment:
python3 -m venv .venv
.venv/bin/python -m pip install fastembed==0.8.1
HF_HUB_DISABLE_IMPLICIT_TOKEN=1 .venv/bin/python rebuild_embeddings.py --cache-dir .model-cache
.venv/bin/python rag_pipeline.py
The first run needs network access to download the open model and uses local CPU inference. The script uses the same model's passage and query paths and records model-file SHA-256 digests, library version, and corpus digest. Later replay needs no inference library. Platforms may differ in floating-point results, so retain your own artifacts and run conditions.
Connect live generation
build_prompt supplies actually retrieved evidence. Connect the request to the live API step in the model-response lesson and retain the candidate, usage, and versions. Classify completed, refused, and incomplete responses before reusing the grader. Current report candidates are authored_fixtures, and natural-language support remains not_scored.
Run experiments and observe counterexamples
Practical parsing of fictional English Markdown, running SQL filtering, cosine, BM25, RRF, request assembly and limited contract evaluation. Document and fixed query vectors are derived from real local Embedding inference snapshots; generated candidates are written by the authors and free text semantics are not scored.
Python 3.10+ · Runs by default using only the standard library · Runs on your computer
- Download the Agent application entry experimental package on this page, unzip it and enter the agent-application-lab-v1 directory.
- Use Python 3.10+ to execute the above command; the default playback only requires the standard library and package data.
- Compare the output with the checkpoint, then run python3 -m unittest test_application -v and complete the current layer task.
python3 rag_pipeline.pyView the entry-point script
"""Actual parsing, SQL filtering, vector search, BM25, RRF and contract evaluation.
Embeddings are replayed from an actual local-model inference snapshot. The generation
stage uses explicitly authored fixtures. No network or inference is needed for replay.
"""
from collections import Counter
import json
import math
import re
import sqlite3
from application_data import FIXTURES, corpus_hash, load_chunks, read_cases
from prompt_iteration import CANDIDATES, build_prompt, grade
def load_snapshot(chunks, path=None):
snapshot = json.loads((path or FIXTURES / "embeddings.json").read_text(encoding="utf-8"))
if snapshot.get("schemaVersion") != 1 or snapshot.get("corpusSha256") != corpus_hash(chunks):
raise ValueError("embedding snapshot is stale; rebuild after changing documents")
if set(snapshot["documents"]) != {chunk["id"] for chunk in chunks}:
raise ValueError("embedding snapshot has different chunk identities")
dimension = snapshot["dimension"]
if type(dimension) is not int or dimension < 1 or not snapshot.get("model"):
raise ValueError("invalid embedding contract")
for vector in [*snapshot["documents"].values(), *snapshot["queries"].values()]:
if (len(vector) != dimension or not all(type(x) in (int, float) and math.isfinite(x) for x in vector)
or not any(vector)):
raise ValueError("invalid embedding vector")
return snapshot
def tokenize(text):
return re.findall(r"[a-z0-9]+", text.lower())
def cosine(left, right):
return sum(a * b for a, b in zip(left, right)) / (math.sqrt(sum(a * a for a in left)) * math.sqrt(sum(b * b for b in right)))
def keyword_scores(query, chunks):
terms = set(tokenize(query))
bags = [Counter(tokenize(chunk["text"])) for chunk in chunks]
average = sum(sum(bag.values()) for bag in bags) / len(bags)
scores = []
for chunk, bag in zip(chunks, bags):
score = 0.0
for term in terms:
frequency = bag[term]
if frequency:
df = sum(term in other for other in bags)
idf = math.log(1 + (len(bags) - df + 0.5) / (df + 0.5))
score += idf * frequency * 2.2 / (frequency + 1.2 * (0.25 + 0.75 * sum(bag.values()) / average))
scores.append((chunk["id"], score))
return sorted((row for row in scores if row[1] > 0), key=lambda row: (-row[1], row[0]))
def index_chunks(chunks):
db = sqlite3.connect(":memory:")
db.execute("CREATE TABLE chunks(id TEXT PRIMARY KEY, tenant TEXT NOT NULL, current INTEGER NOT NULL, body TEXT NOT NULL)")
db.executemany("INSERT INTO chunks VALUES(?,?,?,?)", [(chunk["id"], chunk["tenant"], int(chunk["current"]), json.dumps(chunk)) for chunk in chunks])
return db
def retrieve(query, chunks, snapshot, tenant="shop-a", top_k=3):
if type(top_k) is not int or top_k < 1:
raise ValueError("top_k must be a positive integer")
if query not in snapshot["queries"]:
raise ValueError("query was not recorded; regenerate embeddings for the changed query")
db = index_chunks(chunks)
try:
allowed = [json.loads(row[0]) for row in db.execute("SELECT body FROM chunks WHERE tenant=? AND current=1 ORDER BY id", (tenant,))]
finally:
db.close()
if not allowed:
return {"allowed": 0, "dense": [], "keyword": [], "context": []}
query_vector = snapshot["queries"][query]
dense = sorted([(chunk["id"], cosine(query_vector, snapshot["documents"][chunk["id"]])) for chunk in allowed], key=lambda row: (-row[1], row[0]))
keyword = keyword_scores(query, allowed)
fused = Counter()
for ranking in (dense, keyword):
for rank, (chunk_id, _) in enumerate(ranking, 1):
fused[chunk_id] += 1 / (60 + rank)
selected = [cid for cid, _ in sorted(fused.items(), key=lambda row: (-row[1], row[0]))[:top_k]]
by_id = {chunk["id"]: chunk for chunk in allowed}
return {"allowed": len(allowed), "dense": [cid for cid, _ in dense], "keyword": [cid for cid, _ in keyword], "context": [by_id[cid] for cid in selected]}
def run_pipeline(top_k=3):
chunks = load_chunks()
snapshot = load_snapshot(chunks)
rows = []
for case in read_cases():
retrieved = retrieve(case["query"], chunks, snapshot, top_k=top_k)
request = build_prompt(case, retrieved["context"])
answer = CANDIDATES["v2"][case["id"]]
context_ids = [chunk["id"] for chunk in retrieved["context"]]
required = set(case["requiredEvidence"])
rows.append({"case": case["id"], "contextIds": context_ids,
"request": request, "candidate": answer,
"evidenceRecall": len(required.intersection(context_ids)) / len(required) if required else None,
**grade(case, answer, retrieved["context"])})
return {"chunks": chunks, "snapshot": snapshot, "rows": rows}
def demo():
report = run_pipeline()
answerable = [row["evidenceRecall"] for row in report["rows"] if row["evidenceRecall"] is not None]
return {"chunks": len(report["chunks"]), "dimension": report["snapshot"]["dimension"],
"embeddingModel": report["snapshot"]["model"], "queries": len(report["rows"]),
"contextIds": {row["case"]: row["contextIds"] for row in report["rows"]},
"contractPasses": sum(row["contractPassed"] for row in report["rows"]),
"meanNecessaryEvidenceRecallAt3": sum(answerable) / len(answerable),
"semanticQuality": "not_scored", "generation": "authored_fixtures",
"embedding": "real_local_inference_snapshot"}
if __name__ == "__main__":
print(json.dumps(demo(), ensure_ascii=False, sort_keys=True))
Expected output when running locally
{"chunks": 7, "contextIds": {"battery": ["returns:v2:battery", "returns:v2:general", "warranty:v1:defects"], "ordinary": ["returns:v2:general", "returns:v2:battery", "shipping:v1:contact"], "unknown": ["returns:v2:general", "returns:v2:battery", "warranty:v1:defects"]}, "contractPasses": 3, "dimension": 384, "embedding": "real_local_inference_snapshot", "embeddingModel": "BAAI/bge-small-en-v1.5", "generation": "authored_fixtures", "meanNecessaryEvidenceRecallAt3": 1.0, "queries": 3, "semanticQuality": "not_scored"}- chunks=7, dimension=384. See embeddings.json for the model and artifact digests.
- Three queries are actually run and ranked, with current tenant and version filtering occurring at the top of the ranking.
- The meanNecessaryEvidenceRecallAt3=1.0 of the fixed corpus is calculated, and only two answerable questions are counted.
- generation=authored_fixtures、semanticQuality=not_scored。
Acceptance task for this level
Deliver application-report.json, which points out the original text, chunk, vector model, two rankings, actual context, model input and references one by one. Change top_k to 1 and save a failure report.
Check each item after completion
- Document parsing does come from in-package Markdown.
- The document and query vector versions and dimensions are consistent.
- Old versions and other tenants do not enter the current search ranking.
- There is no expected tag in the final request.
- Top_k changes preserve intermediate results and acceptance differences.
Save your own processes, code and results. Acceptance requirements are provided here, and course mastery status will not be automatically graded or saved at this time.
Hide the answer and check your understanding
What will happen if you only modify the 30 days in the corpus to 20 days and then run the default snapshot?
Expand reference derivation
The document and corpus summary change, the loader reports that the snapshot is expired and stops. After modification, regenerate documents and query vectors, synchronize labels and candidates, and then compare the results. Currently, the entire corpus is conservatively bound, and production can further manage invalidations according to content, permissions, and versions.
Further explanations and practice
When encountering unfamiliar principles, first read the implementation, continuous questioning and migration cases, and then independently explain the premise and boundaries. Answers and notes are saved to the original account record.
All linked explanations and exercises (5 )
- Evidence flow and stage-by-stage RAG diagnosis · answer independently
- Semantic completeness and evidence location in chunking · answer independently
- Complementary retrieval and rank fusion in hybrid search · answer independently
- Evidence support, partial answers, and calibrated abstention · answer independently
- Version closure and incremental consistency during embedding migration · answer independently
Sources and verification scope
The principles are based on public information; the numbers, cases and tasks are the teaching design of this website. Offline experiments verify the range noted on this page, and the learning effect still needs to be judged through independent tasks and feedback.