Agent Application DevelopmentAccount
Knowledge catalogChoose core direction and segmented content

SYSTEMATIC LEARNING / FOUR-LEVEL COURSE

From document ingestion to cited RAG answers

Run Markdown parsing, replay recorded embedding vectors, apply SQL permission filters, and inspect BM25, vector retrieval, RRF, and answer acceptance checks.

Learning objectives: Inspect each intermediate RAG artifact, compare how top-k affects required evidence, and distinguish retrieval, generation, and semantic evaluation results.

Content checked: 2026-10-04 · Each level has independent explanations, tasks and inspections

Choose a starting point based on your familiarity with this topic. Current level: Foundation · Understand the concepts. After completing the task, continue to the next level. Reading and self-checks alone do not establish mastery.

On this level

Review the prerequisites

Suitable for: Understand model requests, hints and tool contracts, and do RAG for the first time.

Warehouse
Parse documents, determine tiles and metadata, and record versions and summaries for content and sources.
Embedding
Text is mapped to vectors by the specified model; documents and queries need to use compatible models and processing.
Retrieval candidates
Vector or keyword ranking provides potentially relevant blocks, filtered by current permissions and versions.
RRF
The results are fused according to the ranking of each ranking; the fusion constant in this experiment is 60, and the final number of contexts is controlled by top_k.
Necessary evidence recall
How many of the required human-annotated sources enter the context, calculated per answerable question.
vector snapshot
The real local model inference artifacts of this batch of texts, with models, dimensions, corpus summaries and model file summaries, support offline playback without model download.

How does the mechanism work?

  1. Parse and preserve sources

    Five copies of fictional English Markdown generate seven independent blocks, retaining tenants, versions, chapters, body, and abstracts.

  2. Generate compatibility vector

    The authors actually ran bge-small-en-v1.5, recording 384-dimensional vectors of seven document blocks and three queries.

  3. Filter first then rank

    SQLite parameterizes filtering of the current tenant and valid versions, and then performs keyword and cosine searches.

  4. Fusion and assembly

    BM25 is fused with vector ranking using RRF, intercepting the final context and writing the request with the source ID.

  5. Acceptance Candidate

    Author candidates are checked for citation membership, permissions, version and limited policy tags, and free text semantics are individually marked.

Foundation · Understand the concepts

Why can an answer miss an exception when the relevant passages are present?

Objectives of this level: Draw the evidence path from document to answer, and explain the responsibilities of warehousing, retrieval and generation.

Start with a question requiring two sources

A user asks whether a damaged battery purchased 12 days ago can be returned. The general rule allows 30 days; the battery exception allows 7 days followed by support review. A complete explanation needs both rules and their relationship. Finding a relevant paragraph about returns does not establish sufficient context.

The project supplies five fictional English Markdown documents containing current rules, an old version, another tenant's VIP rule, and shipping and warranty information. The teaching format begins with JSON metadata and splits text at level-two headings into seven chunks. Documents and passages have SHA-256 digests. Chunk IDs identify document, version, and section so the original can be inspected.

Connect ingestion and answering

Ingestion parses documents, defines boundaries, stores metadata, generates document vectors, and saves the index. Answering filters by trusted identity, generates a query vector, retrieves candidates, fuses rankings, selects context, and sends evidence to the model. After generation, check shape, citations, and actual support. Retain each stage's results to locate the first evidence loss.

Know which vector work actually ran

The author ran FastEmbed 0.8.1 with BAAI/bge-small-en-v1.5 to generate 384-dimensional document and query vectors. The package includes those vectors, model-file digests, versions, and corpus digests. By default, the program reads the snapshot and computes cosine similarity, BM25, and RRF locally. Rebuilding downloads the model and runs inference on your computer.

The model and preprocessing serve this lab's English documents. Chinese or mixed-language corpora require an appropriate model, rebuilt document/query vectors, and evaluation in those languages. Equal dimensions do not establish compatible vector spaces.

Evaluate similarity and answerability separately

A customs-tax question may retrieve return-policy passages even though the corpus contains no tax evidence. Vector distance measures similarity within the chosen representation. Generation should allow an insufficient-evidence response. Evaluate correct abstention separately from refusing supported questions. The lab uses authored candidates to demonstrate these checks; free-form semantic support requires separate review.

Run experiments and observe counterexamples

Practical parsing of fictional English Markdown, running SQL filtering, cosine, BM25, RRF, request assembly and limited contract evaluation. Document and fixed query vectors are derived from real local Embedding inference snapshots; generated candidates are written by the authors and free text semantics are not scored.

Python 3.10+ · Runs by default using only the standard library · Runs on your computer

  1. Download the Agent application entry experimental package on this page, unzip it and enter the agent-application-lab-v1 directory.
  2. Use Python 3.10+ to execute the above command; the default playback only requires the standard library and package data.
  3. Compare the output with the checkpoint, then run python3 -m unittest test_application -v and complete the current layer task.
Download the complete application experiment package (including data and dependent scripts) ↓
python3 rag_pipeline.py
View the entry-point script
"""Actual parsing, SQL filtering, vector search, BM25, RRF and contract evaluation.

Embeddings are replayed from an actual local-model inference snapshot. The generation
stage uses explicitly authored fixtures. No network or inference is needed for replay.
"""
from collections import Counter
import json
import math
import re
import sqlite3

from application_data import FIXTURES, corpus_hash, load_chunks, read_cases
from prompt_iteration import CANDIDATES, build_prompt, grade


def load_snapshot(chunks, path=None):
    snapshot = json.loads((path or FIXTURES / "embeddings.json").read_text(encoding="utf-8"))
    if snapshot.get("schemaVersion") != 1 or snapshot.get("corpusSha256") != corpus_hash(chunks):
        raise ValueError("embedding snapshot is stale; rebuild after changing documents")
    if set(snapshot["documents"]) != {chunk["id"] for chunk in chunks}:
        raise ValueError("embedding snapshot has different chunk identities")
    dimension = snapshot["dimension"]
    if type(dimension) is not int or dimension < 1 or not snapshot.get("model"):
        raise ValueError("invalid embedding contract")
    for vector in [*snapshot["documents"].values(), *snapshot["queries"].values()]:
        if (len(vector) != dimension or not all(type(x) in (int, float) and math.isfinite(x) for x in vector)
                or not any(vector)):
            raise ValueError("invalid embedding vector")
    return snapshot


def tokenize(text):
    return re.findall(r"[a-z0-9]+", text.lower())


def cosine(left, right):
    return sum(a * b for a, b in zip(left, right)) / (math.sqrt(sum(a * a for a in left)) * math.sqrt(sum(b * b for b in right)))


def keyword_scores(query, chunks):
    terms = set(tokenize(query))
    bags = [Counter(tokenize(chunk["text"])) for chunk in chunks]
    average = sum(sum(bag.values()) for bag in bags) / len(bags)
    scores = []
    for chunk, bag in zip(chunks, bags):
        score = 0.0
        for term in terms:
            frequency = bag[term]
            if frequency:
                df = sum(term in other for other in bags)
                idf = math.log(1 + (len(bags) - df + 0.5) / (df + 0.5))
                score += idf * frequency * 2.2 / (frequency + 1.2 * (0.25 + 0.75 * sum(bag.values()) / average))
        scores.append((chunk["id"], score))
    return sorted((row for row in scores if row[1] > 0), key=lambda row: (-row[1], row[0]))


def index_chunks(chunks):
    db = sqlite3.connect(":memory:")
    db.execute("CREATE TABLE chunks(id TEXT PRIMARY KEY, tenant TEXT NOT NULL, current INTEGER NOT NULL, body TEXT NOT NULL)")
    db.executemany("INSERT INTO chunks VALUES(?,?,?,?)", [(chunk["id"], chunk["tenant"], int(chunk["current"]), json.dumps(chunk)) for chunk in chunks])
    return db


def retrieve(query, chunks, snapshot, tenant="shop-a", top_k=3):
    if type(top_k) is not int or top_k < 1:
        raise ValueError("top_k must be a positive integer")
    if query not in snapshot["queries"]:
        raise ValueError("query was not recorded; regenerate embeddings for the changed query")
    db = index_chunks(chunks)
    try:
        allowed = [json.loads(row[0]) for row in db.execute("SELECT body FROM chunks WHERE tenant=? AND current=1 ORDER BY id", (tenant,))]
    finally:
        db.close()
    if not allowed:
        return {"allowed": 0, "dense": [], "keyword": [], "context": []}
    query_vector = snapshot["queries"][query]
    dense = sorted([(chunk["id"], cosine(query_vector, snapshot["documents"][chunk["id"]])) for chunk in allowed], key=lambda row: (-row[1], row[0]))
    keyword = keyword_scores(query, allowed)
    fused = Counter()
    for ranking in (dense, keyword):
        for rank, (chunk_id, _) in enumerate(ranking, 1):
            fused[chunk_id] += 1 / (60 + rank)
    selected = [cid for cid, _ in sorted(fused.items(), key=lambda row: (-row[1], row[0]))[:top_k]]
    by_id = {chunk["id"]: chunk for chunk in allowed}
    return {"allowed": len(allowed), "dense": [cid for cid, _ in dense], "keyword": [cid for cid, _ in keyword], "context": [by_id[cid] for cid in selected]}


def run_pipeline(top_k=3):
    chunks = load_chunks()
    snapshot = load_snapshot(chunks)
    rows = []
    for case in read_cases():
        retrieved = retrieve(case["query"], chunks, snapshot, top_k=top_k)
        request = build_prompt(case, retrieved["context"])
        answer = CANDIDATES["v2"][case["id"]]
        context_ids = [chunk["id"] for chunk in retrieved["context"]]
        required = set(case["requiredEvidence"])
        rows.append({"case": case["id"], "contextIds": context_ids,
                     "request": request, "candidate": answer,
                     "evidenceRecall": len(required.intersection(context_ids)) / len(required) if required else None,
                     **grade(case, answer, retrieved["context"])})
    return {"chunks": chunks, "snapshot": snapshot, "rows": rows}


def demo():
    report = run_pipeline()
    answerable = [row["evidenceRecall"] for row in report["rows"] if row["evidenceRecall"] is not None]
    return {"chunks": len(report["chunks"]), "dimension": report["snapshot"]["dimension"],
            "embeddingModel": report["snapshot"]["model"], "queries": len(report["rows"]),
            "contextIds": {row["case"]: row["contextIds"] for row in report["rows"]},
            "contractPasses": sum(row["contractPassed"] for row in report["rows"]),
            "meanNecessaryEvidenceRecallAt3": sum(answerable) / len(answerable),
            "semanticQuality": "not_scored", "generation": "authored_fixtures",
            "embedding": "real_local_inference_snapshot"}


if __name__ == "__main__":
    print(json.dumps(demo(), ensure_ascii=False, sort_keys=True))

Expected output when running locally

{"chunks": 7, "contextIds": {"battery": ["returns:v2:battery", "returns:v2:general", "warranty:v1:defects"], "ordinary": ["returns:v2:general", "returns:v2:battery", "shipping:v1:contact"], "unknown": ["returns:v2:general", "returns:v2:battery", "warranty:v1:defects"]}, "contractPasses": 3, "dimension": 384, "embedding": "real_local_inference_snapshot", "embeddingModel": "BAAI/bge-small-en-v1.5", "generation": "authored_fixtures", "meanNecessaryEvidenceRecallAt3": 1.0, "queries": 3, "semanticQuality": "not_scored"}
  • chunks=7, dimension=384. See embeddings.json for the model and artifact digests.
  • Three queries are actually run and ranked, with current tenant and version filtering occurring at the top of the ranking.
  • The meanNecessaryEvidenceRecallAt3=1.0 of the fixed corpus is calculated, and only two answerable questions are counted.
  • generation=authored_fixtures、semanticQuality=not_scored。
View the running environment, output and verification records →

Acceptance task for this level

Draw the two processes of warehousing and answering, mark the tenant, version, vector, candidate, final context and reference, and list the necessary evidence collection for the battery problem.

Check each item after completion

  • Mappings of seven chunks to five originals are retained.
  • Battery Issues lists general rules and battery exceptions.
  • Permissions and versions are checked before the candidate enters the context.
  • Description snapshots are from actual Embedding inference, and generated candidates are written by the authors.

Save your own processes, code and results. Acceptance requirements are provided here, and course mastery status will not be automatically graded or saved at this time.

Hide the answer and check your understanding

Both the old rules and the new rules mention returns, why can't both versions be given to the model to choose?

Further explanations and practice

When encountering unfamiliar principles, first read the implementation, continuous questioning and migration cases, and then independently explain the premise and boundaries. Answers and notes are saved to the original account record.

All linked explanations and exercises (5 )

Sources and verification scope

The principles are based on public information; the numbers, cases and tasks are the teaching design of this website. Offline experiments verify the range noted on this page, and the learning effect still needs to be judged through independent tasks and feedback.