Agent Application DevelopmentAccount
Knowledge catalogChoose core direction and segmented content

SYSTEMATIC LEARNING / FOUR-LEVEL COURSE

From document ingestion to cited RAG answers

Run Markdown parsing, replay recorded embedding vectors, apply SQL permission filters, and inspect BM25, vector retrieval, RRF, and answer acceptance checks.

Learning objectives: Inspect each intermediate RAG artifact, compare how top-k affects required evidence, and distinguish retrieval, generation, and semantic evaluation results.

Content checked: 2026-10-04 · Each level has independent explanations, tasks and inspections

Choose a starting point based on your familiarity with this topic. Current level: Design · Explain the trade-offs. After completing the task, continue to the next level. Reading and self-checks alone do not establish mastery.

On this level

Review the prerequisites

Suitable for: Document versions, permissions, indexing and publishing reviews need to be maintained.

Warehouse
Parse documents, determine tiles and metadata, and record versions and summaries for content and sources.
Embedding
Text is mapped to vectors by the specified model; documents and queries need to use compatible models and processing.
Retrieval candidates
Vector or keyword ranking provides potentially relevant blocks, filtered by current permissions and versions.
RRF
The results are fused according to the ranking of each ranking; the fusion constant in this experiment is 60, and the final number of contexts is controlled by top_k.
Necessary evidence recall
How many of the required human-annotated sources enter the context, calculated per answerable question.
vector snapshot
The real local model inference artifacts of this batch of texts, with models, dimensions, corpus summaries and model file summaries, support offline playback without model download.

How does the mechanism work?

  1. Parse and preserve sources

    Five copies of fictional English Markdown generate seven independent blocks, retaining tenants, versions, chapters, body, and abstracts.

  2. Generate compatibility vector

    The authors actually ran bge-small-en-v1.5, recording 384-dimensional vectors of seven document blocks and three queries.

  3. Filter first then rank

    SQLite parameterizes filtering of the current tenant and valid versions, and then performs keyword and cosine searches.

  4. Fusion and assembly

    BM25 is fused with vector ranking using RRF, intercepting the final context and writing the request with the source ID.

  5. Acceptance Candidate

    Author candidates are checked for citation membership, permissions, version and limited policy tags, and free text semantics are individually marked.

Design · Explain the trade-offs

Scale from seven chunks to a maintainable knowledge base

Objectives of this level: Design document lifecycles and comparable retrieval strategies, and illustrate the scope of sample metric support.

Version documents, vectors, and permissions separately

Document IDs identify stable sources; content versions and chunk IDs identify particular text. Embedding models and preprocessing define vector spaces. Permissions and revocations determine current access. Build artifacts for a target index, verify samples, then switch queries. On deletion or revocation, inspect candidate caches, vectors, context, and resumed tasks too.

This lab conservatively binds the whole corpus digest, requiring rebuilds even for metadata changes. Production can separate text vectors from authorization versions to avoid unnecessary inference, while checking current authorization at answer time. Index migration needs an old-version recovery strategy and query vectors compatible with the target index.

Add comparable stages individually

Establish BM25, exact-vector, and RRF baselines before comparing query rewriting, adjacent-chunk expansion, and model reranking. Fix tasks and labels and retain candidates and final context. Compare required evidence, noise, correct abstention, latency, and actual generation usage. Revalidate filtering and recall semantics when adopting ANN or a managed index.

Interpret the lab numbers within their scope

The two answerable fixtures have mean required-evidence recall of 1.0 at top_k=3. A third checks abstention, and three authored candidates pass a limited contract. These numbers describe this fixed teaching corpus. Real evaluation needs independent labels covering document sources, item exceptions, languages, permissions, versions, and time boundaries, plus live generation and semantic review.

Retain evidence for large results and long tasks

Bind source IDs, versions, selected chunks, actual requests, and candidates to a run. Record truncation and summarization when context is tight and keep exceptions traceable. Recheck current authorization and document state after recovery; an old checkpoint summary is not automatically approved evidence today. Research-report and memory-revocation labs explore this boundary further.

State the engineering scope precisely

The teaching parser supports only the supplied Markdown format. Documents are English, the index is in-memory SQLite, and generation uses authored candidates. Local embedding inference is recorded. Live generation, PDF/OCR, online ACLs, distributed indexes, semantic grading, and production load each need separate verification artifacts.

Run experiments and observe counterexamples

Practical parsing of fictional English Markdown, running SQL filtering, cosine, BM25, RRF, request assembly and limited contract evaluation. Document and fixed query vectors are derived from real local Embedding inference snapshots; generated candidates are written by the authors and free text semantics are not scored.

Python 3.10+ · Runs by default using only the standard library · Runs on your computer

  1. Download the Agent application entry experimental package on this page, unzip it and enter the agent-application-lab-v1 directory.
  2. Use Python 3.10+ to execute the above command; the default playback only requires the standard library and package data.
  3. Compare the output with the checkpoint, then run python3 -m unittest test_application -v and complete the current layer task.
Download the complete application experiment package (including data and dependent scripts) ↓
python3 rag_pipeline.py
View the entry-point script
"""Actual parsing, SQL filtering, vector search, BM25, RRF and contract evaluation.

Embeddings are replayed from an actual local-model inference snapshot. The generation
stage uses explicitly authored fixtures. No network or inference is needed for replay.
"""
from collections import Counter
import json
import math
import re
import sqlite3

from application_data import FIXTURES, corpus_hash, load_chunks, read_cases
from prompt_iteration import CANDIDATES, build_prompt, grade


def load_snapshot(chunks, path=None):
    snapshot = json.loads((path or FIXTURES / "embeddings.json").read_text(encoding="utf-8"))
    if snapshot.get("schemaVersion") != 1 or snapshot.get("corpusSha256") != corpus_hash(chunks):
        raise ValueError("embedding snapshot is stale; rebuild after changing documents")
    if set(snapshot["documents"]) != {chunk["id"] for chunk in chunks}:
        raise ValueError("embedding snapshot has different chunk identities")
    dimension = snapshot["dimension"]
    if type(dimension) is not int or dimension < 1 or not snapshot.get("model"):
        raise ValueError("invalid embedding contract")
    for vector in [*snapshot["documents"].values(), *snapshot["queries"].values()]:
        if (len(vector) != dimension or not all(type(x) in (int, float) and math.isfinite(x) for x in vector)
                or not any(vector)):
            raise ValueError("invalid embedding vector")
    return snapshot


def tokenize(text):
    return re.findall(r"[a-z0-9]+", text.lower())


def cosine(left, right):
    return sum(a * b for a, b in zip(left, right)) / (math.sqrt(sum(a * a for a in left)) * math.sqrt(sum(b * b for b in right)))


def keyword_scores(query, chunks):
    terms = set(tokenize(query))
    bags = [Counter(tokenize(chunk["text"])) for chunk in chunks]
    average = sum(sum(bag.values()) for bag in bags) / len(bags)
    scores = []
    for chunk, bag in zip(chunks, bags):
        score = 0.0
        for term in terms:
            frequency = bag[term]
            if frequency:
                df = sum(term in other for other in bags)
                idf = math.log(1 + (len(bags) - df + 0.5) / (df + 0.5))
                score += idf * frequency * 2.2 / (frequency + 1.2 * (0.25 + 0.75 * sum(bag.values()) / average))
        scores.append((chunk["id"], score))
    return sorted((row for row in scores if row[1] > 0), key=lambda row: (-row[1], row[0]))


def index_chunks(chunks):
    db = sqlite3.connect(":memory:")
    db.execute("CREATE TABLE chunks(id TEXT PRIMARY KEY, tenant TEXT NOT NULL, current INTEGER NOT NULL, body TEXT NOT NULL)")
    db.executemany("INSERT INTO chunks VALUES(?,?,?,?)", [(chunk["id"], chunk["tenant"], int(chunk["current"]), json.dumps(chunk)) for chunk in chunks])
    return db


def retrieve(query, chunks, snapshot, tenant="shop-a", top_k=3):
    if type(top_k) is not int or top_k < 1:
        raise ValueError("top_k must be a positive integer")
    if query not in snapshot["queries"]:
        raise ValueError("query was not recorded; regenerate embeddings for the changed query")
    db = index_chunks(chunks)
    try:
        allowed = [json.loads(row[0]) for row in db.execute("SELECT body FROM chunks WHERE tenant=? AND current=1 ORDER BY id", (tenant,))]
    finally:
        db.close()
    if not allowed:
        return {"allowed": 0, "dense": [], "keyword": [], "context": []}
    query_vector = snapshot["queries"][query]
    dense = sorted([(chunk["id"], cosine(query_vector, snapshot["documents"][chunk["id"]])) for chunk in allowed], key=lambda row: (-row[1], row[0]))
    keyword = keyword_scores(query, allowed)
    fused = Counter()
    for ranking in (dense, keyword):
        for rank, (chunk_id, _) in enumerate(ranking, 1):
            fused[chunk_id] += 1 / (60 + rank)
    selected = [cid for cid, _ in sorted(fused.items(), key=lambda row: (-row[1], row[0]))[:top_k]]
    by_id = {chunk["id"]: chunk for chunk in allowed}
    return {"allowed": len(allowed), "dense": [cid for cid, _ in dense], "keyword": [cid for cid, _ in keyword], "context": [by_id[cid] for cid in selected]}


def run_pipeline(top_k=3):
    chunks = load_chunks()
    snapshot = load_snapshot(chunks)
    rows = []
    for case in read_cases():
        retrieved = retrieve(case["query"], chunks, snapshot, top_k=top_k)
        request = build_prompt(case, retrieved["context"])
        answer = CANDIDATES["v2"][case["id"]]
        context_ids = [chunk["id"] for chunk in retrieved["context"]]
        required = set(case["requiredEvidence"])
        rows.append({"case": case["id"], "contextIds": context_ids,
                     "request": request, "candidate": answer,
                     "evidenceRecall": len(required.intersection(context_ids)) / len(required) if required else None,
                     **grade(case, answer, retrieved["context"])})
    return {"chunks": chunks, "snapshot": snapshot, "rows": rows}


def demo():
    report = run_pipeline()
    answerable = [row["evidenceRecall"] for row in report["rows"] if row["evidenceRecall"] is not None]
    return {"chunks": len(report["chunks"]), "dimension": report["snapshot"]["dimension"],
            "embeddingModel": report["snapshot"]["model"], "queries": len(report["rows"]),
            "contextIds": {row["case"]: row["contextIds"] for row in report["rows"]},
            "contractPasses": sum(row["contractPassed"] for row in report["rows"]),
            "meanNecessaryEvidenceRecallAt3": sum(answerable) / len(answerable),
            "semanticQuality": "not_scored", "generation": "authored_fixtures",
            "embedding": "real_local_inference_snapshot"}


if __name__ == "__main__":
    print(json.dumps(demo(), ensure_ascii=False, sort_keys=True))

Expected output when running locally

{"chunks": 7, "contextIds": {"battery": ["returns:v2:battery", "returns:v2:general", "warranty:v1:defects"], "ordinary": ["returns:v2:general", "returns:v2:battery", "shipping:v1:contact"], "unknown": ["returns:v2:general", "returns:v2:battery", "warranty:v1:defects"]}, "contractPasses": 3, "dimension": 384, "embedding": "real_local_inference_snapshot", "embeddingModel": "BAAI/bge-small-en-v1.5", "generation": "authored_fixtures", "meanNecessaryEvidenceRecallAt3": 1.0, "queries": 3, "semanticQuality": "not_scored"}
  • chunks=7, dimension=384. See embeddings.json for the model and artifact digests.
  • Three queries are actually run and ranked, with current tenant and version filtering occurring at the top of the ranking.
  • The meanNecessaryEvidenceRecallAt3=1.0 of the fixed corpus is calculated, and only two answerable questions are counted.
  • generation=authored_fixtures、semanticQuality=not_scored。
View the running environment, output and verification records →

Continue to do advanced research experiments

Why can’t the old context continue to be used after the policy changes?

Transferring pre-retrieval filtered ideas to memory versions and recovery. Observe how existing drafts become invalid after cancellation.

Read full text and fault analysis → · Download Reliability Experiment v3 ↓

python3 cli.py memory-put --db memory.sqlite
python3 cli.py submit --db memory.sqlite
python3 cli.py run --db memory.sqlite --lease-seconds 2 --fault after_draft
python3 cli.py memory-forget --db memory.sqlite
# 等待至少 2 秒后分别执行
python3 cli.py run --db memory.sqlite
python3 cli.py inspect --db memory.sqlite

Keep evidence and check item by item

  • The draft checkpoint already exists after the first exit.
  • Recovery after memory-forget gets failed with memory_changed_or_expired.
  • No publish checkpoint; explain the difference between fail blocking and complete deletion.

Verify local scope, version, and undo blocking; old checkpoints remain, no complete deletion of logs, backups, or checkpoints is provided.

Acceptance task for this level

Write an upgrade plan for the policy knowledge base, listing the acceptance of documents/vectors/authorized versions, index switching, search comparison, rejection, revocation and restoration.

Check each item after completion

  • The same dimensions and the same vector space are judged separately.
  • Permissions and content updates have their own invalid objects.
  • The old and new strategies compare the same set of independent tasks.
  • Report sample size, language, necessary evidence, and unscored items.
  • Deletion or deprivation overwrites cache, context and old task recovery.

Save your own processes, code and results. Acceptance requirements are provided here, and course mastery status will not be automatically graded or saved at this time.

Hide the answer and check your understanding

Both models output 384-dimensional vectors. Can new query vectors search the old index directly?

Further explanations and practice

When encountering unfamiliar principles, first read the implementation, continuous questioning and migration cases, and then independently explain the premise and boundaries. Answers and notes are saved to the original account record.

All linked explanations and exercises (5 )

Sources and verification scope

The principles are based on public information; the numbers, cases and tasks are the teaching design of this website. Offline experiments verify the range noted on this page, and the learning effect still needs to be judged through independent tasks and feedback.