Agent Application DevelopmentAccount
Knowledge catalogChoose core direction and segmented content
knowledge unit 20AdvancedImplementationAbout 18 minutes

Understand → Implement → Debug → Design

Complementary retrieval and rank fusion in hybrid search

Examine hybrid retrieval, rank fusion, deduplication and multi-stage evaluation.

BM25Hybrid SearchRRF

Knowledge content check2026-10-03 · Check the source of the original question2026-10-02

Which step do you want to learn from this knowledge point?

Select the starting point based on the current basis, or you can go deeper one by one. When you encounter an unfamiliar concept, go back to the core principles first; use the knowledge exercises to check your understanding when you are finished.

Understand first

New to this knowledge point

Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.

Start with core principles →

Implement next

Prepare to write the principles into code

Understand implementation steps and boundaries, complete small tasks, and check results against acceptance requirements.

View the code example →

Debug failures

Need to handle failures and changes in conditions

Follow the continuous questioning to locate the failure premise, and then compare the migration cases to explain how the plan should be adjusted.

Continue to delve deeper into the problem →

Compare designs

Need to design or review plans

Combine engineering deductions and senior self-evaluation standards to explain the applicable conditions, costs and alternatives of the plan.

Analyze engineering scenarios →
Knowledge unit directory

LEARN · PRACTICE · REFLECT

Knowledge learning and personal records

My notes and review ↗

First read along the principles, Q&A and migration cases. When you need to check your understanding, switch to reinforcement exercises or start personal recording.

Answers and personal notes

Each modified commit will be kept as an independent history. Your level of mastery is up to you to evaluate yourself against the standards.

Core concept · Complementary retrieval and rank fusion in hybrid search

Understand the core principles first

Preparatory concepts:Inverted index and BM25, vector similarity, Candidates and rearrangements

Lexical retrieval and semantic retrieval address different omissions, and fusion allows complementary evidence to enter candidates. reranking can only rearrange existing evidence, and no fusion algorithm can retrieve precise entities that have never been recalled.

Exact identifiers and semantic similarity provide different signals

XQ-104 and XQ-140 look similar but identify different products. “Pause renewal” and “stop charges after cancellation” may express similar needs. Lexical fields preserve identifiers; vectors help with varied phrasing. Verify tokenization and normalization: BM25 alone does not guarantee meaningful treatment of dashes or case.

Fuse rankings within one allowed scope

Each channel retrieves under the same authorization and version constraints, identifying shared chunks by stable IDs. RRF adds rank contributions: score(d) = Σ 1 / (c + rank_i(d)). Sum only lists containing d, with ranks starting at 1 and fusion constant c. Raw score scales need not be directly comparable.

For c=60, let A rank first and third, B second and first, and C second only in the second list:

Evidence Contributions Approximate score
A 1/61 + 1/63 0.03227
B 1/62 + 1/61 0.03252
C 0 + 1/62 0.01613

The fused order is B, A, C; applicability still needs verification. The constant c differs from vector candidate count k, the fusion window, and final result size. In Elastic, k beyond the fusion window is truncated. Increasing final size cannot recover evidence missing from every candidate list.

Reranking cannot repair every upstream loss

A reranker cannot recreate an identifier removed during rewriting or evidence missed during retrieval. Rich old descriptions may outrank current facts, so retain version, entity, and access constraints. Deduplicate adjacent chunks or aggregate by source without blending incompatible versions.

Measure whether fusion helps

Use exact IDs, paraphrases, typos, unanswerable questions, and combined conditions. Disable each retrieval channel in turn for comparison. Measure evidence coverage, reranking, answer support, and latency. RRF scores depend on the participating list count and cannot serve as an uncalibrated factual threshold. Query distribution determines useful fusion.

Check understanding with a question

Vector search always misses product numbers. How to design BM25, vector recall and reranking?

Use lexical or exact retrieval for product IDs and names, and vectors for varied semantic expressions. Retrieve both within authorization scope, deduplicate stable chunk IDs, fuse rankings with methods such as RRF, then rerank within budget. Do not arbitrarily add BM25 and vector scores. Measure candidate coverage separately from final ranking.

Implementation and trade-offs

Distinguish retrieval signals

"ZX-104 failure" requires the exact model, and approximate semantics may recall ZX-105; "Unable to start" may need to match "boot failure". First normalize the numbers, times, and entities in the query, preserving the original question. Both the lexical and vector paths use current permissions and document validity filtering. You cannot first hand over unauthorized text to the reranking model and then filter and display it.

Rank fusion differs from adding raw scores

Ranking fusion can be used to establish a baseline: contribute 1/(k+r) to the r-th document in each path and merge by unified ID. k and each candidate window affect the results and need to be calibrated on the data set. Score fusion is also possible, but score scales, distributions, and weights should be dealt with, and 12 points from BM25 cannot be simply added to 0.8 points from the vector and considered fair.

Reranking cannot recover missing candidates

The reranker can only choose among input candidates. First check whether the candidate set contains necessary evidence, and then analyze whether the sorting ranks it outside the context window. Too many adjacent duplicate blocks will fill up the quota. Duplication and diversity control can be done by document or topic. There must be a clear degradation plan when the reranking service times out, such as using fused ranking, and recording that reranking was skipped for this request.

Evaluation and Choice

Prepare separate exact numbering, synonymous paraphrasing, long questions and unanswered questions. Check Recall@K, nDCG or human relevance judgments, final reference correctness and P95 latency. Perform small-scale ablation on the number of recall and reranking candidates, and report which types of problems the improvement comes from. RRF is a measurable engineering baseline, not an algorithm that can guarantee optimality without data.

code example

RRF ranking fusion and deduplication

Demonstrates ranking fusion; no retrievers, permission filtering, or parameter tuning are implemented.

def rrf(rankings, k=60):
    scores = {}
    for ranking in rankings:
        unique = dict.fromkeys(ranking)
        for rank, doc in enumerate(unique, 1):
            scores[doc] = scores.get(doc, 0) + 1 / (k + rank)
    return sorted(scores, key=lambda doc: (-scores[doc], doc))

print(rrf([["A", "B", "C"], ["B", "D", "A"]]))

expected output

['B', 'A', 'D', 'C']

Engineering deduction

scene
Interview hypothesis: The equipment knowledge base mismatched the ZX-104 troubleshooting questions to similar models.
design decisions
Add precise number matching and lexical recall, and then integrate semantic search results.
Verify target
Model evidence coverage is improved, and reorder timeouts have traceable downgrades.
applicable boundary
Validate retrieval parameters and improvement claims with a local task set.

Continuous questions and answers

Continue reading along with the premises and constraints of the problem. Understand the reference answers first, then try to put away the answers and explain the cause and effect and trade-offs in your own words.

Draw inferences from one example: If the conditions change, how to deduce it?

First find out the conditions for change, and then determine which premises in the original plan still hold true. The following cases are teaching deductions to facilitate the transfer of principles to new problems.

User expressions have no keyword overlap

Changing conditions:From number query to natural language synonyms

Extended question:If the keyword is not hit all the way, should we still increase its weight?

Derivation and reference solutions

It is not possible to give uniform weight to single-category questions. Preserve the semantic branch and evaluate whether it recalls the target evidence; then preserve the semantic hits through fusion and reranking. Look at the changes hierarchically by question. Accurate identification of questions and synonyms may require different query plans. Weighting is a choice based on data, and one path should not be considered permanently more reliable.

The principles that remain unchanged:Different retrieval signals complement each other, and the final judgment is based on whether the target evidence is covered rather than the high score of a single path.

Product numbers appear in multi-version manuals

Changing conditions:The identity is the same, but the applicable version is different

Extended question:Is it stronger evidence that both paths hit the old version?

Derivation and reference solutions

It's just that both search methods consider it relevant. Filter by user-specified version first, clarify or explicitly list version differences when not specified and would change the answer. Incorporate document_version into the evidence identity and do not support cumulative support across versions; candidate consensus cannot cover business applicability.

The principles that remain unchanged:Ranking consistency does not equal factual applicability, version is the condition for the evidence to be established.

Easy to make mistakes

  • Direct addition of fractions with different dimensions
  • No permission filtering is performed before reranking
  • Only final answers are tested, candidate coverage is not tested

References

It is designed based on public technical information; the reference materials support the technical mechanism, and the scenarios and scoring standards are designed by this website and do not represent the original interview questions of a certain company. New Q&A and migration cases are added for principle explanation, and source verification and case operation verification are recorded separately.

Check how far you understand

After reading, you can explain the principles, boundaries, and trade-offs against these standards. It is up to you to evaluate your mastery; if further verification is needed, complete the small tasks below.

Basic standards met
Can design lexical and vector two-way recall, and remove duplicates based on stable document ID.
Intermediate and advanced signals
Explain rank fusion, candidate windows, and reranking constraints.
Senior criteria
Can use layered evaluation to find sources of improvement and design timeout degradation.

View verification records for independent examples

Continue to do advanced research experiments

Why can’t the old context continue to be used after the policy changes?

Transferring pre-retrieval filtered ideas to memory versions and recovery. Observe how existing drafts become invalid after cancellation.

Read full text and fault analysis → · Download Reliability Experiment v3 ↓

python3 cli.py memory-put --db memory.sqlite
python3 cli.py submit --db memory.sqlite
python3 cli.py run --db memory.sqlite --lease-seconds 2 --fault after_draft
python3 cli.py memory-forget --db memory.sqlite
# 等待至少 2 秒后分别执行
python3 cli.py run --db memory.sqlite
python3 cli.py inspect --db memory.sqlite

Keep evidence and check item by item

  • The draft checkpoint already exists after the first exit.
  • Recovery after memory-forget gets failed with memory_changed_or_expired.
  • No publish checkpoint; explain the difference between fail blocking and complete deletion.

Verify local scope, version, and undo blocking; old checkpoints remain, no complete deletion of logs, backups, or checkpoints is provided.

Hands-on verificationComplete on demand · Suggestions15 minutes

Hand-calculated RRF for two-way sorting A,B,C and B,D,A, illustrating how duplicate documents are merged.

Expand acceptance requirements and checkpoints
  • Merge scores for the same document
  • Documents that are not candidates cannot be rearranged and filled in.
  • Rankings are counted from a unified starting point

Key inspections

  • Know that lexical and vector signals are complementary
  • Can explain fusion and reranking of boundaries
  • Evaluate candidate coverage individually