Understand first
New to this knowledge point
Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.
Start with core principles →Understand → Implement → Debug → Design
Examine hybrid retrieval, rank fusion, deduplication and multi-stage evaluation.
Knowledge content check2026-10-03 · Check the source of the original question2026-10-02
It is recommended to understand first:
Prompts, structured output, and iterative acceptance checks →Context and generation budgets →Tool calls: structure, authorization, and business contracts →Select the starting point based on the current basis, or you can go deeper one by one. When you encounter an unfamiliar concept, go back to the core principles first; use the knowledge exercises to check your understanding when you are finished.
New to this knowledge point
Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.
Start with core principles →Prepare to write the principles into code
Understand implementation steps and boundaries, complete small tasks, and check results against acceptance requirements.
View the code example →Need to handle failures and changes in conditions
Follow the continuous questioning to locate the failure premise, and then compare the migration cases to explain how the plan should be adjusted.
Continue to delve deeper into the problem →Need to design or review plans
Combine engineering deductions and senior self-evaluation standards to explain the applicable conditions, costs and alternatives of the plan.
Analyze engineering scenarios →LEARN · PRACTICE · REFLECT
First read along the principles, Q&A and migration cases. When you need to check your understanding, switch to reinforcement exercises or start personal recording.
Can be practiced directly. After logging in, answers, favorites, and notes will be saved to your account.
Log in and saveEach modified commit will be kept as an independent history. Your level of mastery is up to you to evaluate yourself against the standards.
Core concept · Complementary retrieval and rank fusion in hybrid search
Preparatory concepts:Inverted index and BM25, vector similarity, Candidates and rearrangements
Lexical retrieval and semantic retrieval address different omissions, and fusion allows complementary evidence to enter candidates. reranking can only rearrange existing evidence, and no fusion algorithm can retrieve precise entities that have never been recalled.
XQ-104 and XQ-140 look similar but identify different products. “Pause renewal” and “stop charges after cancellation” may express similar needs. Lexical fields preserve identifiers; vectors help with varied phrasing. Verify tokenization and normalization: BM25 alone does not guarantee meaningful treatment of dashes or case.
Each channel retrieves under the same authorization and version constraints, identifying shared chunks by stable IDs. RRF adds rank contributions: score(d) = Σ 1 / (c + rank_i(d)). Sum only lists containing d, with ranks starting at 1 and fusion constant c. Raw score scales need not be directly comparable.
For c=60, let A rank first and third, B second and first, and C second only in the second list:
| Evidence | Contributions | Approximate score |
|---|---|---|
| A | 1/61 + 1/63 | 0.03227 |
| B | 1/62 + 1/61 | 0.03252 |
| C | 0 + 1/62 | 0.01613 |
The fused order is B, A, C; applicability still needs verification. The constant c differs from vector candidate count k, the fusion window, and final result size. In Elastic, k beyond the fusion window is truncated. Increasing final size cannot recover evidence missing from every candidate list.
A reranker cannot recreate an identifier removed during rewriting or evidence missed during retrieval. Rich old descriptions may outrank current facts, so retain version, entity, and access constraints. Deduplicate adjacent chunks or aggregate by source without blending incompatible versions.
Use exact IDs, paraphrases, typos, unanswerable questions, and combined conditions. Disable each retrieval channel in turn for comparison. Measure evidence coverage, reranking, answer support, and latency. RRF scores depend on the participating list count and cannot serve as an uncalibrated factual threshold. Query distribution determines useful fusion.
Use lexical or exact retrieval for product IDs and names, and vectors for varied semantic expressions. Retrieve both within authorization scope, deduplicate stable chunk IDs, fuse rankings with methods such as RRF, then rerank within budget. Do not arbitrarily add BM25 and vector scores. Measure candidate coverage separately from final ranking.
"ZX-104 failure" requires the exact model, and approximate semantics may recall ZX-105; "Unable to start" may need to match "boot failure". First normalize the numbers, times, and entities in the query, preserving the original question. Both the lexical and vector paths use current permissions and document validity filtering. You cannot first hand over unauthorized text to the reranking model and then filter and display it.
Ranking fusion can be used to establish a baseline: contribute 1/(k+r) to the r-th document in each path and merge by unified ID. k and each candidate window affect the results and need to be calibrated on the data set. Score fusion is also possible, but score scales, distributions, and weights should be dealt with, and 12 points from BM25 cannot be simply added to 0.8 points from the vector and considered fair.
The reranker can only choose among input candidates. First check whether the candidate set contains necessary evidence, and then analyze whether the sorting ranks it outside the context window. Too many adjacent duplicate blocks will fill up the quota. Duplication and diversity control can be done by document or topic. There must be a clear degradation plan when the reranking service times out, such as using fused ranking, and recording that reranking was skipped for this request.
Prepare separate exact numbering, synonymous paraphrasing, long questions and unanswered questions. Check Recall@K, nDCG or human relevance judgments, final reference correctness and P95 latency. Perform small-scale ablation on the number of recall and reranking candidates, and report which types of problems the improvement comes from. RRF is a measurable engineering baseline, not an algorithm that can guarantee optimality without data.
Demonstrates ranking fusion; no retrievers, permission filtering, or parameter tuning are implemented.
def rrf(rankings, k=60):
scores = {}
for ranking in rankings:
unique = dict.fromkeys(ranking)
for rank, doc in enumerate(unique, 1):
scores[doc] = scores.get(doc, 0) + 1 / (k + rank)
return sorted(scores, key=lambda doc: (-scores[doc], doc))
print(rrf([["A", "B", "C"], ["B", "D", "A"]]))
expected output
['B', 'A', 'D', 'C']Continue reading along with the premises and constraints of the problem. Understand the reference answers first, then try to put away the answers and explain the cause and effect and trade-offs in your own words.
Level 1How does RRF handle the same block being recalled by two ways?
The identity of the same evidence is further implemented from the two complementary paths.
Use the same stable chunk_id to aggregate the two-way results and accumulate its ranking contribution in different lists; repeated votes in the same list cannot be counted repeatedly. The ranking of each channel is retained to facilitate explanation and debugging. Different document versions with the same text cannot be merged solely by text hashing, otherwise the old version and the new version may be mistakenly regarded as supporting each other.
Follow this answer further
Level 2If the two paths hit different blocks of the same text, should they be treated as the same item?
After parents are merged by ID, duplicate text in real data may amplify the sorting contribution.
First, distinguish between evidence identity and text duplication. Overlapping blocks within the same source version can be aggregated or deduplicated, but different terms, versions, and permissions cannot be merged based solely on the same text. Keep the original block positioning, and select the evidence according to the diversity of sources after reranking; otherwise, ten copied documents will appear to form a consensus of ten votes.
Follow this answer further
Level 3After deduplication, the target falls out of the results. How to judge whether the deduplication is wrong or the fusion parameters are inappropriate?
After introducing deduplication to change the ranking, it is necessary to isolate semantic loss and parameter effects.
Compare the source location and necessary evidence content before and after deduplication; if the exception clause disappears with deduplication, revise the identity rules first. If the evidence is still there but the ranking is lowered, adjust the candidate window or fusion constant in the annotation set. The two changes are made separately to avoid parameter tuning covering up incorrect merging; finally verify evidence support for the answer, rather than ranking alone.
Level 1Why do we still get wrong answers when the rearranged model is stronger?
Distinguish between candidate coverage and ranking capabilities to avoid misattribution.
reranking only optimizes the order of given candidates. Candidates cannot be sorted without correct evidence; when the text is missing units, versions or numbers, it may also rank related but wrong blocks first. Candidate coverage and hard constraints are checked first before rearrangements are evaluated. A stronger model may also smoothly complete missing information, and fluency cannot be regarded as support.
Level 1What should I do if the number is deleted by the query rewriter?
Recall improvements rely on upstream intent not being broken by overwriting.
The user's original sentence is retained as an independent search branch, and the extracted number is a protected entity slot. The rewritten result must pass the retention check; the rewrite is rejected or rolled back when the number conflicts or is omitted. Don't let semantic expansion replace precise qualification. Only by recording the original sentence, rewriting and hits in each way can we find the change in the recall direction before and after rewriting.
First find out the conditions for change, and then determine which premises in the original plan still hold true. The following cases are teaching deductions to facilitate the transfer of principles to new problems.
Changing conditions:From number query to natural language synonyms
Extended question:If the keyword is not hit all the way, should we still increase its weight?
It is not possible to give uniform weight to single-category questions. Preserve the semantic branch and evaluate whether it recalls the target evidence; then preserve the semantic hits through fusion and reranking. Look at the changes hierarchically by question. Accurate identification of questions and synonyms may require different query plans. Weighting is a choice based on data, and one path should not be considered permanently more reliable.
The principles that remain unchanged:Different retrieval signals complement each other, and the final judgment is based on whether the target evidence is covered rather than the high score of a single path.
Changing conditions:The identity is the same, but the applicable version is different
Extended question:Is it stronger evidence that both paths hit the old version?
It's just that both search methods consider it relevant. Filter by user-specified version first, clarify or explicitly list version differences when not specified and would change the answer. Incorporate document_version into the evidence identity and do not support cumulative support across versions; candidate consensus cannot cover business applicability.
The principles that remain unchanged:Ranking consistency does not equal factual applicability, version is the condition for the evidence to be established.
It is designed based on public technical information; the reference materials support the technical mechanism, and the scenarios and scoring standards are designed by this website and do not represent the original interview questions of a certain company. New Q&A and migration cases are added for principle explanation, and source verification and case operation verification are recorded separately.
After reading, you can explain the principles, boundaries, and trade-offs against these standards. It is up to you to evaluate your mastery; if further verification is needed, complete the small tasks below.
View verification records for independent examples
Transferring pre-retrieval filtered ideas to memory versions and recovery. Observe how existing drafts become invalid after cancellation.
Read full text and fault analysis → · Download Reliability Experiment v3 ↓
python3 cli.py memory-put --db memory.sqlite
python3 cli.py submit --db memory.sqlite
python3 cli.py run --db memory.sqlite --lease-seconds 2 --fault after_draft
python3 cli.py memory-forget --db memory.sqlite
# 等待至少 2 秒后分别执行
python3 cli.py run --db memory.sqlite
python3 cli.py inspect --db memory.sqliteVerify local scope, version, and undo blocking; old checkpoints remain, no complete deletion of logs, backups, or checkpoints is provided.
Hand-calculated RRF for two-way sorting A,B,C and B,D,A, illustrating how duplicate documents are merged.