Understand first
New to this knowledge point
Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.
Start with core principles →Understand → Implement → Debug → Design
Investigate vector space compatibility, dual indexing, incremental changes, and deletion consistency.
Knowledge content check2026-10-03 · Check the source of the original question2026-10-02
It is recommended to understand first:
Prompts, structured output, and iterative acceptance checks →Context and generation budgets →Tool calls: structure, authorization, and business contracts →Select the starting point based on the current basis, or you can go deeper one by one. When you encounter an unfamiliar concept, go back to the core principles first; use the knowledge exercises to check your understanding when you are finished.
New to this knowledge point
Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.
Start with core principles →Prepare to write the principles into code
Understand implementation steps and boundaries, complete small tasks, and check results against acceptance requirements.
Reading implementation and trade-offs →Need to handle failures and changes in conditions
Follow the continuous questioning to locate the failure premise, and then compare the migration cases to explain how the plan should be adjusted.
Continue to delve deeper into the problem →Need to design or review plans
Combine engineering deductions and senior self-evaluation standards to explain the applicable conditions, costs and alternatives of the plan.
Analyze engineering scenarios →LEARN · PRACTICE · REFLECT
First read along the principles, Q&A and migration cases. When you need to check your understanding, switch to reinforcement exercises or start personal recording.
Can be practiced directly. After logging in, answers, favorites, and notes will be saved to your account.
Log in and saveEach modified commit will be kept as an independent history. Your level of mastery is up to you to evaluate yourself against the standards.
Core concept · Version closure and incremental consistency during embedding migration
Preparatory concepts:versioned index, Change streams and watermarks, Canary release
The text preprocessing, query vector and document vector of a query must belong to compatible representation contracts; index switching must also maintain update, delete and permission constraints, not just move the vector.
Two models may both produce 768-dimensional vectors while using different coordinate semantics. Latitude/longitude and projected coordinates likewise share numeric shapes without direct comparability. Model, normalization, chunking, and preprocessing define a retrieval version; query vectors must match it.
Backfill from a snapshot watermark while retaining subsequent updates, deletions, and ACL changes. Each document carries a source revision. Accept writes only when their revision is at least as new as the indexed revision, preventing slow backfill from overwriting updates. Version deletion markers too; copying currently existing documents once can miss changes or resurrect deleted data.
Elastic supports atomic alias-action groups, but applications still check action results. Alias switching does not change the query embedding model. Bind model, index, preprocessing, and cache under one request-pinned configuration version. Canaries and shadow reads use matching model vectors, comparing coverage, permissions, and refusals rather than cross-model similarity scores.
Retained old indexes must continue receiving deletions, revocations, and required updates. Reverting to a three-day-old index can otherwise expose deleted documents. Switch only after complete backfill, caught-up increments, authorization validation, and acceptable quality and cost. A few passing queries cannot establish consistency across millions of records.
Equal vector dimensions do not establish compatibility across embedding models. Version model, chunking, and preprocessing; backfill a snapshot while applying incremental changes and deletions. Queries use the matching model and index. Switch after shadow and canary evaluation. Retain rollback indexes while propagating deletions and access revocations to both versions.
Version metadata includes model identification, dimensions, normalization, distance metric, tokenization or preprocessing method, chunking strategy and document version. The same dimensionality does not prove that the two models are in the same semantic space. The query vector must be generated from a matching version, otherwise the index may return results normally but have completely wrong relevance. This type of silent failure is more difficult to detect than interface errors.
Create a document snapshot at time point T to record the starting point of the change stream; the new index backfills the content of T while replaying subsequent additions, updates, and deletions. Events are deduplicated according to document versions, and old events cannot overwrite the updated results. Keep trackable tombstones or equivalent version rules for deletions to prevent late backfills from rewriting deleted content. Permission revocation needs to affect visibility quickly and cannot wait for a complete rebuild.
Verify document count, missing blocks, version watermark, and deletion propagation, and then compare retrieval and answer quality using the same query set. Shadow queries are not user facing but should also respect data permissions. Canary release groups by stable users or tasks to avoid switching spaces back and forth in the same session. Observe quality, latency, and cost, and switch default aliases only when predefined thresholds are reached.
Rolling back not only changes the alias, but also restores the corresponding query encoder, cache namespace and retrieval configuration. The old index still needs to be synchronized for key revocation and deletion during the retention period and cannot become a backup channel for expired data leakage. Confirm that the new index is stable and there are no unfinished tasks for the old version before recycling, and record the recoverable range and changes that cannot be rolled back.
Continue reading along with the premises and constraints of the problem. Understand the reference answers first, then try to put away the answers and explain the cause and effect and trade-offs in your own words.
Level 1Can old vectors be directly shared with the same dimensions?
Explain the most misused types of compatibility from migration constraints.
No. The same dimension only satisfies the storage shape, and model training and vector coordinate semantics may be different. Unless the provider explicitly guarantees spatial compatibility, or has a validated mapping, document vectors are reconstructed and matched to the query model. Normalization methods and text preprocessing changes must also be verified together; "the interface does not report an error" is not evidence of compatibility.
Level 1What should I do if a document is deleted during backfill?
Snapshot migration introduces concurrent changes, and deletion requires independent consistency rules.
Deletion events enter the persistent change stream along with document revisions. The new index remembers the tombstone or deletion watermark and rejects late-arriving old backfills; the old index also handles deletions. Before switching, reconcile live/deleted state and change-stream lag between the source and both indexes; do not compare counts alone, because if the quantity is the same, one deletion and one excess may cancel each other out.
Follow this answer further
Level 2The deletion event arrives first, and the new index does not yet have the document. Can I just ignore it?
The father asked to delete the spread, further exposing the resurrection window after "deleting non-existent records".
No, subsequent backfilling may bring older versions. Deleted revisions or watermarks are saved even if the body does not exist, and any older writes are rejected. If the index cannot save tombstone, it can be maintained at the trusted metadata layer and checked before writing; batch retries are also handled to ensure that old snapshots do not bypass checks.
Follow this answer further
Level 3How can a user re-create a document with the same ID after deletion so that it is not permanently blocked by the old tombstone?
Tombstone anti-resurrection also needs to allow legal reconstruction and verify the complete semantics of version comparison.
Distinguish between business identification and record generation. To recreate with a higher revision or new generation, write conditions are compared against the version by generation; old tombstones only prevent old facts from being resurrected. If the source system cannot guarantee the version order, a new stable identity will be adopted and the association will be clear. The arrival time cannot be used to infer which creation is valid.
Level 1Does the cache need to be changed when rolling back?
Switching is not only about the index, but the derived results also belong to the version contract.
The cache needs to be bound to the retrieval configuration version as well as the knowledge and permission versions. When rolling back, the cache corresponding to the old configuration is used, provided that it still meets the current source and permissions; the retrieval results of the new model cannot be reused just because the questions are the same. Answer caches that are missing dependency information should be invalidated and regenerated to avoid references pointing to different chunks or revoked data.
First find out the conditions for change, and then determine which premises in the original plan still hold true. The following cases are teaching deductions to facilitate the transfer of principles to new problems.
Changing conditions:In addition to the model, also change the block ID and evidence location
Extended question:How to retain old citations and user favorites?
Establish stable positioning for the original document_id, source_version, and paragraph anchors. The old chunk_id is only used as a historical index identity. The new block retains the source range, using a mapping to jump back to the original text rather than hard mapping to the "most similar" new block; old references that cannot be matched are explicitly invalidated. Migration assessments need to redefine the scope of necessary evidence.
The principles that remain unchanged:The search implementation can change, and source facts and reference positioning cannot be guessed based on similarity.
Changing conditions:Migration adds pressure on downstream budgets
Extended question:Should the full switch be made if the recall rate increases?
Compare answer support, latency, and cost by question type, including unanswerable and access-restricted questions. Use canaries grouped by query type, retaining the old configuration where it already meets the goal. Check capacity and budget before expansion. Pin the entire query configuration; never mix candidates from old and new embedding spaces directly.
The principles that remain unchanged:Changes need to verify full business benefits, and the consistency of the representation space cannot be compromised by cost tradeoffs.
It is designed based on public technical information; the reference materials support the technical mechanism, and the scenarios and scoring standards are designed by this website and do not represent the original interview questions of a certain company. New Q&A and migration cases are added for principle explanation, and source verification and case operation verification are recorded separately.
After reading, you can explain the principles, boundaries, and trade-offs against these standards. It is up to you to evaluate your mastery; if further verification is needed, complete the small tasks below.
Transferring pre-retrieval filtered ideas to memory versions and recovery. Observe how existing drafts become invalid after cancellation.
Read full text and fault analysis → · Download Reliability Experiment v3 ↓
python3 cli.py memory-put --db memory.sqlite
python3 cli.py submit --db memory.sqlite
python3 cli.py run --db memory.sqlite --lease-seconds 2 --fault after_draft
python3 cli.py memory-forget --db memory.sqlite
# 等待至少 2 秒后分别执行
python3 cli.py run --db memory.sqlite
python3 cli.py inspect --db memory.sqliteVerify local scope, version, and undo blocking; old checkpoints remain, no complete deletion of logs, backups, or checkpoints is provided.
Draw the migration timing of the three events of adding, updating, and deleting after the snapshot at time T.