Understand first
New to this knowledge point
Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.
Start with core principles →Understand → Implement → Debug → Design
Examine double-write issues, transaction outboxes, duplicate deliveries, and recovery scans.
Knowledge content check2026-10-03 · Check the source of the original question2026-10-02
It is recommended to understand first:
Idempotency, unknown outcomes, and task recovery →Select the starting point based on the current basis, or you can go deeper one by one. When you encounter an unfamiliar concept, go back to the core principles first; use the knowledge exercises to check your understanding when you are finished.
New to this knowledge point
Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.
Start with core principles →Prepare to write the principles into code
Understand implementation steps and boundaries, complete small tasks, and check results against acceptance requirements.
View the code example →Need to handle failures and changes in conditions
Follow the continuous questioning to locate the failure premise, and then compare the migration cases to explain how the plan should be adjusted.
Continue to delve deeper into the problem →Need to design or review plans
Combine engineering deductions and senior self-evaluation standards to explain the applicable conditions, costs and alternatives of the plan.
Analyze engineering scenarios →LEARN · PRACTICE · REFLECT
First read along the principles, Q&A and migration cases. When you need to check your understanding, switch to reinforcement exercises or start personal recording.
Can be practiced directly. After logging in, answers, favorites, and notes will be saved to your account.
Log in and saveEach modified commit will be kept as an independent history. Your level of mastery is up to you to evaluate yourself against the standards.
Core concept · Atomic intent and duplicate delivery with the outbox pattern
Preparatory concepts:local database transaction, Queue delivery semantics, Idempotent consumption
The same transaction submits the task and the intention to be delivered, eliminating the double-write window of "task exists but notification is permanently lost"; actual delivery may still be repeated, and consumers must independently ensure that business processing can be reentrant.
Write the task first and a crash can prevent message delivery. Send the message first and a consumer may find no task record. try/catch cannot handle sudden process death. An outbox commits business rows and pending events in one database transaction.
Scan unconfirmed events, send them, then mark delivery. A successful send followed by a failed mark causes redelivery, requiring stable event_id, business revisions, and idempotent consumers. The outbox does not combine the database and queue into one transaction or guarantee one external effect. If ordering matters, use per-object sequence numbers rather than network arrival order.
ACK semantics depend on the queue and may establish only acknowledged delivery. Consumers first claim work durably or record idempotent handling, then ACK at the agreed boundary. Long agent tasks commonly use messages to start durable runs instead of holding messages throughout inference. Run state and business receipts establish the terminal outcome.
Monitor oldest undelivered events, retries, unscheduled tasks, and dead letters. Reconciliation scans identify state discrepancies. Cleanup depends on confirmed delivery, retention, and redelivery needs, rather than creation time alone. Kill processes during rollback, after commit before send, after send before marking, and after consumption before ACK. Verify no lost tasks and explicitly defined duplicate handling.
Commit tasks and pending outbox events in one database transaction. A separate dispatcher retries delivery; consumers deduplicate stable event or operation IDs. Delivery can repeat, while committed intent survives temporary send failures. Monitor oldest undelivered events and stuck tasks, using reconciliation scans to identify discrepancies.
If you submit the task first and then send the message, the process may crash in the middle, and the task will never be executed; if you send the message first and then submit the task, the consumer may not be able to find the task, or the transaction may be rolled back in the end. Writing "try again in case of exception" in the request thread cannot cover the situation when the process disappears. A to-be-delivered fact needs to be persisted at the same time as task submission.
Write runs and outbox in the same transaction. The event contains event_id, run_id, type, payload version and creation time. The deliverer receives unsent events and updates the status after sending. If the sending is successful but the status update fails, the sending will be repeated next time, so the consumer must have deduplication or idempotent business transfer. Do not wait for model calls or remote message acknowledgments while holding long database transactions.
The consumer transfers and receives the task according to the legal status, and repeated event_id does not repeatedly create logical operations. Processing business results and consumption records can be submitted in the same storage transaction; external side effects still require business idempotency and receipt verification. The queue's ACK is only a confirmation of delivery processing and cannot replace the business success status. Sequence-sensitive tasks use version or sequence checks, and late events cannot return completed tasks to the queue.
Four downtime points are injected: before and after the transaction is committed, after the message is sent, and before the mark is sent. Asserts that all submitted tasks are eventually discoverable and that duplicate messages will not duplicate side effects. Monitor Outbox oldest event age, retries, dead letters, and runtime. Compensating scans require the use of the same idempotency keys, which cannot fix lost tasks and create double executions.
SQLite in-memory transaction demo; does not include real queue posters, idempotent consumption, or distributed stress testing.
import sqlite3
db = sqlite3.connect(":memory:")
db.executescript("""
CREATE TABLE runs(id TEXT PRIMARY KEY);
CREATE TABLE outbox(event_id TEXT PRIMARY KEY, run_id TEXT NOT NULL);
""")
def submit(run_id, fail=False):
with db:
db.execute("INSERT INTO runs VALUES (?)", (run_id,))
if fail:
raise RuntimeError("crash before outbox")
db.execute("INSERT INTO outbox VALUES (?, ?)", (run_id + ":created", run_id))
try:
submit("r1", fail=True)
except RuntimeError:
pass
print("after rollback:", db.execute("SELECT count(*) FROM runs").fetchone()[0],
db.execute("SELECT count(*) FROM outbox").fetchone()[0])
submit("r2")
print("after commit:", db.execute("SELECT count(*) FROM runs").fetchone()[0],
db.execute("SELECT count(*) FROM outbox").fetchone()[0])
db.close()
expected output
after rollback: 0 0
after commit: 1 1Continue reading along with the premises and constraints of the problem. Understand the reference answers first, then try to put away the answers and explain the cause and effect and trade-offs in your own words.
Level 1What happens if the sending succeeds but the tag fails?
Second failure window from atomic commit intent into external delivery.
The same event_id will be sent again when the deliverer cannot acknowledge the local tag. The consumer side checks the processed ID or receives the same run atomically, returns the existing status, and does not create tasks repeatedly. Even if the queue comes with a deduplication window, the service remains idempotent because the window and playback period may be different. Record the number of re-investments for troubleshooting.
Follow this answer further
Level 2If two deliverers scan the same unsent event at the same time, will they both be sent?
Concurrent delivery adds another source of duplication after the parent asks to acknowledge re-delivery.
Possibly. Use short transaction conditional acquisition, lease or row lock to reduce concurrent duplication, message sending is completed outside the transaction, and the acquisition order is verified when the mark is submitted. Even so, timeout takeovers and response losses may still be repeated, and consumers cannot omit idempotency. Receipt control improves efficiency, and business correctness does not rely on "never reissue".
Follow this answer further
Level 3The consumer first checks that event_id does not exist and then executes it. Why is it still possible to repeat it?
Delivery reaches the consumer repeatedly, and the race condition of post-read execution requires atomic processing.
Two consumers can check that it does not exist at the same time. With unique constraints and atomic state acquisition, business changes and idempotent confirmations are completed in the same local transaction; if the business action is on the remote end, create a unique intention and execute it by pressing the stable operation key. You cannot use a SELECT to act as an idempotent lock, the transaction boundaries must cover the truly protected state.
Level 1Does the message ACK mean the business is successful?
Distinguish between transport-level success and task-level success.
Doesn't mean. ACK represents the acknowledgment of the queue layer. The business may not have been completed or only the pending execution status has been reliably saved. ACK boundaries should be defined: persistent scheduling intent can be ACKed after success, and task completion is queried by independent status and receipt. If the ACK is sent before the business result is persisted and there is no recovery clue, processing may still be lost.
Level 1How to avoid deleting undelivered events when cleaning Outbox?
Reliable links finally have a life cycle of retention and cleanup.
Only events that meet the confirmed delivery and retention conditions are cleared, undelivered and unknown are processed separately; event_id, business version and necessary receipts are archived and retained to support redelivery. When deleting by partition, unfinished events are first checked and alarmed, and old events are not regarded as completed. Cleanup and delivery use conditional states to avoid deleting rows being processed.
First find out the conditions for change, and then determine which premises in the original plan still hold true. The following cases are teaching deductions to facilitate the transfer of principles to new problems.
Changing conditions:Multiple revisions of the same object enter the queue
Extended question:Can repeated deduplication prevent old versions from being overwritten?
No, deduplication prevents the same event from being repeated, but different events may be out of order. Record the object revision and reject the old revision from overwriting the new state; detect the gap according to the object serial number and fill it when it needs to be processed step by step. The consumer processes the business status and records the processed ID as much as possible in the same local transaction.
The principles that remain unchanged:Event identity and status sequence are independent dimensions, and idempotency cannot replace version verification.
Changing conditions:Business side effects occur outside of consumer affairs
Extended question:Is it guaranteed to be exactly once after inserting the ID in processed_events?
If the ID is inserted, it will go down and the action will be missed. If the action is successful, the action will be repeated if the ID is inserted again. Save the intention first, use the stable operation_id to perform remote idempotency or reconciliation, and then confirm the result; the processed ID can indicate that it has been reliably handed over to the running state machine, and cannot be falsely claimed that the remote end has been completed. Remain unknown and manually verified without target support.
The principles that remain unchanged:Local atomic intentions cannot span remote transactions, and side effects still need to independently restore the contract.
It is designed based on public technical information; the reference materials support the technical mechanism, and the scenarios and scoring standards are designed by this website and do not represent the original interview questions of a certain company. New Q&A and migration cases are added for principle explanation, and source verification and case operation verification are recorded separately.
After reading, you can explain the principles, boundaries, and trade-offs against these standards. It is up to you to evaluate your mastery; if further verification is needed, complete the small tasks below.
View verification records for independent examples
Draw the transaction boundaries of task submission, Outbox delivery and consumption, and mark recurrence points.