Understand first
New to this knowledge point
Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.
Start with core principles →Understand → Implement → Debug → Design
Examine failure points, business invariants, deterministic tool stand-ins, and recovery acceptance.
Knowledge content check2026-10-03 · Check the source of the original question2026-10-02
It is recommended to understand first:
Idempotency, unknown outcomes, and task recovery →Agent evaluation: outcomes, constraints, and evidence →Select the starting point based on the current basis, or you can go deeper one by one. When you encounter an unfamiliar concept, go back to the core principles first; use the knowledge exercises to check your understanding when you are finished.
New to this knowledge point
Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.
Start with core principles →Prepare to write the principles into code
Understand implementation steps and boundaries, complete small tasks, and check results against acceptance requirements.
Reading implementation and trade-offs →Need to handle failures and changes in conditions
Follow the continuous questioning to locate the failure premise, and then compare the migration cases to explain how the plan should be adjusted.
Continue to delve deeper into the problem →Need to design or review plans
Combine engineering deductions and senior self-evaluation standards to explain the applicable conditions, costs and alternatives of the plan.
Analyze engineering scenarios →LEARN · PRACTICE · REFLECT
First read along the principles, Q&A and migration cases. When you need to check your understanding, switch to reinforcement exercises or start personal recording.
Can be practiced directly. After logging in, answers, favorites, and notes will be saved to your account.
Log in and saveEach modified commit will be kept as an independent history. Your level of mastery is up to you to evaluate yourself against the standards.
Core concept · Failure windows and recovery invariants
Preparatory concepts:transaction commit, Idempotent operations, state machine
Recovery correctness is determined by the business facts before and after the failure. The test must cut off the window between "submission, receipt, and status advancement" to prove that replay will not break idempotency, permissions, and budget constraints.
Failure before commit and lost acknowledgment after commit can both appear as timeouts. The first may allow a safe retry; blindly repeating the second may duplicate an effect. Injection must control whether the target committed, rather than always returning the same error.
A fixed model supplies a tool plan; a tool test double records business keys, effects, and receipts. Terminate the process before local persistence, after external commit, and after receiving a receipt. Inspect target records and local state after recovery. Repeated requests may be valid; effects must satisfy the contract.
An unverifiable outcome remains unknown. Avoiding duplicate writes may satisfy safety while missing a completion deadline. Record both conclusions. A safe pause does not establish successful recovery, and the system must not guess success. These additional fault paths are proposed tests, not new experiments.
Inject faults around persistence, tool commits, response delivery, and state advancement. Stateful test doubles reproduce lost receipts, duplicates, revocations, and worker takeover. Verify effects, access, budget conservation, and truthful terminal state. Fixed model outputs isolate runtime logic; real-model runs then evaluate end-to-end behavior.
First, list the conditions that must always be true: the same logical action will not have repeated effects, new actions will not be started after cancellation, recovery will not increase the used budget, revocation of permissions will immediately affect execution, and there must be evidence of acceptance for success. Each invariant maps to a state boundary that may be broken. You cannot just check whether the final page is displayed successfully.
It crashes before writing the graph, after external submission, before receiving the response, and before saving the result. Then add lease expiration, old workers continuing to write, queue duplication and disorder, and configuration changes during the recovery process. Tool stand-ins record all requests, idempotency keys, and actual effects so that tests can distinguish repeated requests from repeated business side effects.
Runtime reliability testing uses a fixed model response sequence, such as continuously repeating the same call, claiming completion early, and raising override parameters. Such a failure can accurately locate the scheduling and policy logic. The real model test is run separately to observe task completion and adaptability. The two cannot replace each other. The test environment must be isolated from the real sending, payment and publishing tools.
Check the operation ledger, final artifact, failure classification, receipts and audit events, not just function returns. For each injection point, record whether the recovery was successful, how time-consuming it was, and the reasons why manual intervention was required. An unknown state is not necessarily an implementation error; explicitly stopping when it cannot be safely checked is more consistent with business invariants than forging success.
Continue reading along with the premises and constraints of the problem. Understand the reference answers first, then try to put away the answers and explain the cause and effect and trade-offs in your own words.
Level 1What is the difference between repeat requests and repeat effects?
Recovery often relies on retransmissions, so first differentiate between network attempts and real effects.
A repeated request is to try the same business action again, and the repeated effect is that the target system actually does it twice. Systems that allow at least one delivery can request the same key multiple times, but the receiving end needs to limit the effect through unique business keys, status checks, or idempotent mechanisms; just counting HTTP calls cannot determine whether idempotency is correct.
Follow this answer further
Level 2The deduplication on the receiving side is successful, but the local memory has not been completed. What should I do during recovery?
The parent question distinguishes repeated requests and effects, and the child question increases the fault window when the local state is lagging behind.
Query the saved results of the receiver along the same business key, persist the verification receipt, and then advance it to the local state. If the deduplication interface only says "duplicate" but does not return results, you still need to query independently; you cannot create a new business key to bypass deduplication. The test should assert that the local after recovery is consistent with the target final state.
Follow this answer further
Level 3The query also times out. Is it possible to use a new key to ensure final completion?
The parent asks the dependent result query, and then lets the query fail to check the processing of the unknown state.
Safety cannot be guaranteed on this basis. The new key may perform the completed action again; save the original key and unknown status, and continue to reconcile or pay labor according to the budget. Only when credible evidence of non-execution is obtained, or the business clearly allows repeated effects, will new actions be discussed; lack of acknowledgment from the network is not evidence of non-implementation.
Level 1Why is fixed model response still valuable for testing?
Malfunction experiments require repeatable control variables that account for the value boundaries of stand-ins.
Fixed output allows control over tool selection and parameters, reliably reproducing runtime errors such as checkpoints, timeouts, cancellations, and privilege revocations. It cannot prove that the model will take the correct next step when faced with real receipts, so subsequent behavioral evaluation of the real model is required; the two-tier tests answer different questions and should not be substituted for each other.
Level 1Does stopping at unknown when it cannot be verified count as a test failure?
In the absence of a target receipt it must be clear what exactly the test is asking for.
Depends on the contract. If a safety pause is required when it cannot be verified, if it stops at unknown and does not continue writing, the safety test has passed; if it is also required to be completed within ten minutes, the availability has not met the standard. Report security results, recovery status, and unfinished reasons respectively, retain the manual reconciliation entry, and do not change unknown to unexecuted to try again.
First find out the conditions for change, and then determine which premises in the original plan still hold true. The following cases are teaching deductions to facilitate the transfer of principles to new problems.
Changing conditions:There are no write side effects, but there are query charges and permission changes.
Extended question:Can I just retry infinitely?
No. Read-only reduces the risk of duplicate effects, but still requires bounded backoff, cost budgeting, and each authorization check; failure final status is recorded. If a new version of the data is used in a repeated search, answers should be rechecked rather than mixing old evidence with new results.
The principles that remain unchanged:Recovery remains subject to fact, permission, and resource invariants.
Changing conditions:From a single tool to multiple submissions for different services.
Extended question:Is it enough to inject one failure at the beginning?
Not enough. Test executed, unexecuted, and unknown for each submission and compensation window, checking that compensation itself is idempotent and has real receipts. Business compensation does not erase the original facts; if it cannot be undone, it will be handled manually, and it does not claim automatic atomic rollback across services.
The principles that remain unchanged:The failure window determines recovery actions, and compensation is also a controlled side effect.
It is designed based on public technical information; the reference materials support the technical mechanism, and the scenarios and scoring standards are designed by this website and do not represent the original interview questions of a certain company. New Q&A and migration cases are added for principle explanation, and source verification and case operation verification are recorded separately.
After reading, you can explain the principles, boundaries, and trade-offs against these standards. It is up to you to evaluate your mastery; if further verification is needed, complete the small tasks below.
List four downtime points for a release action and write the expected recovery behavior at each point.