Agent Application DevelopmentAccount
Knowledge catalogChoose core direction and segmented content
knowledge unit 43AdvancedImplementationAbout 18 minutes

Understand → Implement → Debug → Design

Failure windows and recovery invariants

Examine failure points, business invariants, deterministic tool stand-ins, and recovery acceptance.

fault injectionRecovery testinvariant

Knowledge content check2026-10-03 · Check the source of the original question2026-10-02

Which step do you want to learn from this knowledge point?

Select the starting point based on the current basis, or you can go deeper one by one. When you encounter an unfamiliar concept, go back to the core principles first; use the knowledge exercises to check your understanding when you are finished.

Understand first

New to this knowledge point

Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.

Start with core principles →

Realize again

Prepare to write the principles into code

Understand implementation steps and boundaries, complete small tasks, and check results against acceptance requirements.

Reading implementation and trade-offs →

Will troubleshoot

Need to handle failures and changes in conditions

Follow the continuous questioning to locate the failure premise, and then compare the migration cases to explain how the plan should be adjusted.

Continue to delve deeper into the problem →

Able to choose

Need to design or review plans

Combine engineering deductions and senior self-evaluation standards to explain the applicable conditions, costs and alternatives of the plan.

Analyze engineering scenarios →
Knowledge unit directory

LEARN · PRACTICE · REFLECT

Knowledge learning and personal records

My notes and review ↗

First read along the principles, Q&A and migration cases. When you need to check your understanding, switch to reinforcement exercises or start personal recording.

Answers and personal notes

Each modified commit will be kept as an independent history. Your level of mastery is up to you to evaluate yourself against the standards.

Core concept · Failure windows and recovery invariants

Understand the core principles first

Preparatory concepts:transaction commit, Idempotent operations, state machine

Recovery correctness is determined by the business facts before and after the failure. The test must cut off the window between "submission, receipt, and status advancement" to prove that replay will not break idempotency, permissions, and budget constraints.

A 500 response is not every failure

Failure before commit and lost acknowledgment after commit can both appear as timeouts. The first may allow a safe retry; blindly repeating the second may duplicate an effect. Injection must control whether the target committed, rather than always returning the same error.

Isolate runtime with stateful test doubles

A fixed model supplies a tool plan; a tool test double records business keys, effects, and receipts. Terminate the process before local persistence, after external commit, and after receiving a receipt. Inspect target records and local state after recovery. Repeated requests may be valid; effects must satisfy the contract.

Distinguish safety from availability

An unverifiable outcome remains unknown. Avoiding duplicate writes may satisfy safety while missing a completion deadline. Record both conclusions. A safe pause does not establish successful recovery, and the system must not guess success. These additional fault paths are proposed tests, not new experiments.

Check understanding with a question

How to prove that the Agent can recover from failures instead of just passing the normal process test?

Inject faults around persistence, tool commits, response delivery, and state advancement. Stateful test doubles reproduce lost receipts, duplicates, revocations, and worker takeover. Verify effects, access, budget conservation, and truthful terminal state. Fixed model outputs isolate runtime logic; real-model runs then evaluate end-to-end behavior.

Realization and trade-offs

Backwards testing from invariants

First, list the conditions that must always be true: the same logical action will not have repeated effects, new actions will not be started after cancellation, recovery will not increase the used budget, revocation of permissions will immediately affect execution, and there must be evidence of acceptance for success. Each invariant maps to a state boundary that may be broken. You cannot just check whether the final page is displayed successfully.

Inject real failure window

It crashes before writing the graph, after external submission, before receiving the response, and before saving the result. Then add lease expiration, old workers continuing to write, queue duplication and disorder, and configuration changes during the recovery process. Tool stand-ins record all requests, idempotency keys, and actual effects so that tests can distinguish repeated requests from repeated business side effects.

Control model randomness

Runtime reliability testing uses a fixed model response sequence, such as continuously repeating the same call, claiming completion early, and raising override parameters. Such a failure can accurately locate the scheduling and policy logic. The real model test is run separately to observe task completion and adaptability. The two cannot replace each other. The test environment must be isolated from the real sending, payment and publishing tools.

Accept recovery results

Check the operation ledger, final artifact, failure classification, receipts and audit events, not just function returns. For each injection point, record whether the recovery was successful, how time-consuming it was, and the reasons why manual intervention was required. An unknown state is not necessarily an implementation error; explicitly stopping when it cannot be safely checked is more consistent with business invariants than forging success.

Engineering deduction

scene
Interview hypothesis: The publishing tool has idempotency keys, but the recovery logic generates new keys every time.
design decisions
Tool test doubles record effects, inject responses when lost and then recover.
Verify target
Testing can reveal duplicate releases even if a single normal call passes.
applicable boundary
Isolated experimental passes do not equate to full conformance guarantees from real vendors.

Continuous questions and answers

Continue reading along with the premises and constraints of the problem. Understand the reference answers first, then try to put away the answers and explain the cause and effect and trade-offs in your own words.

Draw inferences from one example: If the conditions change, how to deduce it?

First find out the conditions for change, and then determine which premises in the original plan still hold true. The following cases are teaching deductions to facilitate the transfer of principles to new problems.

Read-only retrieval failed

Changing conditions:There are no write side effects, but there are query charges and permission changes.

Extended question:Can I just retry infinitely?

Derivation and reference solutions

No. Read-only reduces the risk of duplicate effects, but still requires bounded backoff, cost budgeting, and each authorization check; failure final status is recorded. If a new version of the data is used in a repeated search, answers should be rechecked rather than mixing old evidence with new results.

The principles that remain unchanged:Recovery remains subject to fact, permission, and resource invariants.

Multi-step refunds and inventory compensation

Changing conditions:From a single tool to multiple submissions for different services.

Extended question:Is it enough to inject one failure at the beginning?

Derivation and reference solutions

Not enough. Test executed, unexecuted, and unknown for each submission and compensation window, checking that compensation itself is idempotent and has real receipts. Business compensation does not erase the original facts; if it cannot be undone, it will be handled manually, and it does not claim automatic atomic rollback across services.

The principles that remain unchanged:The failure window determines recovery actions, and compensation is also a controlled side effect.

Easy to make mistakes

  • Only test normal paths and HTTP 500
  • Recovery testing really sends a message out
  • Only look at the return code and not the side effect ledger

References

It is designed based on public technical information; the reference materials support the technical mechanism, and the scenarios and scoring standards are designed by this website and do not represent the original interview questions of a certain company. New Q&A and migration cases are added for principle explanation, and source verification and case operation verification are recorded separately.

Check how far you understand

After reading, you can explain the principles, boundaries, and trade-offs against these standards. It is up to you to evaluate your mastery; if further verification is needed, complete the small tasks below.

Basic standards met
Ability to propose timeout, restart and duplicate message scenarios.
Intermediate and advanced signals
Locate the precise commit window and check business invariants.
Senior Signal
Build controllable fault surrogates, fixed model sequences, and end-to-end hierarchical evaluation.

Hands-on verificationComplete on demand · Suggestions15 minutes

List four downtime points for a release action and write the expected recovery behavior at each point.

Expand acceptance requirements and checkpoints
  • At least overwrite after submission and before response
  • Assert the actual number of effects
  • Unknown status will not be successfully disguised

Key inspections

  • Test by commit window instead of just status code
  • Ability to log requests and real side effects
  • Separate verification of runtime and model quality