Understand first
New to this knowledge point
Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.
Start with core principles →Understand → Implement → Debug → Design
Choose a recoverable orchestration structure around defined processes, dynamic exploration, checkpoints, and manual approvals.
Knowledge content check2026-10-03 · Check the source of the original question2026-10-02
It is recommended to understand first:
Control and completion criteria in agent loops →Select the starting point based on the current basis, or you can go deeper one by one. When you encounter an unfamiliar concept, go back to the core principles first; use the knowledge exercises to check your understanding when you are finished.
New to this knowledge point
Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.
Start with core principles →Prepare to write the principles into code
Understand implementation steps and boundaries, complete small tasks, and check results against acceptance requirements.
Reading implementation and trade-offs →Need to handle failures and changes in conditions
Follow the continuous questioning to locate the failure premise, and then compare the migration cases to explain how the plan should be adjusted.
Continue to delve deeper into the problem →Need to design or review plans
Combine engineering deductions and senior self-evaluation standards to explain the applicable conditions, costs and alternatives of the plan.
Analyze engineering scenarios →LEARN · PRACTICE · REFLECT
First read along the principles, Q&A and migration cases. When you need to check your understanding, switch to reinforcement exercises or start personal recording.
Can be practiced directly. After logging in, answers, favorites, and notes will be saved to your account.
Log in and saveEach modified commit will be kept as an independent history. Your level of mastery is up to you to evaluate yourself against the standards.
Core concept · Combine deterministic workflows with bounded exploration
Preparatory concepts:state machine, data contract, Idempotent execution
Known business rules should determine which paths are allowed, and the model only explores unknown steps within the allowed range. Graphs improve process observability, and loops provide local adaptability; neither replaces the idempotency and version acceptance of external writes.
Refund eligibility, approval, and execution follow business rules. Investigating inconsistent information may require new searches based on emerging evidence. Explicit transitions suit the former; controlled loops suit the latter. Leaving mandatory checks to the model makes them optional. Encoding every exploratory search in a graph can overconstrain investigation.
An outer graph can define evidence gathering, drafting, review, and publication. The evidence node may allow several read-only searches, returning evidence IDs, coverage gaps, and input versions. Branch on those fields rather than guessing from prose. Failed exploration returns gaps without bypassing review.
A checkpoint cannot atomically commit graph state and send an email. Recovery may re-enter a node, so its effects still need protection. LangGraph’s interrupt recovery restarts the node, illustrating this boundary. Compare recovery points and external action boundaries when choosing a framework.
Use state graphs for known steps and acceptance rules, and controlled agent loops when results determine the next investigation. They can coexist: an outer retrieval–draft–review–publish workflow can contain exploratory read-only calls. Graphs and persistence do not guarantee idempotent external writes. Define replay points, write protection, approval bindings, and state versions.
If the sequence of steps is determined by the business, such as data storage, permission checking, draft generation, and manual review, the control process should be explicitly modeled. If the next step is determined by evidence, such as locating unknown faults or searching multiple materials, the model needs to explore the space. LangGraph officially distinguishes between workflows with predetermined paths and Agents that dynamically determine actions; there is no need to choose one or the other in engineering. The outer workflow carries the compliance and life cycle, the inner loop of the node handles the variable research process, and the node output must go through a stable data contract to enter the next stage.
Status shouldn't just be an array of messages. At least include task identification, input version, current stage, evidence reference, artifact reference, budget and approval records. Nodes read status and return changes, and branches are based on structured values to avoid guessing "pass" or "fail" from natural language. When writing the same field in parallel branches, clear merging rules need to be specified, such as deduplication based on document identification rather than relying on arrival order. For certain permission denials and budget exhaustion, directly determine the final state; do not let the model repeatedly strive for prohibited actions.
Persistent checkpoints record the running status of threads, and long-term storage saves cross-thread knowledge. The two have different purposes. Production recovery uses persistent storage and stable thread_id; memory checkpoints are for demonstration only. LangGraph interrupt will pause and wait for recovery input. The node where it is located will be re-executed during recovery, so the operations before approval must be replayable safely or be split to an independent node. Approval should be bound to the artifact version or summary, and permissions and data versions should be re-verified after recovery. Old approval should not be allowed to approve subsequent changes. External write operations still require business idempotency keys and result queries.
Three types of drills are designed: restarting the process after a node is abnormal, modifying the artifact while waiting for approval, and completing two branches at the same time. Check whether the completed evidence is retained, whether the approval has expired, whether the aggregation results are stable, and whether external side effects are repeated. Then compare the number of graph nodes, debugging cost and model rounds. There is no need to split a simple question and answer into dozens of nodes; nor can long-term tasks rely solely on saving conversations in the database. Being able to clearly explain the recovery point and failure boundary is the basis for selection.
Continue reading along with the premises and constraints of the problem. Understand the reference answers first, then try to put away the answers and explain the cause and effect and trade-offs in your own words.
Level 1Parallel nodes modify the evidence list at the same time. How should they be merged?
Parallel graphs turn serial reading and writing into a merging problem, and the semantics of the same field need to be clarified.
Have nodes return deltas instead of overwriting the full list, and define reducers that merge by evidence ID and version. The same ID, the same version and the same content can be deduplicated. If the same ID has different content, conflicts will remain or merge will be refused. Silent overwriting cannot be done based on the order of arrival. If the display requires order, sort by stable fields first to avoid taking the order returned by the network as the order of fact.
Level 1Recovery will rerun the node. What will happen if a notification is sent before approval?
Recovery points determine which code is replayed, and persistence does not mean that external actions are executed only once.
Notifications may be sent again after recovery because the code before interrupt will be re-run. Change the notification to an independent action with a stable business key, or move it to the node after the approval is successful; it is still necessary to prevent the node from crashing after execution and before the checkpoint is saved. Simply moving the sending statement cannot guarantee no duplication. The notification gateway needs to save the delivery status and check the receipt.
Follow this answer further
Level 2After the notification becomes an independent node, what should I do if it is successfully sent but fails to save the checkpoint?
Splitting nodes reduces the replay range, but still leaves a window between remote execution and local persistence.
Recovery will treat the node as incomplete, so use the same notification action key to query or replay old results. If the remote end has idempotency support, the key will be reused; if there is no idempotency support, the key will be verified first based on the business number. If it cannot be verified, the key will remain unknown. The checkpoint records the running progress, and the notification ledger records external facts. Both perform their respective duties.
Follow this answer further
Level 3What should I do if I manually determine that an unknown notification has not been sent and only receive a successful receipt later?
The results must correct the facts when they finally arrive, rather than just maintain a consistent appearance of the process state.
Keep the timeline of the original operation and manual judgment, press the same key for late receipts to update the facts and mark the judgment conflict; if a second item has been issued, record the actual duplication and press the business remedy. Late events cannot be discarded to maintain the appearance of "send only once", and avoiding automatic retries is often more appropriate for high-impact unknown actions.
Level 1How to prevent the draft you see during approval from being different from the published content?
There is a content change window between the review node and the publishing node.
Approval of binding draft content summary, attachment versions and release targets, recalculated and compared before release. Any changes that affect actual actions invalidate old approvals; whether pure display metadata can be reused should be clarified by policy. The best approved is an immutable published object, which is only read by the publishing node, avoiding another writable field replacing the body.
First find out the conditions for change, and then determine which premises in the original plan still hold true. The following cases are teaching deductions to facilitate the transfer of principles to new problems.
Changing conditions:The mission path is unknown in advance, and the next step depends on observation.
Extended question:Do I still need an outer image?
A small outer state machine can be used to manage running, waiting permissions, completion and budget, and the internal Agent can independently select read-only diagnostic tools. Each observation saves the version and writes the repair action to exit the exploration and enter the clear approval and acceptance steps. It is not necessary to write a fixed path for each error, but it is still necessary to fix which actions can change the production status.
The principles that remain unchanged:What is uncertain is the exploration path, and the execution boundary and final state can still be determined.
Changing conditions:The path is basically fixed, and the model is only responsible for explanation
Extended question:Do we still need to plan the settlement steps cyclically?
No need. Qualification, amount and transaction submission are completed by the determination process, and the model translates the verified results into user instructions; when the information is missing, it returns to a fixed state. Allowing the model to rearrange the order of settlement increases risk without any exploration benefit, and only adds restricted model nodes where material interpretation cannot be structured.
The principles that remain unchanged:Put models where real judgment is needed and business rules in auditable control processes.
It is designed based on public technical information; the reference materials support the technical mechanism, and the scenarios and scoring standards are designed by this website and do not represent the original interview questions of a certain company. New Q&A and migration cases are added for principle explanation, and source verification and case operation verification are recorded separately.
After reading, you can explain the principles, boundaries, and trade-offs against these standards. It is up to you to evaluate your mastery; if further verification is needed, complete the small tasks below.
Choose between fixed steps and autonomous exploration steps for the report generation process, and explain why.