Understand first
New to this knowledge point
Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.
Start with core principles →Understand → Implement → Debug → Design
Check workflow version, persistence state schema and playback compatibility.
Knowledge content check2026-10-03 · Check the source of the original question2026-10-02
It is recommended to understand first:
Idempotency, unknown outcomes, and task recovery →Select the starting point based on the current basis, or you can go deeper one by one. When you encounter an unfamiliar concept, go back to the core principles first; use the knowledge exercises to check your understanding when you are finished.
New to this knowledge point
Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.
Start with core principles →Prepare to write the principles into code
Understand implementation steps and boundaries, complete small tasks, and check results against acceptance requirements.
Reading implementation and trade-offs →Need to handle failures and changes in conditions
Follow the continuous questioning to locate the failure premise, and then compare the migration cases to explain how the plan should be adjusted.
Continue to delve deeper into the problem →Need to design or review plans
Combine engineering deductions and senior self-evaluation standards to explain the applicable conditions, costs and alternatives of the plan.
Analyze engineering scenarios →LEARN · PRACTICE · REFLECT
First read along the principles, Q&A and migration cases. When you need to check your understanding, switch to reinforcement exercises or start personal recording.
Can be practiced directly. After logging in, answers, favorites, and notes will be saved to your account.
Log in and saveEach modified commit will be kept as an independent history. Your level of mastery is up to you to evaluate yourself against the standards.
Core concept · Structural and business compatibility during workflow upgrades
Preparatory concepts:Schema and serialization, version routing, Persistence state invariants
The old state can be deserialized only for structural compatibility; safe recovery also requires that the new code retains the original steps, approvals, budgets, and tool semantics. Releases require defining version ownership and migration boundaries for outstanding runs.
Changing amount from yuan to fen retains a numeric JSON field but changes its meaning by a factor of 100. Renaming a node can remove a waiting task’s recovery target; a changed tool contract can invalidate old receipts. Record workflow, state_schema, tool_contract, and relevant configuration versions. Old tasks cannot automatically use the newest semantics.
Optional fields and defaults suit small changes. Semantic changes need deterministic migrations; hard-to-migrate runs may finish under old executors. LangGraph documents different topology restrictions for completed and interrupted threads. Renaming a state key loses its saved value, and incompatible types can fail. Framework allowance does not establish business compatibility; approval and operation identity still require checks.
Read an original snapshot and emit new state plus a migration record, without external writes or repeated budget consumption. Conditional source_version writes and atomic snapshot switching protect retries and crashes; retain the original until success. Fixtures cover running, waiting approval, unknown, and completed states. Preserve receipts, cancellation, remaining budget, and approval bindings.
Code rollback neither reverses new state nor undoes external effects. Expand compatibility before a dual-read phase and later contraction, or let new-version executors finish their runs while pausing new admissions. Inventory unfinished runs before removing old versions. A canary can pass new tasks but repeat an action when resuming a three-day-old approval. Include recovery of old runs in release validation.
Persisted checkpoints do not guarantee future code understands their meaning. Record workflow, schema, tool, and configuration versions. Compatible runs can resume; incompatible ones need deterministic migration or old executors. Retain original snapshots and verify invariants. Models cannot guess old field semantics. Rollback must account for new-state compatibility.
Node renaming, status field splitting, enumeration meaning changes, and tool result format changes may affect recovery. Even though JSON can still be deserialized, the old amount unit or status semantics may no longer apply. First define the workflow version and status Schema version. The version is fixed when the task is created. The deployment cannot silently map all old tasks to the latest logic.
Small compatibility changes adopt new optional fields and default values; write deterministic migrations when the meaning needs to change; high-risk long-term tasks can be left to run on the old worker until completion. The migration function inputs the old state, outputs the new state and migration records, and does not perform business side effects. Explicitly reject unknown fields and unsupported versions instead of continuing with null values.
Build recovery fixtures from masked snapshots, including running, awaiting approval, unknown results, and completed tasks. Verification step completion records, budget consumption, approval binding and tool receipts are not lost. The migration itself needs to be repeatable or only executed once through version conditions. Keep the old state until the new state is saved successfully to avoid being unable to recover after the migration is interrupted.
Code rollback does not automatically reverse data migration. Explain in advance which versions can be read in both directions, which ones must continue to be terminated by new workers, and how to pause new tasks. Statistics of unfinished runs before release are distributed by version, and old executors are removed only after clearing or migrating. The interviewer should bring up the realistic constraint that "there are tasks still running when incompatible changes are released."
Continue reading along with the premises and constraints of the problem. Understand the reference answers first, then try to put away the answers and explain the cause and effect and trade-offs in your own words.
Level 1What will be the impact of changing the node name?
Enter the compatibility of the topology recovery address from the status field compatibility.
A waiting task's saved next-node pointer may still refer to the old name. Renaming or deleting it can leave recovery without a target or send it down the wrong path. Framework migration depends on thread state: retain alias adapters or route by the old workflow version before migrating the saved execution position. Changes in node semantics also require business-version checks; giving a new node the old name is insufficient.
Level 1What should I do if the migration execution is down halfway?
The migration process itself also needs to be revertive and idempotent.
Small snapshot migrations are committed atomically in local transactions, and the new state and migrated version are successful at the same time; large-scale migrations are processed one by one, using unique migration IDs and conditional writes to record progress. Unfinished items will be redone after a crash, and completed items will return to the existing results. The original snapshot is retained until the verification is completed, and the migration function does not contain business side effects.
Follow this answer further
Level 2Migration requires changing three tables. Can I save them in three times?
After solving the problem of single snapshot atomicity, multi-table constraints expand transaction boundaries.
If three tables jointly express approval, budget, and execution position, split exposure will create an inconsistent state. If a transaction can be placed, submit it together; when atomic submission across storage is not possible, first write a new version of invisible data, then use a trusted pointer to switch, and record compensation or retry progress. Migration intermediate states cannot be treated as resumable tasks by Workers.
Follow this answer further
Level 3The new status has been switched, but the verification found that the approval was lost. Can you add approved=true?
Validation failures after the switch touch authorization invariants, requiring fact repair rather than field completion.
It cannot be inferred that an approval already exists by virtue of the run ever waiting. Go back to the retained original snapshot and approval ledger to check the actual approval; if it cannot be proved, block recovery and re-approval. Fix to preserve actor, action versions and validity periods, and verify other invariants. Migration errors cannot be masked with a default success status.
Level 1Why is code rollback not equal to state rollback?
Release rollback involves running the data rather than just deploying the package.
The data may have changed into a structure or semantics that the old code does not understand, or the old code may misinterpret new fields. Verify the read compatibility matrix before rolling back; after one-way migration, keep the new executor to finish or pause the affected tasks without forcing loading. External completed actions are recorded independently and cannot be retried as if they have not been executed through rollback.
First find out the conditions for change, and then determine which premises in the original plan still hold true. The following cases are teaching deductions to facilitate the transfer of principles to new problems.
Changing conditions:Structure expansion without changing business steps
Extended question:Do you want all old tasks to be migrated again?
If the default value is clear and does not affect approval, budget, or control flow, you can read it by default without expensive full migration; still confirm that serializers, validators, and legacy code tolerate the field. Unknown observation values cannot be filled in as "Verification Successful". Compatible changes are also versioned for tracking purposes.
The principles that remain unchanged:Migration costs vary with semantic changes, and structural extensions cannot create false facts.
Changing conditions:The field structure remains unchanged, but the side effects change
Extended question:Can old snapshots still be directly entered into the node with the same name?
Compatibility cannot be determined based on the same name and type. The original task may only authorize reading, and the new logic adds external write actions. The old path should be maintained or reviewed and approved again. Record the tool semantic version and the process behavior version separately; migration cannot secretly expand the scope of operations.
The principles that remain unchanged:Structural compatibility does not replace behavior and authorization compatibility, and persistent state cannot update user intent.
It is designed based on public technical information; the reference materials support the technical mechanism, and the scenarios and scoring standards are designed by this website and do not represent the original interview questions of a certain company. New Q&A and migration cases are added for principle explanation, and source verification and case operation verification are recorded separately.
After reading, you can explain the principles, boundaries, and trade-offs against these standards. It is up to you to evaluate your mastery; if further verification is needed, complete the small tasks below.
Split design v1 to v2 migration for waiting state, retaining original approvals and budget.