Understand first
New to this knowledge point
Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.
Start with core principles →Understand → Implement → Debug → Design
The lease is responsible for taking over, the monotonically increasing token is responsible for rejecting expired writes, and the business side effects must still be independently constrained.
Knowledge content check2026-10-03 · Check the source of the original question2026-10-02
It is recommended to understand first:
Tool calls: structure, authorization, and business contracts →Select the starting point based on the current basis, or you can go deeper one by one. When you encounter an unfamiliar concept, go back to the core principles first; use the knowledge exercises to check your understanding when you are finished.
New to this knowledge point
Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.
Start with core principles →Prepare to write the principles into code
Understand implementation steps and boundaries, complete small tasks, and check results against acceptance requirements.
View the code example →Need to handle failures and changes in conditions
Follow the continuous questioning to locate the failure premise, and then compare the migration cases to explain how the plan should be adjusted.
Continue to delve deeper into the problem →Need to design or review plans
Combine engineering deductions and senior self-evaluation standards to explain the applicable conditions, costs and alternatives of the plan.
Analyze engineering scenarios →LEARN · PRACTICE · REFLECT
First read along the principles, Q&A and migration cases. When you need to check your understanding, switch to reinforcement exercises or start personal recording.
Can be practiced directly. After logging in, answers, favorites, and notes will be saved to your account.
Log in and saveEach modified commit will be kept as an independent history. Your level of mastery is up to you to evaluate yourself against the standards.
Core concept · Lease takeover and fencing stale workers
Preparatory concepts:Row level condition updates, Lease and clock, Side effects are idempotent
When the lease expires, only others are allowed to take over, but the old process cannot be stopped. When old and new executors coexist, the actual submitter must verify the execution generation to prevent old results from overwriting the new state.
A lease enables takeover after a crash, but the old worker may merely be paused or partitioned and later resume. A longer TTL reduces takeover frequency without proving that the old process stopped. Redis guidance recommends fencing for long operations; a living process does not necessarily hold a valid lock.
Allocate a monotonically increasing fencing token atomically during acquisition. Local commits check owner, token, lease, and task state together. Renewal must match the current token. A token represents an execution generation; an operation ID represents business intent. Takeover must not change an existing action’s idempotency key.
A pre-call lease read leaves a race before the write. A conditional database UPDATE can validate and commit together. Remote idempotency deduplicates the same action but does not automatically reject other writes from an expired worker. Highest-seen-token fencing rejects generations below the registered value. Immediate rejection requires registering the new generation at takeover or atomically checking authority at a trusted gateway.
Acquire work in a short transaction, releasing row locks after updating owner and lease. Keep inference and network calls outside long transactions. PostgreSQL SKIP LOCKED suits multi-consumer queue tables, not general consistent reads. Pause an old worker, let a new worker take over, then resume the old one. Verify stale commits fail at every relevant target.
Leases allow takeover without stopping a paused old worker. Allocate monotonically increasing tokens atomically; renewal and commits must match owner, token, and validity. Use short claim transactions rather than holding locks throughout inference. External targets also need appropriate fencing or idempotency protection; local checks alone cannot prevent remote effects.
Worker A picks up the task, and then a long pause occurs. The scheduler sees that the heartbeat has expired and hands the task to B; A may still continue to submit results after recovery. This is a dual execution problem and cannot be answered with "we set a Redis lock". The lease mainly provides takeover after loss of contact. The official Redis distributed lock document also indicates that long-term tasks require fencing tokens, and it cannot be assumed that the process will always hold the lock while it is alive. The security goal should be written as follows: After a new task claim has succeeded, the protected resources cannot be modified in the old round.
Task table design run_id, owner, lease_until, fence, status. Claim the task in a short database transaction, only if it is queued or its running lease has expired; final states such as canceled and failed cannot be automatically taken over; after success, the fence is increased by one and the token of this generation is returned. PostgreSQL's FOR UPDATE SKIP LOCKED is suitable for multiple consumers to claim queued tasks, but it provides the behavior of skipping locked rows and does not automatically generate leases or resolve old process resurrections. Renewal must match both owner and fence; if update fails, new tool calls will be stopped immediately. Do not keep the claim transaction open during ten minutes of task execution, otherwise it will occupy the lock for a long time and affect the takeover.
The UPDATE submitted as a result requires conditions such as run_id, fence, owner, and unexpired lease. Zero rows affected indicate that execution rights have been lost. Token checking and writing must be completed in the same transaction or the same atomic operation. SELECT first and then unconditional UPDATE will leave a race condition. For the external ticketing system, it is not enough to add a token field to the request. The target writer must verify the generation; if it cannot be supported, use business unique key idempotency and reconciliation to limit the impact of the old Worker. Batch writes are also checked batch by batch, and long CPU calculations can continue to end, but the final submission is still constrained.
The acceptance sequence is that A claims fence=1 and is suspended until expiration, B claims fence=2, and then A submits. A is required to update zero rows, and B can submit. Then test the renewal of the old owner, concurrent claims by old and new workers, clock offset, retry submission and completion request after cancellation. Use database time consistently to evaluate leases in production, and conservatively stop writing when heartbeat fails; the wall clocks of all node machines cannot be regarded as completely consistent. The example demonstrates rejecting old commits using injected logical timing, without simulating network partitions, actual row locks, or a complete cluster.
The Python standard library is runnable and only demonstrates local commits subject to conditions; logical time is injected by the test, and production uses a unified time source and implements renewal.
import sqlite3
c = sqlite3.connect(':memory:')
c.execute('CREATE TABLE jobs(id TEXT PRIMARY KEY, owner TEXT, until INTEGER, fence INTEGER, state TEXT)')
c.execute("INSERT INTO jobs VALUES ('r', '', 0, 0, 'ready')")
def claim(owner, now):
with c:
cur = c.execute("UPDATE jobs SET owner=?, until=?, fence=fence+1, state='running' WHERE id='r' AND until<=? AND state IN ('ready','running')", (owner, now+5, now))
if cur.rowcount != 1: return None
return c.execute("SELECT fence FROM jobs WHERE id='r'").fetchone()[0]
def finish(owner, token, now):
with c:
return c.execute("UPDATE jobs SET state='done' WHERE id='r' AND owner=? AND fence=? AND until>? AND state='running'", (owner, token, now)).rowcount
a = claim('A', 0)
b = claim('B', 6)
print(a, b)
print('old:', finish('A', a, 7))
print('new:', finish('B', b, 7))
expected output
1 2
old: 0
new: 1Continue reading along with the premises and constraints of the problem. Understand the reference answers first, then try to put away the answers and explain the cause and effect and trade-offs in your own words.
Level 1If the lease is still valid but the user has canceled it, how can the submission conditions be extended?
After obtaining execution rights, the user's cancellation can also change the submission conditions.
In addition to owner, token, and lease_until, a conditional commit checks the run state, cancel_epoch, or input revision. Cancellation and commit must have a defined order in trusted storage: if cancellation succeeds first, reject subsequent commits; external actions already committed enter reconciliation rather than disappearing. Stopping scheduling is insufficient if a running worker can still write done.
Follow this answer further
Level 2The cancellation was written to the database first, but the old Worker still issued remote requests. Can I just reject the final submission?
After the parent asks to expand the local condition, the remote action may still cross the boundary.
Do not treat local final state rejection as if it did not occur at the remote end. Call the sending gateway to check the cancellation version and token again. If the target supports it, use the generation or operation key to verify; the request that has been sent is saved as pending verification. The final state of the product distinguishes between cancellation requests, stopping new actions and effects that have occurred to avoid giving false full rollbacks.
Follow this answer further
Level 3If the remote end does not recognize the token, can the idempotency key completely replace fencing?
When the remote protection capability is insufficient, action identities and execution rights identities are further distinguished.
No. Idempotent keys protect the same operation from being repeated, and the old worker may still write if it generates different actions or parameters. Approved intents can be pinned, using a trusted gateway to dispatch only allowed actions in the current generation, or using target business version conditions. The absence of these capabilities can only limit impact and reconciliation, and clearly cannot prevent all expired side effects.
Level 1Should each batch advance the fencing token, and how does that differ from the acquisition generation?
In addition to anti-aging executors, distinguish execution generations and business progress.
Fences usually increase as execution rights are reacquired or taken over, expressing generations; consecutive batches of the same Worker can use independent steps or revisions to record progress. Adding fences to each batch is a more fine-grained design, but it must be recognized consistently by downstream users and cannot be added arbitrarily to cause old requests to be rejected. operation_id remains stable for retries of the same business action.
Level 1How to avoid being taken over by other workers between SELECT and UPDATE in PostgreSQL?
Put the abstract atomic retrieval into the database concurrency mechanism.
In the same short transaction, SELECT FOR UPDATE SKIP LOCKED selects the claimable rows, then updates the owner, increments the token, sets the lease, and finally commits and returns the token; or use the UPDATE RETURNING scheme with atomic conditions. You cannot do a lock-free SELECT followed by another UPDATE. After committing the claim transaction and executing the work, the commit still uses conditional updates to verify that the rights have not expired.
First find out the conditions for change, and then determine which premises in the original plan still hold true. The following cases are teaching deductions to facilitate the transfer of principles to new problems.
Changing conditions:Submitting will overwrite the remote report file
Extended question:The database rejects the old done, why is the file still overwritten by the old Worker?
Rejecting the local commit after the side effect is too late. Remote writes need to use version conditions, target-side fencing, or write the immutable version first and then switch the reference by the trusted submitter. Old workers can produce orphan drafts, but cannot overwrite current official references; garbage collection cleans up by state. Protecting only local rows does not protect object storage.
The principles that remain unchanged:Execution right verification must cover the resource end where the effect actually occurs.
Changing conditions:A model call is longer than the lease
Extended question:Is it safe to change the TTL to one hour?
A longer lease reduces ordinary expirations but also delays takeover after loss of contact; a pause can still exceed an hour. Use independent heartbeat renewal and stop new actions when losing authority, and the submitter continues to check the token. If an issued write cannot be interrupted, target idempotency or conditional commits are still needed. TTL alone cannot establish safety.
The principles that remain unchanged:The lease length adjusts the liveness and fault tolerance speed, and the security of the old executor relies on submission verification.
It is designed based on public technical information; the reference materials support the technical mechanism, and the scenarios and scoring standards are designed by this website and do not represent the original interview questions of a certain company. New Q&A and migration cases are added for principle explanation, and source verification and case operation verification are recorded separately.
After reading, you can explain the principles, boundaries, and trade-offs against these standards. It is up to you to evaluate your mastery; if further verification is needed, complete the small tasks below.
View verification records for independent examples
From the idempotent counterexample of two databases, continue to verify the lease, checkpoint and independent server receipt. Unzip the reliability experiment v3 and execute it in a separate directory.
Read full text and fault analysis → · Download Reliability Experiment v3 ↓
python3 cli.py submit --db crash.sqlite
python3 cli.py run --db crash.sqlite --lease-seconds 2 --fault after_collect
python3 cli.py inspect --db crash.sqlite
# 首次运行预期退出码 75;等待至少 2 秒后分别执行
python3 cli.py run --db crash.sqlite
python3 evaluate.py --db crash.sqliteFixed collect → draft → verify → publish flow; verifying local persistence protocol with real process exit does not prove that any remote service executes exactly once.
Give old workers pause and resume timing, marking which writes should be rejected.