Agent Application DevelopmentAccount
Knowledge catalogChoose core direction and segmented content
knowledge unit 31AdvancedSystem designAbout 15 minutes

Understand → Implement → Debug → Design

Lease takeover and fencing stale workers

The lease is responsible for taking over, the monotonically increasing token is responsible for rejecting expired writes, and the business side effects must still be independently constrained.

WorkerleaseFencing TokenConcurrency control

Knowledge content check2026-10-03 · Check the source of the original question2026-10-02

Which step do you want to learn from this knowledge point?

Select the starting point based on the current basis, or you can go deeper one by one. When you encounter an unfamiliar concept, go back to the core principles first; use the knowledge exercises to check your understanding when you are finished.

Understand first

New to this knowledge point

Complete the prerequisite concepts, read the principles and counterexamples, and then explain why in your own words.

Start with core principles →

Implement next

Prepare to write the principles into code

Understand implementation steps and boundaries, complete small tasks, and check results against acceptance requirements.

View the code example →

Debug failures

Need to handle failures and changes in conditions

Follow the continuous questioning to locate the failure premise, and then compare the migration cases to explain how the plan should be adjusted.

Continue to delve deeper into the problem →

Compare designs

Need to design or review plans

Combine engineering deductions and senior self-evaluation standards to explain the applicable conditions, costs and alternatives of the plan.

Analyze engineering scenarios →
Knowledge unit directory

LEARN · PRACTICE · REFLECT

Knowledge learning and personal records

My notes and review ↗

First read along the principles, Q&A and migration cases. When you need to check your understanding, switch to reinforcement exercises or start personal recording.

Answers and personal notes

Each modified commit will be kept as an independent history. Your level of mastery is up to you to evaluate yourself against the standards.

Core concept · Lease takeover and fencing stale workers

Understand the core principles first

Preparatory concepts:Row level condition updates, Lease and clock, Side effects are idempotent

When the lease expires, only others are allowed to take over, but the old process cannot be stopped. When old and new executors coexist, the actual submitter must verify the execution generation to prevent old results from overwriting the new state.

Liveness and safety need different mechanisms

A lease enables takeover after a crash, but the old worker may merely be paused or partitioned and later resume. A longer TTL reduces takeover frequency without proving that the old process stopped. Redis guidance recommends fencing for long operations; a living process does not necessarily hold a valid lock.

Include generations in the commit contract

Allocate a monotonically increasing fencing token atomically during acquisition. Local commits check owner, token, lease, and task state together. Renewal must match the current token. A token represents an execution generation; an operation ID represents business intent. Takeover must not change an existing action’s idempotency key.

Enforce fencing at the resource

A pre-call lease read leaves a race before the write. A conditional database UPDATE can validate and commit together. Remote idempotency deduplicates the same action but does not automatically reject other writes from an expired worker. Highest-seen-token fencing rejects generations below the registered value. Immediate rejection requires registering the new generation at takeover or atomically checking authority at a trusted gateway.

Do not hold a lock for the whole task

Acquire work in a short transaction, releasing row locks after updating owner and lease. Keep inference and network calls outside long transactions. PostgreSQL SKIP LOCKED suits multi-consumer queue tables, not general consistent reads. Pause an old worker, let a new worker take over, then resume the old one. Verify stale commits fail at every relevant target.

Check understanding with a question

Two Workers resume an Agent task at the same time. How to use lease and Fencing Token to prevent the old Worker from writing?

Leases allow takeover without stopping a paused old worker. Allocate monotonically increasing tokens atomically; renewal and commits must match owner, token, and validity. Use short claim transactions rather than holding locks throughout inference. External targets also need appropriate fencing or idempotency protection; local checks alone cannot prevent remote effects.

Implementation and trade-offs

Expiration of the lease does not mean that the old process has stopped

Worker A picks up the task, and then a long pause occurs. The scheduler sees that the heartbeat has expired and hands the task to B; A may still continue to submit results after recovery. This is a dual execution problem and cannot be answered with "we set a Redis lock". The lease mainly provides takeover after loss of contact. The official Redis distributed lock document also indicates that long-term tasks require fencing tokens, and it cannot be assumed that the process will always hold the lock while it is alive. The security goal should be written as follows: After a new task claim has succeeded, the protected resources cannot be modified in the old round.

Atomic claims generate monotonically increasing execution generations

Task table design run_id, owner, lease_until, fence, status. Claim the task in a short database transaction, only if it is queued or its running lease has expired; final states such as canceled and failed cannot be automatically taken over; after success, the fence is increased by one and the token of this generation is returned. PostgreSQL's FOR UPDATE SKIP LOCKED is suitable for multiple consumers to claim queued tasks, but it provides the behavior of skipping locked rows and does not automatically generate leases or resolve old process resurrections. Renewal must match both owner and fence; if update fails, new tool calls will be stopped immediately. Do not keep the claim transaction open during ten minutes of task execution, otherwise it will occupy the lock for a long time and affect the takeover.

Enforce stale-write rejection at the final resource

The UPDATE submitted as a result requires conditions such as run_id, fence, owner, and unexpired lease. Zero rows affected indicate that execution rights have been lost. Token checking and writing must be completed in the same transaction or the same atomic operation. SELECT first and then unconditional UPDATE will leave a race condition. For the external ticketing system, it is not enough to add a token field to the request. The target writer must verify the generation; if it cannot be supported, use business unique key idempotency and reconciliation to limit the impact of the old Worker. Batch writes are also checked batch by batch, and long CPU calculations can continue to end, but the final submission is still constrained.

Test an old worker resuming after losing its lease

The acceptance sequence is that A claims fence=1 and is suspended until expiration, B claims fence=2, and then A submits. A is required to update zero rows, and B can submit. Then test the renewal of the old owner, concurrent claims by old and new workers, clock offset, retry submission and completion request after cancellation. Use database time consistently to evaluate leases in production, and conservatively stop writing when heartbeat fails; the wall clocks of all node machines cannot be regarded as completely consistent. The example demonstrates rejecting old commits using injected logical timing, without simulating network partitions, actual row locks, or a complete cluster.

code example

SQLite demo expired generation cannot complete the task

The Python standard library is runnable and only demonstrates local commits subject to conditions; logical time is injected by the test, and production uses a unified time source and implements renewal.

import sqlite3
c = sqlite3.connect(':memory:')
c.execute('CREATE TABLE jobs(id TEXT PRIMARY KEY, owner TEXT, until INTEGER, fence INTEGER, state TEXT)')
c.execute("INSERT INTO jobs VALUES ('r', '', 0, 0, 'ready')")
def claim(owner, now):
    with c:
        cur = c.execute("UPDATE jobs SET owner=?, until=?, fence=fence+1, state='running' WHERE id='r' AND until<=? AND state IN ('ready','running')", (owner, now+5, now))
        if cur.rowcount != 1: return None
        return c.execute("SELECT fence FROM jobs WHERE id='r'").fetchone()[0]
def finish(owner, token, now):
    with c:
        return c.execute("UPDATE jobs SET state='done' WHERE id='r' AND owner=? AND fence=? AND until>? AND state='running'", (owner, token, now)).rowcount
a = claim('A', 0)
b = claim('B', 6)
print(a, b)
print('old:', finish('A', a, 7))
print('new:', finish('B', b, 7))

expected output

1 2
old: 0
new: 1

Engineering deduction

scene
Assume an engineering scenario: the long article research task is taken over by two Workers.
design decisions
The claim generation is incremented, and article draft submission verifies the generation and valid lease at the same time; remote operations use separate idempotency keys.
Verify target
The acceptance goal is that the old Worker cannot overwrite the new Worker's draft, and the takeover status can be tracked.
applicable boundary
When the external resource does not check the token, local fencing cannot prevent old requests from being executed remotely.

Continuous questions and answers

Continue reading along with the premises and constraints of the problem. Understand the reference answers first, then try to put away the answers and explain the cause and effect and trade-offs in your own words.

Draw inferences from one example: If the conditions change, how to deduce it?

First find out the conditions for change, and then determine which premises in the original plan still hold true. The following cases are teaching deductions to facilitate the transfer of principles to new problems.

Task status is local, file object is remote

Changing conditions:Submitting will overwrite the remote report file

Extended question:The database rejects the old done, why is the file still overwritten by the old Worker?

Derivation and reference solutions

Rejecting the local commit after the side effect is too late. Remote writes need to use version conditions, target-side fencing, or write the immutable version first and then switch the reference by the trusted submitter. Old workers can produce orphan drafts, but cannot overwrite current official references; garbage collection cleans up by state. Protecting only local rows does not protect object storage.

The principles that remain unchanged:Execution right verification must cover the resource end where the effect actually occurs.

Long task lease short

Changing conditions:A model call is longer than the lease

Extended question:Is it safe to change the TTL to one hour?

Derivation and reference solutions

A longer lease reduces ordinary expirations but also delays takeover after loss of contact; a pause can still exceed an hour. Use independent heartbeat renewal and stop new actions when losing authority, and the submitter continues to check the token. If an issued write cannot be interrupted, target idempotency or conditional commits are still needed. TTL alone cannot establish safety.

The principles that remain unchanged:The lease length adjusts the liveness and fault tolerance speed, and the security of the old executor relies on submission verification.

Easy to make mistakes

  • Think TTL expiration kills old workers
  • Tokens are randomly generated rather than monotonically increasing
  • Only check token locally on Worker
  • Put the entire long task in a database lock transaction

References

It is designed based on public technical information; the reference materials support the technical mechanism, and the scenarios and scoring standards are designed by this website and do not represent the original interview questions of a certain company. New Q&A and migration cases are added for principle explanation, and source verification and case operation verification are recorded separately.

Check how far you understand

After reading, you can explain the principles, boundaries, and trade-offs against these standards. It is up to you to evaluate your mastery; if further verification is needed, complete the small tasks below.

Basic standards met
An expired lease does not mean the old worker has stopped.
Intermediate and advanced signals
Reject stale writes with incrementing tokens and conditional updates.
Senior criteria
Explain which resources support fencing tokens and what additional controls are needed for external systems that do not.

View verification records for independent examples

Continue to do advanced research experiments

After the process really exits, how can it continue?

From the idempotent counterexample of two databases, continue to verify the lease, checkpoint and independent server receipt. Unzip the reliability experiment v3 and execute it in a separate directory.

Read full text and fault analysis → · Download Reliability Experiment v3 ↓

python3 cli.py submit --db crash.sqlite
python3 cli.py run --db crash.sqlite --lease-seconds 2 --fault after_collect
python3 cli.py inspect --db crash.sqlite
# 首次运行预期退出码 75;等待至少 2 秒后分别执行
python3 cli.py run --db crash.sqlite
python3 evaluate.py --db crash.sqlite

Keep evidence and check item by item

  • First inspect shows running, collect checkpoint saved.
  • After the lease expires, succeeded, generation=2, collect is only submitted once.
  • Press README and run after_effect to check that there is only one receipt in the independent publisher database.

Fixed collect → draft → verify → publish flow; verifying local persistence protocol with real process exit does not prove that any remote service executes exactly once.

Hands-on verificationComplete on demand · Suggestions15 minutes

Give old workers pause and resume timing, marking which writes should be rejected.

Expand acceptance requirements and checkpoints
  • The new owner has a higher generation
  • Storage side execution condition check
  • You cannot rely solely on the client to determine the lease

Key inspections

  • Can reproduce the problem of old Worker resurrection after lease expiration
  • Ability to design atomic claims, renewals and conditional submissions
  • Checks that make it clear that fencing must be performed by the resource writer
  • Ability to distinguish between task preemption and side effects idempotency