WritingEngineering

A Redis lock with a TTL needs a fencing token the database checks

Why a Redis lock with a TTL can let two workers write the same rows, and how a fencing token checked by the store makes the late writer fail instead.

A worker picks up the settlement job for merchant m_4471: 1,200 refunds totalling 84,300.00 EUR, to be sent to the gateway one by one. It takes the key lock:settle:m_4471 with SET NX PX 30000 and starts at refund r_001. At second 26, on r_815, it stalls: a stop-the-world pause, a slow hop to the database, a container frozen mid-deploy. At second 30 Redis deletes the key. At second 31 a second worker, redelivered the same job by the queue, acquires the lock and starts on the rows still marked pending. At second 34 the first worker wakes up, believes it still holds the lock, and carries on writing.

A lock with a TTL is a liveness mechanism and nothing more: it guarantees that a worker which crashes cannot hold the merchant's settlement forever. It guarantees nothing about two workers writing the same rows at once. The only thing that gives you mutual exclusion is a fencing token, a number issued with the lock that the database checks on every write the lock is supposed to protect. A Redis lock without a fencing token is not a lock; it is a hint that usually works.

The TTL answers a different question

The TTL exists because processes die. Without it, a worker killed at r_815 leaves lock:settle:m_4471 in Redis until somebody notices, and every later run of the job fails to acquire. So the TTL is sized for recovery after a crash, and there is a tension in sizing it: 30 seconds means a dead worker blocks the merchant for 30 seconds, and a live worker that takes 31 seconds loses its lock without being told.

Two fixes look tempting and both leave the hole open. Renewing the lease from a background timer keeps a long job alive, but a paused process cannot renew, and when it resumes it has no idea that time passed. Checking GET lock:settle:m_4471 before each write narrows the window but does not close it: the key can expire between the check and the UPDATE, and under a pause that gap is exactly where the stall lands. Both fixes act on the worker's belief about the lock. Only the store can act on the truth.

What a fencing token is

The token is a counter that increases every time the lock is acquired. After SET NX PX succeeds, the holder runs INCR fence:settle:m_4471 and receives, say, 33. Because only the holder runs INCR, tokens rise in the order the lock was granted. The database keeps a fence column on every row the lock protects, and every write carries the holder's token and succeeds only if that token is at least the stored one. A late writer carrying 33 against a row already at 34 is turned away by the row itself, however sure the worker is that it still holds the lock.

flowchart TD
  W[Worker] --> L[Acquire lock with TTL]
  L --> T[Read token from INCR]
  T --> C[Claim row with token]
  C --> F{Token at least stored fence}:::accent
  F -->|yes| S[Store token and apply write]
  F -->|no| X[Reject and abort batch]
  S --> G[Call gateway with refund id]
  G --> M[Confirm with same token]
  M --> F
Every write carries the token and the store, not the worker, decides whether it is current

The write that matters is the claim. A refund moves from pending to sending in one fenced UPDATE before the gateway is called, and from sending to sent in a second fenced UPDATE after it answers. The claim is what stops two workers sending the same refund; the confirm is what stops a stale worker recording a result it no longer owns.

TypeScript
// simplified: fenced state transitions on a refund row
async function claimRefund(db: Db, refundId: string, token: number) {
  const res = await db.execute(
    `UPDATE refund
        SET state = 'sending', fence = $1
      WHERE id = $2
        AND state IN ('pending', 'sending')
        AND fence < $1`,
    [token, refundId],
  );
  if (res.rowCount === 0) throw new FenceViolation(refundId, token);
}

async function confirmRefund(db: Db, refundId: string, token: number) {
  const res = await db.execute(
    `UPDATE refund
        SET state = 'sent'
      WHERE id = $2 AND state = 'sending' AND fence = $1`,
    [token, refundId],
  );
  if (res.rowCount === 0) throw new FenceViolation(refundId, token);
}

The claim accepts a row already in sending as long as its fence is lower. That is deliberate: a row claimed under token 33 while the lock is now held under 34 belongs to a holder whose lease has ended, and the new holder is entitled to take it over. The confirm is stricter, fence = $1, because a result should only ever be recorded by the holder that claimed the row.

  1. Acquire the lease

    SET lock:settle:m_4471 NX PX 30000. On failure, stop; another holder is live or its lease has not expired yet.
  2. Take the token

    INCR fence:settle:m_4471 and keep the number for the whole batch. Never reuse a token across acquisitions.
  3. Claim before the side effect

    Move the row to sending with the token in the WHERE clause. Zero rows updated means abort, not retry.
  4. Confirm with the same token

    After the gateway answers, move the row to sent only if the fence still equals your token.

The worked example, second by second

Worker A acquires the lock at 10:00:00.000 and receives token 33. It claims r_001 through r_814, calls the gateway for each, and confirms each. At 10:00:26 it claims r_815, setting fence = 33, state = sending, and then stalls before the gateway call. At 10:00:30 the key expires. At 10:00:31 worker B, handed the same job by the queue after A's visibility timeout, acquires the lock and receives token 34.

Worker B selects every refund for m_4471 whose state is not sent. It finds r_815 in sending with fence 33 and r_816 onwards in pending. Before touching r_815 it asks the gateway whether a refund with reference r_815 already exists, because a sending row means a call may have gone out. The gateway says no. B claims r_815 with token 34, calls the gateway, confirms it, and moves on.

At 10:00:34 worker A resumes. It calls the gateway for r_815 with idempotency key r_815, and the gateway returns the refund B already created rather than making a second one. A then runs confirmRefund('r_815', 33). The row's fence is 34, zero rows are updated, FenceViolation is thrown, and A abandons the batch without claiming r_816. The customer behind r_815 sees one refund of 70.25 EUR. Without the fence, A would have confirmed r_815, claimed r_816 in the same instant as B, and the two workers would have raced down the remaining 385 rows together.

sequenceDiagram
  participant A as Worker A
  participant R as Redis
  participant B as Worker B
  participant D as Database
  A->>R: SET lock NX PX 30000
  R-->>A: OK token 33
  A->>D: claim r815 fence 33
  Note over A: pause of 8 seconds
  Note over R: key expires at 30s
  B->>R: SET lock NX PX 30000
  R-->>B: OK token 34
  B->>D: claim r815 fence 34
  A->>D: confirm r815 fence 33
  D-->>A: 0 rows updated
The pause lands between claim and gateway call, and the row rejects the late confirm

Where the token stops helping

The fence protects the rows it is stored on. It cannot protect the gateway call that sits between the claim and the confirm, which is why worker A's second call for r_815 was harmless only because the refund id was the idempotency key. Any side effect outside the fenced store needs its own deduplication: an idempotency key on the payment API, a unique reference on the email, a conditional put on the object store. I have built payment orchestration across 15+ gateways, and the rule I apply is that the fence decides who owns the row and the idempotency key decides what the outside world does about a duplicate.

The counter itself is a second weak point. If Redis restarts without the counter persisted, INCR starts again at 1 and a new holder receives a token lower than the fence already on the rows, so every claim it makes fails and the batch stalls. Seed the counter at startup from SELECT MAX(fence) FROM refund for that merchant, or issue tokens from a database sequence and keep Redis only for the lease. A batch that spans two stores has the same problem twice: each store needs its own fence column, and a write to a store without one is unprotected however careful the other store is.

What each mechanism actually guarantees

SituationTTL aloneTTL with fencing token
Worker crashes holding the lockLock frees after the TTLLock frees; next holder reclaims rows in sending with a lower fence
Worker pauses past the TTLTwo holders write the same rowsLate writer's UPDATE matches zero rows and aborts
Queue redelivers the job earlySecond worker waits on SET NX, then racesSecond worker gets a higher token; first is fenced out on resume
Redis restarts without persistenceLock disappears, both workers proceedCounter reset detected by seeding from the store's max fence

The middle two rows are the ones that produce a double refund, and neither is fixed by a longer TTL. Making the TTL 5 minutes turns a 30-second race into a 5-minute outage for every merchant whose worker died, and still leaves the paused worker free to write when it wakes.

Add the fence column and the two fenced UPDATE statements before anything else; the lock acquisition and the counter are an afternoon's work once the writes refuse a stale token. Then make the gateway call idempotent on the refund id, so the one write the fence cannot see is harmless when it happens twice.

Written by Md Nasim Anjum, senior full-stack engineer in Manchester. He builds payment orchestration, KYC and KYB compliance platforms and conversational AI.

Get in touchAll writingRSS

More writing