WritingEngineering
A Redis lock with a TTL needs a fencing token the database checks
Why a Redis lock with a TTL can let two workers write the same rows, and how a fencing token checked by the store makes the late writer fail instead.
A worker picks up the settlement job for merchant m_4471: 1,200 refunds totalling 84,300.00 EUR, to be sent to the gateway one by one. It takes the key lock:settle:m_4471 with SET NX PX 30000 and starts at refund r_001. At second 26, on r_815, it stalls: a stop-the-world pause, a slow hop to the database, a container frozen mid-deploy. At second 30 Redis deletes the key. At second 31 a second worker, redelivered the same job by the queue, acquires the lock and starts on the rows still marked pending. At second 34 the first worker wakes up, believes it still holds the lock, and carries on writing.
A lock with a TTL is a liveness mechanism and nothing more: it guarantees that a worker which crashes cannot hold the merchant's settlement forever. It guarantees nothing about two workers writing the same rows at once. The only thing that gives you mutual exclusion is a fencing token, a number issued with the lock that the database checks on every write the lock is supposed to protect. A Redis lock without a fencing token is not a lock; it is a hint that usually works.
The TTL answers a different question
The TTL exists because processes die. Without it, a worker killed at r_815 leaves lock:settle:m_4471 in Redis until somebody notices, and every later run of the job fails to acquire. So the TTL is sized for recovery after a crash, and there is a tension in sizing it: 30 seconds means a dead worker blocks the merchant for 30 seconds, and a live worker that takes 31 seconds loses its lock without being told.
Two fixes look tempting and both leave the hole open. Renewing the lease from a background timer keeps a long job alive, but a paused process cannot renew, and when it resumes it has no idea that time passed. Checking GET lock:settle:m_4471 before each write narrows the window but does not close it: the key can expire between the check and the UPDATE, and under a pause that gap is exactly where the stall lands. Both fixes act on the worker's belief about the lock. Only the store can act on the truth.
What a fencing token is
The token is a counter that increases every time the lock is acquired. After SET NX PX succeeds, the holder runs INCR fence:settle:m_4471 and receives, say, 33. Because only the holder runs INCR, tokens rise in the order the lock was granted. The database keeps a fence column on every row the lock protects, and every write carries the holder's token and succeeds only if that token is at least the stored one. A late writer carrying 33 against a row already at 34 is turned away by the row itself, however sure the worker is that it still holds the lock.
flowchart TD
W[Worker] --> L[Acquire lock with TTL]
L --> T[Read token from INCR]
T --> C[Claim row with token]
C --> F{Token at least stored fence}:::accent
F -->|yes| S[Store token and apply write]
F -->|no| X[Reject and abort batch]
S --> G[Call gateway with refund id]
G --> M[Confirm with same token]
M --> FThe write that matters is the claim. A refund moves from pending to sending in one fenced UPDATE before the gateway is called, and from sending to sent in a second fenced UPDATE after it answers. The claim is what stops two workers sending the same refund; the confirm is what stops a stale worker recording a result it no longer owns.
// simplified: fenced state transitions on a refund row
async function claimRefund(db: Db, refundId: string, token: number) {
const res = await db.execute(
`UPDATE refund
SET state = 'sending', fence = $1
WHERE id = $2
AND state IN ('pending', 'sending')
AND fence < $1`,
[token, refundId],
);
if (res.rowCount === 0) throw new FenceViolation(refundId, token);
}
async function confirmRefund(db: Db, refundId: string, token: number) {
const res = await db.execute(
`UPDATE refund
SET state = 'sent'
WHERE id = $2 AND state = 'sending' AND fence = $1`,
[token, refundId],
);
if (res.rowCount === 0) throw new FenceViolation(refundId, token);
}The claim accepts a row already in sending as long as its fence is lower. That is deliberate: a row claimed under token 33 while the lock is now held under 34 belongs to a holder whose lease has ended, and the new holder is entitled to take it over. The confirm is stricter, fence = $1, because a result should only ever be recorded by the holder that claimed the row.
Acquire the lease
SET lock:settle:m_4471 NX PX 30000. On failure, stop; another holder is live or its lease has not expired yet.Take the token
INCR fence:settle:m_4471and keep the number for the whole batch. Never reuse a token across acquisitions.Claim before the side effect
Move the row tosendingwith the token in theWHEREclause. Zero rows updated means abort, not retry.Confirm with the same token
After the gateway answers, move the row tosentonly if the fence still equals your token.
The worked example, second by second
Worker A acquires the lock at 10:00:00.000 and receives token 33. It claims r_001 through r_814, calls the gateway for each, and confirms each. At 10:00:26 it claims r_815, setting fence = 33, state = sending, and then stalls before the gateway call. At 10:00:30 the key expires. At 10:00:31 worker B, handed the same job by the queue after A's visibility timeout, acquires the lock and receives token 34.
Worker B selects every refund for m_4471 whose state is not sent. It finds r_815 in sending with fence 33 and r_816 onwards in pending. Before touching r_815 it asks the gateway whether a refund with reference r_815 already exists, because a sending row means a call may have gone out. The gateway says no. B claims r_815 with token 34, calls the gateway, confirms it, and moves on.
At 10:00:34 worker A resumes. It calls the gateway for r_815 with idempotency key r_815, and the gateway returns the refund B already created rather than making a second one. A then runs confirmRefund('r_815', 33). The row's fence is 34, zero rows are updated, FenceViolation is thrown, and A abandons the batch without claiming r_816. The customer behind r_815 sees one refund of 70.25 EUR. Without the fence, A would have confirmed r_815, claimed r_816 in the same instant as B, and the two workers would have raced down the remaining 385 rows together.
sequenceDiagram participant A as Worker A participant R as Redis participant B as Worker B participant D as Database A->>R: SET lock NX PX 30000 R-->>A: OK token 33 A->>D: claim r815 fence 33 Note over A: pause of 8 seconds Note over R: key expires at 30s B->>R: SET lock NX PX 30000 R-->>B: OK token 34 B->>D: claim r815 fence 34 A->>D: confirm r815 fence 33 D-->>A: 0 rows updated
Where the token stops helping
The fence protects the rows it is stored on. It cannot protect the gateway call that sits between the claim and the confirm, which is why worker A's second call for r_815 was harmless only because the refund id was the idempotency key. Any side effect outside the fenced store needs its own deduplication: an idempotency key on the payment API, a unique reference on the email, a conditional put on the object store. I have built payment orchestration across 15+ gateways, and the rule I apply is that the fence decides who owns the row and the idempotency key decides what the outside world does about a duplicate.
The counter itself is a second weak point. If Redis restarts without the counter persisted, INCR starts again at 1 and a new holder receives a token lower than the fence already on the rows, so every claim it makes fails and the batch stalls. Seed the counter at startup from SELECT MAX(fence) FROM refund for that merchant, or issue tokens from a database sequence and keep Redis only for the lease. A batch that spans two stores has the same problem twice: each store needs its own fence column, and a write to a store without one is unprotected however careful the other store is.
What each mechanism actually guarantees
| Situation | TTL alone | TTL with fencing token |
|---|---|---|
| Worker crashes holding the lock | Lock frees after the TTL | Lock frees; next holder reclaims rows in sending with a lower fence |
| Worker pauses past the TTL | Two holders write the same rows | Late writer's UPDATE matches zero rows and aborts |
| Queue redelivers the job early | Second worker waits on SET NX, then races | Second worker gets a higher token; first is fenced out on resume |
| Redis restarts without persistence | Lock disappears, both workers proceed | Counter reset detected by seeding from the store's max fence |
The middle two rows are the ones that produce a double refund, and neither is fixed by a longer TTL. Making the TTL 5 minutes turns a 30-second race into a 5-minute outage for every merchant whose worker died, and still leaves the paused worker free to write when it wakes.
Add the fence column and the two fenced UPDATE statements before anything else; the lock acquisition and the counter are an afternoon's work once the writes refuse a stale token. Then make the gateway call idempotent on the refund id, so the one write the fence cannot see is harmless when it happens twice.