WritingEngineering

Requeue, dead-letter or retry table for a message that keeps failing

Decide where a RabbitMQ message goes after its handler throws: back on the queue, into TTL retry queues or into a retry table, and set the delays and the cap.

A consumer takes shipment.requested for order 48213 off the work queue, calls the carrier's rate API, and gets a 503 after two seconds. The handler throws. What the consumer does with that message in the next ten milliseconds decides whether the outage costs one delayed label or a stuck queue and a retry storm aimed at a service that is already down. Anyone running RabbitMQ consumers meets this choice the first time a dependency fails in production, and most settle it by accident, by keeping the client library's default.

A failed message can go to one of three places: back onto the same queue (requeue), into a retry queue whose TTL dead-letters it back when the delay is up, or into a database table that a scheduler drains. My position is that requeue is never the right answer to a handler error, and that the real decision is between the other two, settled by one question: does the next attempt need anything besides a delay and a cap? If not, retry queues. If a person or a condition decides when it tries again, a retry table.

What decides it

CriterionRequeueRetry queues with TTLRetry table
Where the message waitsNear the head of the work queueA retry.<delay> queueA message_retries row
Delay before the next attemptNone; redelivered within millisecondsFixed per queue (5s, 1m, 10m)Any value in next_attempt_at, or null
Where the attempt count livesNowhere on a classic queue; only a redelivered flagThe x-death header the broker writesAn attempt column
What one bad message blocksEvery message behind it at prefetch 1NothingNothing
Pausing retries for one tenant or causeImpossibleImpossible without new queuesOne UPDATE setting paused = true

The rows that settle most cases are the delay and the pause. If every failure you expect clears within minutes with no action from you, a delay is all a retry needs, and the broker can hold it. If some wait on a human (a lapsed carrier contract) or on a condition (the carrier's health check turning green), the retry needs a value that is not a clock, and the broker has nowhere to keep it.

How requeue fails, step by step

Take the 503 above with ch.nack(msg) and the client default of requeue: true. The broker puts the message back near its original position, and because the consumer has prefetch 1 and is the only one subscribed, it receives the same message again within milliseconds. It calls the carrier, waits two seconds, gets another 503, nacks again: thirty attempts a minute from one consumer for as long as the outage lasts, each a stack trace in the logs.

Meanwhile every shipment.requested published after order 48213 sits behind it. That customer sees an order confirmed and no tracking number, and so does every customer after them. Nothing in the broker's metrics says poison: ready messages climb, unacked sits at 1, consumer utilisation reads 100 per cent, which looks busy rather than broken. Containment is one flag: the message leaves the work queue on the first failure, with requeue: false or with a publish followed by an ack. I work in TypeScript and NestJS with RabbitMQ, and the rule I apply is that a handler either acks after its side effect or moves the message somewhere it can wait, never neither.

Requeue

Best for A consumer shutting down that must hand back messages it never attempted

nack with requeue: true returns the message to the queue and the broker redelivers it.

  • strengths: no topology and no code; the message keeps its place in the queue; right for a graceful shutdown
  • costs: redelivery within milliseconds makes a hot loop; a classic queue keeps no attempt count; one bad message holds every message behind it at low prefetch

Retry queues with TTL

Best for Transient errors where a wait of seconds to minutes is the only thing that will change

The consumer publishes the message to retry.5s, retry.1m or retry.10m by attempt number and acks the original; each retry queue carries x-message-ttl and an x-dead-letter-exchange pointing back at the work exchange.

  • strengths: the broker does the waiting and the counting through x-death; survives restarts and deploys; no scheduler; backlog per delay visible as queue depth
  • costs: delays are fixed at declaration, so a new backoff step is a new queue; nothing can be paused or edited while a message waits; a crash between publish and ack re-sends the message once

Retry table

Best for Retries a person inspects, or that wait for a condition rather than a clock

The consumer inserts a row with the payload, attempt, next_attempt_at and last_error, acks the original, and a scheduler republishes due rows through the outbox.

  • strengths: next_attempt_at can be any time, or null until an operator or a health check sets it; pause by tenant, error class or dependency with one update; the stuck set is a query, not a queue you drain to look at
  • costs: a second queue in the database, with its own scheduler, lock (FOR UPDATE SKIP LOCKED) and alert on the oldest due row; the payload is stored twice; message and correlation ids must be copied onto the new publish by hand

One queue per delay, never per-message expiry

RabbitMQ lets you set an expiration on each message, so a single retry queue with per-message TTLs looks like it gives any backoff you like. It does not, because the broker only expires messages from the head of a queue: a ten-minute message at the head holds the five-second message behind it for the full ten minutes.

flowchart TD
  W[Work queue] --> C[Consumer]
  C -->|ack after success| D[Done]
  C -->|attempt 1| R1[Retry 5s queue]
  C -->|attempt 2| R2[Retry 1m queue]
  C -->|attempt 3| R3[Retry 10m queue]
  R1 -->|TTL expires| X[Work exchange]
  R2 -->|TTL expires| X
  R3 -->|TTL expires| X
  X --> W
  C -->|cap hit or not retryable| P[Parked queue]:::accent
Retry topology with a queue per delay and a parked queue for the cap

Declare one queue per delay, each with its own x-message-ttl, and choose the queue from the attempt number. That number comes from the x-death header, which the broker appends every time it dead-letters the message, one entry per queue and reason with a count. Sum the counts for the retry queues and you have the delays already served.

TypeScript
// simplified: route a failed message by how many retries it has served
const RETRY_QUEUES = ['retry.5s', 'retry.1m', 'retry.10m'];

type Death = { queue: string; reason: string; count: number };

function retriesServed(msg: ConsumeMessage): number {
  const deaths = (msg.properties.headers?.['x-death'] ?? []) as Death[];
  return deaths
    .filter((d) => d.reason === 'expired' && d.queue.startsWith('retry.'))
    .reduce((n, d) => n + Number(d.count), 0);
}

async function onFailure(ch: Channel, msg: ConsumeMessage, err: Error) {
  const served = retriesServed(msg);
  const headers = { ...msg.properties.headers, 'x-last-error': err.message };
  const target = !isRetryable(err) || served >= RETRY_QUEUES.length
    ? 'parked'
    : RETRY_QUEUES[served];
  ch.publish('', target, msg.content, { ...msg.properties, headers });
  ch.ack(msg); // a crash between publish and ack redelivers once; stay idempotent
}

The parked queue takes the cap and the non-retryable errors, a missing address.country field for instance, and needs an alert on its depth. A parked queue nobody reads is a slower hot loop: the message stops costing CPU and starts costing a customer.

Where the clock is the wrong input

Now let the carrier be down for ninety minutes. With retry queues the message is tried at five seconds, one minute and ten minutes, then parked about eleven minutes in. When the carrier returns there are 1,400 parked messages (an example figure) to shovel back onto the work queue by hand, and they all hit the carrier at once. The cap did what it was told, and what it was told was wrong for this failure.

With a retry table the flow differs at the second step. The consumer inserts the row with dependency = 'carrier' and next_attempt_at ten minutes out, and a health check that has seen three straight failures sets paused = true on every row for that dependency. When the check turns green it runs one update, next_attempt_at = now(), paused = false, and the scheduler drains the rows at its own pace, fifty a second say, rather than in one burst. The table's own failure is a scheduler that stops: rows pile up silently, so alert on the age of the oldest row whose next_attempt_at has passed.

flowchart TD
  S[Handler threw] --> Q1{Same payload would fail again}
  Q1 -->|yes| P[Park with the last error]
  Q1 -->|no| Q2{A delay alone fixes it}:::accent
  Q2 -->|no| T[Retry table]
  Q2 -->|yes| Q3{Must pause or edit retries}
  Q3 -->|yes| T
  Q3 -->|no| R[Retry queues with TTL]
Which home a failed message should get

Whichever side wins, build the parked destination first and alert on it, because both designs end there when retries run out. Then make the handler idempotent on the message id, since both can re-send a message in the gap between publishing it onwards and acknowledging the original. Only then is it worth arguing about queues versus rows.

Written by Md Nasim Anjum, senior full-stack engineer in Manchester. He builds payment orchestration, KYC and KYB compliance platforms and conversational AI.

Get in touchAll writingRSS

More writing