WritingPayments

A gateway's yes to a refund is an acceptance, not a result

Model a refund as its own lifecycle under the capture it reverses, so one that fails days after the gateway accepted it is caught, re-sent and explained.

On Monday a customer pays 84.00 EUR for a jacket and the capture settles. On Thursday they return it, and a support agent clicks refund for the full 84.00 EUR. The gateway answers inside 300 ms: HTTP 200, a refund id, status pending. The agent tells the customer the money is on its way, the order page flips to Refunded, and the ledger posts 84.00 EUR out. Nine days later a refund.failed webhook arrives with the reason account_closed: the card the jacket was bought with no longer exists. Nobody is watching that webhook, and the books say the money left.

The error is treating the gateway's synchronous yes as the outcome. For a payment, the authorisation response is the result: the customer is on the page, the issuer has approved or declined, and you act on it at once. For a refund, the synchronous response only means the gateway accepted the instruction to try. The result arrives later, often days later, as an event, and it can be a failure. So a refund needs its own lifecycle, hung under the capture it reverses, in which accepted and succeeded are different states, and everything you show the customer or post to the ledger is keyed to that difference.

Why a refund fails differently from a payment

A payment is a question put to the issuer while the customer waits. The issuer checks the account, the card's status and its risk rules, and answers in seconds; if the answer is no, the customer tries another card. A refund is an instruction travelling the other way with nobody waiting at the end. The gateway queues it, the acquirer submits it in a later clearing run, and the issuer applies it when that run arrives. Nothing re-checks the card at the moment the agent clicks.

That gap is where the differences live. A card closed since the purchase, a merchant balance that cannot cover the refund, a scheme time limit on refunding an old capture, and a capture charged back in the meantime all surface after the yes. And a refund has no authorisation you can hold and then release: once accepted it is in flight, and the only reliable way to undo a wrong one is to charge the customer again.

The refund lifecycle

A refund row starts as requested the moment the agent asks, before any call to the gateway, so the amount is reserved against the capture even if the process dies on the next line. It becomes submitted when the request goes out. A synchronous refusal (a 4xx, an amount the gateway will not take, a capture it no longer recognises) ends it as rejected. A synchronous yes makes it accepted. A timeout makes it unknown, which is not a failure and must not be retried with a fresh request.

stateDiagram-v2
  [*] --> Requested
  Requested --> Submitted
  Submitted --> Rejected
  Submitted --> Accepted:::accent
  Submitted --> Unknown
  Unknown --> Accepted
  Unknown --> Rejected
  Accepted --> Succeeded
  Accepted --> Failed
  Succeeded --> [*]
  Failed --> [*]
  Rejected --> [*]
Accepted is the state a refund spends days in, and the one most systems skip

From accepted, only a gateway event moves the row: refund.succeeded (some gateways say settled or completed) or refund.failed. No agent action, no timer and no customer complaint may move a refund out of accepted, because nothing except the gateway knows whether the money moved. From unknown, the orchestrator replays the same request with the same idempotency key, or fetches the refund by its own reference, until the gateway says which of accepted or rejected applies.

The refundable amount is derived on every request

The capture is for 84.00 EUR. An agent refunds 30.00 EUR for a damaged sleeve; the gateway accepts. A week later a second agent, handling the return of the rest, asks for 60.00 EUR. The right answer is a refusal carrying the number 54.00, because 84.00 minus the 30.00 still in flight leaves 54.00, and the second agent refunds that. Four days on, the first refund fails: the customer's card was replaced after fraud. Now 84.00 minus the 54.00 in flight leaves 30.00 again, and the first agent can re-send 30.00 by another method.

None of that works if refunded_amount is a counter on the capture, added to on acceptance and subtracted from on failure. Two agents racing read the same counter; a failed webhook delivered twice subtracts twice; a refund stuck in unknown is either counted or not, and both are wrong at different moments. Derive the figure from the refund rows instead, under a row lock on the capture, counting every state that might still move money.

TypeScript
// simplified: the refundable amount is computed, never stored
type RefundState =
  | 'requested' | 'submitted' | 'unknown' | 'accepted'
  | 'succeeded' | 'failed' | 'rejected';

// states that still hold, or have taken, part of the capture
const holding: RefundState[] = [
  'requested', 'submitted', 'unknown', 'accepted', 'succeeded',
];

function refundable(
  capturedMinor: number,
  refunds: { state: RefundState; amountMinor: number }[],
): number {
  const held = refunds
    .filter((r) => holding.includes(r.state))
    .reduce((sum, r) => sum + r.amountMinor, 0);
  return capturedMinor - held;
}

async function requestRefund(tx: Tx, captureId: string, amountMinor: number) {
  const capture = await tx.captures.lockForUpdate(captureId);
  const refunds = await tx.refunds.forCapture(captureId);
  const available = refundable(capture.amountMinor, refunds);
  if (amountMinor > available) throw new RefundExceedsCapture(available);
  return tx.refunds.insert({
    captureId, amountMinor, state: 'requested', idempotencyKey: randomUUID(),
  });
}

The lock stops the two agents both reading 54.00. The limit is the captured amount, not the authorised one: a 120.00 EUR authorisation captured for 84.00 EUR has 84.00 EUR to refund.

The failure, step by step

Take the opening scenario again, with the lifecycle in place. At 10:02 on Thursday the agent asks for 84.00 EUR. The orchestrator locks the capture, computes 84.00 available, inserts the row as requested, calls the gateway, and gets 200 with re_7f3 and pending. The row becomes accepted. The order page reads "Refund requested: 84.00 EUR, usually within 10 working days", the ledger holds a pending refund entry rather than a cash movement, and the customer gets one message saying their bank has been asked.

On day 9 the clearing run reaches the issuer, which rejects the credit because the account is closed. The gateway emits refund.failed for re_7f3 with reason account_closed. The handler stores the event, moves the row from accepted to failed, and the refundable amount on the capture is 84.00 EUR again. The pending ledger entry is reversed, a support case opens carrying the reason, and the customer is told the refund could not reach that card and asked for an account to pay into. A second refund row, method bank_transfer, goes against the same capture when they reply.

flowchart TD
  A[Refund accepted day 0] --> B[Order page says requested]
  A --> C[Ledger holds pending entry]
  A --> D[Clearing run reaches issuer]
  D --> E[Issuer rejects account closed]
  E --> F[refund failed webhook day 9]:::accent
  F --> G[Row accepted to failed]
  G --> H[Refundable back to 84 EUR]
  G --> I[Case opened with reason]
  G --> J[Customer asked for account]
What one failed webhook has to set in motion when accepted and succeeded are distinct

Without the state, day 9 has nothing to update: the page already said Refunded and the ledger already posted cash out. The first signal is the customer writing in three weeks later, or the settlement file arriving without the refund on it, which is why reconciling that file into ledger entries is the backstop for whatever the handler misses.

  1. Store the event

    Insert the raw refund.failed payload with its gateway event id before acknowledging; apply it from a job.
  2. Move the row

    Transition accepted to failed only if the row is still accepted; a second delivery finds failed and stops.
  3. Reverse the ledger entry

    Post the reversal of the pending refund entry, referencing the refund id and the event id.
  4. Open the case

    Route by reason code: a closed account needs new details from the customer, a merchant balance shortfall needs finance.
  5. Tell the customer first

    Send the message before they chase, naming the amount and asking only for what the reason requires.

What each state means to the customer and the ledger

I have integrated 15+ global and regional gateways behind one API, and the rule I apply is that the word Refunded never appears on a customer-facing page until a succeeded event has been stored for that row. Every other state has its own wording, its own ledger treatment and its own list of what an agent may do.

StateOrder pageLedgerAgent may
requested, submitted, unknownRefund being processedNothing posted; amount held on the captureNothing; wait for the resolver
acceptedRefund requested, with the expected windowPending refund entry against the captureAttempt a cancel where the gateway offers one
succeededRefunded, with the date from the eventRefund posted on the settlement lineClose the case
failedRefund could not be completedPending entry reversedRe-send, or refund by another method
rejectedNot shownNothingFix the amount or reason and resubmit

The window shown under accepted is stored per gateway and payment method; a card refund and a wallet refund arrive on different timescales, and one number across the platform is a promise broken for one of them.

Where the model breaks

Some gateways never send a failure event for a refund: the terminal status is only visible by polling, or the refund is simply missing from the settlement file. For those, a sweep job fetches every accepted refund older than the method's expected window and applies what it finds, and reconciliation opens an exception for any accepted refund absent from its payout.

Some payment methods have no instrument to push money back to. A bank transfer or a voucher means the refund is a payout: you collect account details, and accepted means the payout provider created the transfer. The lifecycle holds, but the row carries a different method and gateway from the capture, and the refundable check still runs against the original capture.

And a chargeback can arrive on a capture with a refund in accepted. Left alone, the customer is credited twice. The chargeback handler should look for refunds in flight on the capture and put the refund id and date in the evidence it submits.

Build the refund table first, with the state column and the derived refundable check under a capture lock, and stop writing Refunded anywhere until succeeded is stored. Then wire the refund.failed handler to open a case and message the customer, and add the sweep for the gateways that will never send it.

Written by Md Nasim Anjum, senior full-stack engineer in Manchester. He builds payment orchestration, KYC and KYB compliance platforms and conversational AI.

Get in touchAll writingRSS

More writing