WritingPayments
Apply payment webhooks in gateway order, never in arrival order
A playbook for storing, ordering and reducing gateway webhooks so a capture that arrives before its authorisation ends in the right state and amount.
A gateway posts payment.captured for pay_8c21, EUR 120.00, at 10:02:15. The payment.authorised event for the same payment reaches you at 10:02:49, because the first delivery hit a pod that was restarting and the gateway retried after 30 seconds. By the end of this playbook your handler will store both, work out that captured follows authorised whatever the clock on your side says, and leave pay_8c21 in Captured with 12000 minor units recorded. The rule it builds towards: apply webhooks in the gateway's order, reconstructed from an ordering key you choose per gateway, never in the order they reach your server, because a handler that moves state on arrival will leave a captured payment showing as authorised within its first week in production.
Store every event before you acknowledge it
The gateway's retry is the only mechanism that brings a lost event back, and it stops the moment you return 200. If you parse, validate and apply inside the request, any exception between the first byte and the commit loses the event for good, and a slow database turns into a timeout that the gateway counts as a failure, so it retries into the same slow database. Storing first makes the acknowledgement a statement about durability and nothing else.
Create one table, gateway_events, with these columns: gateway, gateway_event_id, payment_ref, event_type, sequence (nullable, because not every gateway sends one), occurred_at, received_at, payload as the raw bytes rather than a parsed object, and applied_at, null until the reducer has consumed the row. Put a unique constraint on the pair of gateway and gateway_event_id, and on conflict return 200 without touching anything else. The handler verifies the signature, inserts the row, enqueues the payment_ref for the apply job, and returns. Nothing in the request path reads or writes the payments table.
You know it worked when a deliberately redelivered event produces exactly one row, and no row with a null applied_at is older than a minute outside an incident.
flowchart TD W[Webhook endpoint] --> S[Store raw event] S --> K[Return 200] S --> Q[Apply job per payment] Q --> O[Sort by ordering key] O --> R[Reducer]:::accent R --> T[Payment state and totals] R --> P[Parked events] P --> F[Fetch payment after 60s] F --> S
Choose one ordering key per gateway
Arrival order is a property of your network, not of the payment. The gateway's own order lives somewhere in the payload, and where it lives differs by gateway, so decide at integration time what the key is, what breaks a tie, and how you will notice a gap. I have built orchestration across 15+ global and regional gateways, and the rule I apply is that an adapter is not finished until its ordering key and its gap detection are stated in code rather than in a wiki page.
| What the gateway gives you | Ordering key | Tie-breaker | How a gap shows up |
|---|---|---|---|
| A per-payment sequence number | sequence | none needed | an integer missing between the highest applied and the newest stored |
A server-side occurred_at with millisecond precision | occurred_at | rank of event type: authorised 1, captured 2, refunded 3, disputed 4 | a transition the state machine cannot make from the current state |
| Only an event id and the time of delivery | rank of event type, then the cumulative amount in the payload | received_at, as a last resort | same as above, and a refund is never applied without a fetch first |
The second row matters more than it looks. Two events with the same occurred_at to the millisecond are rare, but an authorisation and an immediate capture from an auto-capture flow can share a second, and if the gateway only gives you seconds, the rank is what keeps captured after authorised.
You know it worked when, for a gateway with sequence numbers, replaying the stored events of any payment sorted by the key produces the same final state as the gateway's own GET payment response.
Reduce the payment's state from the stored events, not from the incoming one
An on-arrival handler asks whether this event can move the current state, and has two answers, apply or drop, both wrong for an early event. Watch what it does to pay_8c21. The payment.authorised delivery gets a 502 from the restarting pod. Four seconds later payment.captured arrives, the handler checks the transition created to captured, finds it is not allowed, drops the event and returns 200, so the gateway never retries it.
At 10:02:49 the retried payment.authorised lands and the state becomes Authorised. The system now sees an authorised payment with 0 captured; the capture job picks it up that evening and sends a second capture for 12000, which the gateway rejects as already captured, or honours against the remaining authorisation if the gateway allows over-capture. The customer sees an order stuck at awaiting payment while their statement shows EUR 120.00 taken. The settlement file carries a capture the ledger never expected.
A reducer asks a different question: given every stored event for pay_8c21, sorted by the ordering key, what state does the payment end in? The incoming event is only the reason to run it again. Anything that cannot apply yet is parked, with applied_at left null, so when the authorised event finally lands the rerun applies both in order.
// Simplified: runs for one payment after any new event, under a row lock
type State = 'created' | 'authorised' | 'captured' | 'refunded' | 'voided';
type Ev = { id: string; type: string; key: number; amountMinor: number };
const next: Record<State, Partial<Record<string, State>>> = {
created: { authorised: 'authorised' },
authorised: { captured: 'captured', voided: 'voided' },
captured: { captured: 'captured', refunded: 'captured' },
refunded: {},
voided: {},
};
export function reduce(events: Ev[]) {
let pending = [...events].sort((a, b) => a.key - b.key);
let state: State = 'created';
let captured = 0;
let refunded = 0;
const applied: string[] = [];
let progressed = true;
while (progressed && pending.length) {
progressed = false;
for (const ev of pending) {
const to = next[state][ev.type];
const over = ev.type === 'refunded' && refunded + ev.amountMinor > captured;
if (!to || over) continue; // not yet: stays pending, never dropped
if (ev.type === 'captured') captured += ev.amountMinor;
if (ev.type === 'refunded') refunded += ev.amountMinor;
state = refunded > 0 && refunded === captured ? 'refunded' : to;
applied.push(ev.id);
pending = pending.filter((p) => p.id !== ev.id);
progressed = true;
break; // restart from the lowest key after every change
}
}
return { state, captured, refunded, applied, parked: pending };
}The pass restarts from the lowest key after every applied event, so a parked refund gets another chance once the capture it was waiting for goes through. A partial refund keeps the state at Captured and lets the totals carry the detail; only a refund that brings the two totals level moves it to Refunded. The apply job loads all events for the payment, runs reduce, writes the state and totals to the payment row, sets applied_at on the applied ids, and commits, in one transaction that holds a lock on the payment row so two deliveries for the same payment cannot run the reducer concurrently. Replay becomes free: fix a transition, clear applied_at for the affected payments, run the job.
You know it worked when deleting the payment row for a test payment and running the job brings it back identical.
Park early events and fetch when the gap outlives the retry window
Parking is the right first response, because the missing event is normally seconds behind. It is the wrong permanent response, because sometimes the missing event is never coming: the gateway's retry budget ran out during a long outage on your side, or the authorisation happened on the gateway's hosted page before your system knew the payment existed. A parked event older than the gateway's retry interval is a signal to go and ask rather than wait.
Run a job every 30 seconds that selects payments holding a parked event whose received_at is older than 60 seconds, and calls the gateway's GET payment endpoint for each. Write the response into gateway_events as an event of type synced, with gateway_event_id set to sync plus the payment ref plus the fetched state and totals, so a repeated fetch of an unchanged payment collides with the unique constraint and does nothing. Give the reducer one extra rule: a synced event replaces the working state and totals with what the gateway reported, and later events apply on top. Mark the row with source = fetch so reconciliation can tell webhooks from fetches.
You know it worked when you block the webhook endpoint for one test payment for two minutes and the payment still reaches Captured, with a synced row behind it.
Derive amounts from the events, never from deltas on the row
Partial captures and partial refunds are where arrival order does the most damage, because a running total on the payment row moves twice when an event is repeated or reordered. If captured_minor is a column you increment on each payment.captured, a redelivery that carries a fresh event id, which your unique constraint cannot catch, adds 8000 twice. Derived by the reducer from the full event set, the total is a function of the rows, and the duplicate is still visible to a reconciliation query as two rows with the same amount and the same occurred_at.
Take pay_8c21 again with an authorisation of 12000 and two captures, 8000 then 3500. A refund of 3500 arriving between the captures applies, because 3500 is within the 8000 captured. A refund of 10000 arriving in the same gap parks, because the reducer's over-refund check sees 10000 against 8000 and reads it as a missing capture rather than as the gateway refunding more than it took. When the 3500 capture lands, the rerun applies the capture and then the refund, in that order, and the totals end at 11500 captured and 10000 refunded. Write captured_minor and refunded_minor whole from the reducer's output on every run; never add to them.
You know it worked when a nightly query finds no payment whose row totals differ from the sums over its applied events.
Build the gateway_events table and the store-then-acknowledge handler first, because every other step reads from it and nothing is lost while you build the rest. The reducer comes second and replaces the on-arrival handler outright rather than sitting beside it. The fetch job can follow a week later; until then, parked events older than a minute go to a dashboard and a person.
Before the next gateway goes live
- Every webhook handler returns 200 only after the raw event row is committed
- The pair of gateway and gateway event id has a unique constraint, and a redelivery produces one row
- The adapter names its ordering key and tie-breaker in code, with a test for a capture that arrives first
- The reducer runs under a lock per payment and writes state and totals whole on every run
- An event that cannot apply is parked with a null applied at, never dropped
- A parked event older than 60 seconds triggers a fetch that writes an idempotent synced row
- A refund that exceeds the captured total parks and alerts instead of applying
- A nightly query compares row totals to sums over applied events and finds nothing