WritingPayments
Designing payment routing that fails over safely
What a routing engine in front of many gateways decides on every transaction, how it measures gateway health, and when a failed payment can safely try again.
At Paymid I work on payment orchestration: more than fifteen global and regional gateways behind one API, in a PCI DSS-compliant environment, with Apple Pay and Google Pay alongside cards and open banking. The part of that system I think about most is the Payment Routing Builder, which chooses a gateway for each transaction, balances load across them, and fails over when one stops answering.
Routing sounds like a lookup table. In practice it is a small decision engine that runs on every payment, where a wrong answer costs either a lost sale or, worse, a customer charged twice. This post sets out how I think about that decision: what the engine needs to know, how it should judge a gateway's health, and the rules that make failover safe rather than merely fast.
- 15+payment gateways behind one API
- 10+Apple Pay and Google Pay integrations
- 700+payment channels orchestrated
- 200+currencies handled
Routing is a decision per transaction
Every payment arrives with facts that narrow the choice: the payment method, the currency, the amount, the card's issuing country, the merchant, and whatever that merchant has agreed with each acquirer. Every gateway brings facts of its own: which methods and currencies it supports, which regions it serves, what it costs, and how it has behaved over the last few minutes.
The engine combines the two into an ordered list of routes, not a single answer. The first route carries the payment; the others are where it may go if the first fails for a reason that allows it. Returning the whole list up front keeps the failover decision inside the same logic as the original choice, and makes every decision explainable afterwards.
flowchart TD P[Payment] --> E[Eligibility] E --> R[Ranking] R --> W[Weighted choice] W --> A[Route 1] A --> X[Result] A -.->|technical failure| B[Route 2] B --> X
A payment arrives
Eligibility removes the impossible
Ranking orders the rest
Weight picks the first route
Failover, only when it is safe
Eligibility first, then ranking
I split the decision into two passes. The first removes every gateway that cannot take the payment at all. The second orders what is left.
type Method = "card" | "apple_pay" | "google_pay" | "open_banking";
type Payment = {
method: Method;
currency: string;
amount: number;
issuerCountry?: string;
merchantId: string;
};
type Gateway = {
id: string;
methods: Method[];
currencies: string[];
health: "up" | "degraded" | "down";
p95LatencyMs: number;
costBps: number; // processing cost, in basis points
weight: number; // this merchant's share of traffic for the gateway
};
// Pass one drops what cannot take the payment; pass two orders the rest
export function rank(payment: Payment, gateways: Gateway[]): Gateway[] {
const eligible = gateways.filter(
(g) =>
g.health !== "down" &&
g.methods.includes(payment.method) &&
g.currencies.includes(payment.currency),
);
return eligible.sort(
(a, b) =>
Number(a.health === "degraded") - Number(b.health === "degraded") ||
a.costBps - b.costBps ||
a.p95LatencyMs - b.p95LatencyMs,
);
}This is simplified, but the shape holds. Eligibility rules are hard constraints and are never traded off against anything else: a cheaper gateway that does not support the wallet is not a candidate. Ranking is where preferences live: a healthy route before a degraded one, then cost and speed. Keeping the two apart stops a scoring formula from quietly sending a payment somewhere it cannot succeed.
Health is measured, not configured
A gateway's status page is rarely the first to know it is in trouble. The engine should work it out from its own traffic: a rolling window of recent attempts per gateway and payment method, recording the outcome and latency of each.
The important distinction is between technical failures and business declines. A timeout, a connection error or a 5xx response says something about the gateway. An issuer declining for insufficient funds says nothing about it at all, and counting it would mark a healthy gateway as failing on a day when customers happen to be short of money.
From technical failures, a circuit breaker per gateway and method gives three states:
stateDiagram-v2 [*] --> Closed Closed --> Open: failures cross the threshold Open --> HalfOpen: cooling-off period ends HalfOpen --> Closed: test traffic succeeds HalfOpen --> Open: test traffic fails
| State | What it means | What the engine does |
|---|---|---|
| Closed | Failures are within normal levels | Routes to it as usual |
| Open | Failures crossed the threshold | Skips it for a cooling-off period |
| Half-open | The cooling-off period has passed | Sends it a small share of traffic, then closes or opens again |
Latency belongs in the same window. I watch the 95th percentile rather than the average: a gateway can look fine on average while one payment in twenty hangs long enough for the customer to give up.
Average and 95th-percentile latencyms
Failover without charging anyone twice
Fast failover is easy. Safe failover depends on knowing whether the first attempt could have succeeded.
If the gateway refused the connection, or rejected the request before sending it to the card network, nothing was authorised and the payment can move to the next route. If the request timed out after it was sent, the outcome is unknown: the issuer may have approved it. Sending the payment elsewhere at that point risks a second authorisation on the customer's card. The safe move is to ask the first gateway for the transaction's status, or reverse it, before trying anywhere else.
sequenceDiagram participant E as Routing engine participant A as Gateway A participant B as Gateway B E->>A: Authorise, key pay_123-a1 A--xE: Timeout after the request was sent E->>A: Status of pay_123-a1? A-->>E: Not authorised E->>B: Authorise, key pay_123-b1 B-->>E: Approved
Idempotency keys make this manageable. The merchant's payment carries one key that never changes, and each attempt at a gateway carries its own key derived from it. A retry against the same gateway can then never create a second charge, and every attempt traces back to one payment.
Declines need the same care:
| Outcome of the first attempt | Try the next route? |
|---|---|
| Connection refused, or rejected before reaching the network | Yes |
| Timed out after the request was sent | Only once the first attempt is confirmed as not authorised |
| Soft decline, such as "do not honour" | Sometimes, within the card schemes' reattempt rules |
| Hard decline: stolen card, closed account, invalid number | Never |
A failover route must also accept what the first attempt already collected. A wallet payment can only move to a gateway that supports that wallet, and a payment that has been through 3-D Secure needs a route that can use that authentication, or the customer is asked to authenticate again.
Load balancing keeps failover honest
If ranking always wins, every payment goes to the single best gateway and the others go cold. The cost shows up on the worst day: the failover route has not carried real traffic for weeks, and nobody knows whether it still works.
Splitting traffic by weight among healthy gateways keeps every route in use, gives each one enough volume for its health figures to mean something, and lets a merchant spread risk across acquirers. The weighted choice picks the first route; the ranked list supplies the fallbacks.
// Choose the first route by weight among healthy gateways; the rest follow in ranked order
export function withPrimary(ranked: Gateway[], random = Math.random()): Gateway[] {
const healthy = ranked.filter((g) => g.health === "up" && g.weight > 0);
const total = healthy.reduce((sum, g) => sum + g.weight, 0);
let point = random * total;
const primary = healthy.find((g) => (point -= g.weight) < 0) ?? ranked[0];
return primary ? [primary, ...ranked.filter((g) => g !== primary)] : [];
}Keep the checkout out of it
Cards, Apple Pay, Google Pay and open banking all arrive through one checkout, and the checkout never names a gateway. It asks for a payment and receives a result. That boundary is what lets routing change (a new acquirer, a new rule, a gateway taken out for maintenance) without touching a merchant's integration, and it keeps the PCI-sensitive parts of the flow in one place.
Wallets test that boundary. An Apple Pay or Google Pay token is encrypted for a specific party, and is decrypted either by the gateway or by the platform itself, inside its PCI scope. That choice limits which gateways a wallet payment can route to, so it belongs in the eligibility rules, not in the checkout.
What to measure
A routing engine is only as good as the numbers used to tune it. The ones I would put on the first dashboard:
- Authorisation rate by gateway, payment method and issuing country, so a drop in one corner is not averaged away
- 95th-percentile latency per gateway, alongside its timeout rate
- Failover rate and reasons: how often a second route was used, and whether it succeeded
- Traffic share against target weights, to catch a gateway that is quietly being skipped
- Soft-decline recoveries: payments approved on a second route after a soft decline on the first
Good routing is invisible. The merchant sees one API and the customer sees one checkout, while behind them the engine keeps choosing, measuring and falling back. Most of the engineering is in the unglamorous parts: telling a decline from an outage, knowing when a retry is safe, and keeping every route warm enough to trust.