WritingPayments

Designing payment routing that fails over safely

What a routing engine in front of many gateways decides on every transaction, how it measures gateway health, and when a failed payment can safely try again.

At Paymid I work on payment orchestration: more than fifteen global and regional gateways behind one API, in a PCI DSS-compliant environment, with Apple Pay and Google Pay alongside cards and open banking. The part of that system I think about most is the Payment Routing Builder, which chooses a gateway for each transaction, balances load across them, and fails over when one stops answering.

Routing sounds like a lookup table. In practice it is a small decision engine that runs on every payment, where a wrong answer costs either a lost sale or, worse, a customer charged twice. This post sets out how I think about that decision: what the engine needs to know, how it should judge a gateway's health, and the rules that make failover safe rather than merely fast.

  • 15+payment gateways behind one API
  • 10+Apple Pay and Google Pay integrations
  • 700+payment channels orchestrated
  • 200+currencies handled

Routing is a decision per transaction

Every payment arrives with facts that narrow the choice: the payment method, the currency, the amount, the card's issuing country, the merchant, and whatever that merchant has agreed with each acquirer. Every gateway brings facts of its own: which methods and currencies it supports, which regions it serves, what it costs, and how it has behaved over the last few minutes.

The engine combines the two into an ordered list of routes, not a single answer. The first route carries the payment; the others are where it may go if the first fails for a reason that allows it. Returning the whole list up front keeps the failover decision inside the same logic as the original choice, and makes every decision explainable afterwards.

A payment arrives

Method, currency, amount, issuing country and merchant: the facts that narrow the choice before any gateway is considered.

Eligibility removes the impossible

Gateways that are down, or cannot take this method or currency, are dropped. These are hard constraints, never traded against price or speed.

Ranking orders the rest

Healthy before degraded, then cost and latency as tie-breakers. The result is a list, not an answer.

Weight picks the first route

Traffic is split by weight among healthy gateways, so every route carries real payments and stays trustworthy.

Failover, only when it is safe

If the first route fails for a technical reason the payment moves on, but after a timeout only once the first gateway confirms nothing was authorised.

Eligibility first, then ranking

I split the decision into two passes. The first removes every gateway that cannot take the payment at all. The second orders what is left.

TypeScript
type Method = "card" | "apple_pay" | "google_pay" | "open_banking";

type Payment = {
  method: Method;
  currency: string;
  amount: number;
  issuerCountry?: string;
  merchantId: string;
};

type Gateway = {
  id: string;
  methods: Method[];
  currencies: string[];
  health: "up" | "degraded" | "down";
  p95LatencyMs: number;
  costBps: number; // processing cost, in basis points
  weight: number; // this merchant's share of traffic for the gateway
};

// Pass one drops what cannot take the payment; pass two orders the rest
export function rank(payment: Payment, gateways: Gateway[]): Gateway[] {
  const eligible = gateways.filter(
    (g) =>
      g.health !== "down" &&
      g.methods.includes(payment.method) &&
      g.currencies.includes(payment.currency),
  );
  return eligible.sort(
    (a, b) =>
      Number(a.health === "degraded") - Number(b.health === "degraded") ||
      a.costBps - b.costBps ||
      a.p95LatencyMs - b.p95LatencyMs,
  );
}

This is simplified, but the shape holds. Eligibility rules are hard constraints and are never traded off against anything else: a cheaper gateway that does not support the wallet is not a candidate. Ranking is where preferences live: a healthy route before a degraded one, then cost and speed. Keeping the two apart stops a scoring formula from quietly sending a payment somewhere it cannot succeed.

Health is measured, not configured

A gateway's status page is rarely the first to know it is in trouble. The engine should work it out from its own traffic: a rolling window of recent attempts per gateway and payment method, recording the outcome and latency of each.

The important distinction is between technical failures and business declines. A timeout, a connection error or a 5xx response says something about the gateway. An issuer declining for insufficient funds says nothing about it at all, and counting it would mark a healthy gateway as failing on a day when customers happen to be short of money.

From technical failures, a circuit breaker per gateway and method gives three states:

stateDiagram-v2
  [*] --> Closed
  Closed --> Open: failures cross the threshold
  Open --> HalfOpen: cooling-off period ends
  HalfOpen --> Closed: test traffic succeeds
  HalfOpen --> Open: test traffic fails
StateWhat it meansWhat the engine does
ClosedFailures are within normal levelsRoutes to it as usual
OpenFailures crossed the thresholdSkips it for a cooling-off period
Half-openThe cooling-off period has passedSends it a small share of traffic, then closes or opens again

Latency belongs in the same window. I watch the 95th percentile rather than the average: a gateway can look fine on average while one payment in twenty hangs long enough for the customer to give up.

Average and 95th-percentile latencyms

  • Gateway A, average180 ms
  • Gateway A, p95260 ms
  • Gateway B, average170 ms
  • Gateway B, p951,400 ms
Two gateways with similar averages can have very different tails. Illustrative figures.

Failover without charging anyone twice

Fast failover is easy. Safe failover depends on knowing whether the first attempt could have succeeded.

If the gateway refused the connection, or rejected the request before sending it to the card network, nothing was authorised and the payment can move to the next route. If the request timed out after it was sent, the outcome is unknown: the issuer may have approved it. Sending the payment elsewhere at that point risks a second authorisation on the customer's card. The safe move is to ask the first gateway for the transaction's status, or reverse it, before trying anywhere else.

sequenceDiagram
  participant E as Routing engine
  participant A as Gateway A
  participant B as Gateway B
  E->>A: Authorise, key pay_123-a1
  A--xE: Timeout after the request was sent
  E->>A: Status of pay_123-a1?
  A-->>E: Not authorised
  E->>B: Authorise, key pay_123-b1
  B-->>E: Approved
A timeout is not a failure until the first gateway confirms it

Idempotency keys make this manageable. The merchant's payment carries one key that never changes, and each attempt at a gateway carries its own key derived from it. A retry against the same gateway can then never create a second charge, and every attempt traces back to one payment.

Declines need the same care:

Outcome of the first attemptTry the next route?
Connection refused, or rejected before reaching the networkYes
Timed out after the request was sentOnly once the first attempt is confirmed as not authorised
Soft decline, such as "do not honour"Sometimes, within the card schemes' reattempt rules
Hard decline: stolen card, closed account, invalid numberNever

A failover route must also accept what the first attempt already collected. A wallet payment can only move to a gateway that supports that wallet, and a payment that has been through 3-D Secure needs a route that can use that authentication, or the customer is asked to authenticate again.

Load balancing keeps failover honest

If ranking always wins, every payment goes to the single best gateway and the others go cold. The cost shows up on the worst day: the failover route has not carried real traffic for weeks, and nobody knows whether it still works.

Splitting traffic by weight among healthy gateways keeps every route in use, gives each one enough volume for its health figures to mean something, and lets a merchant spread risk across acquirers. The weighted choice picks the first route; the ranked list supplies the fallbacks.

TypeScript
// Choose the first route by weight among healthy gateways; the rest follow in ranked order
export function withPrimary(ranked: Gateway[], random = Math.random()): Gateway[] {
  const healthy = ranked.filter((g) => g.health === "up" && g.weight > 0);
  const total = healthy.reduce((sum, g) => sum + g.weight, 0);
  let point = random * total;
  const primary = healthy.find((g) => (point -= g.weight) < 0) ?? ranked[0];
  return primary ? [primary, ...ranked.filter((g) => g !== primary)] : [];
}

Keep the checkout out of it

Cards, Apple Pay, Google Pay and open banking all arrive through one checkout, and the checkout never names a gateway. It asks for a payment and receives a result. That boundary is what lets routing change (a new acquirer, a new rule, a gateway taken out for maintenance) without touching a merchant's integration, and it keeps the PCI-sensitive parts of the flow in one place.

Wallets test that boundary. An Apple Pay or Google Pay token is encrypted for a specific party, and is decrypted either by the gateway or by the platform itself, inside its PCI scope. That choice limits which gateways a wallet payment can route to, so it belongs in the eligibility rules, not in the checkout.

What to measure

A routing engine is only as good as the numbers used to tune it. The ones I would put on the first dashboard:

  • Authorisation rate by gateway, payment method and issuing country, so a drop in one corner is not averaged away
  • 95th-percentile latency per gateway, alongside its timeout rate
  • Failover rate and reasons: how often a second route was used, and whether it succeeded
  • Traffic share against target weights, to catch a gateway that is quietly being skipped
  • Soft-decline recoveries: payments approved on a second route after a soft decline on the first

Good routing is invisible. The merchant sees one API and the customer sees one checkout, while behind them the engine keeps choosing, measuring and falling back. Most of the engineering is in the unglamorous parts: telling a decline from an outage, knowing when a retry is safe, and keeping every route warm enough to trust.

Written by Md Nasim Anjum, senior full-stack engineer in Manchester. He builds payment orchestration, KYC and KYB compliance platforms and conversational AI.

Get in touchAll writingRSS

More writing