How Payment Authorization Decisions Are Made in Milliseconds

Content authorBy EGSPublished onReading time14 min read
Close-up of two hands interacting with a payment terminal at a supermarket checkout, with a blurred retail background.

A card authorization is a synchronous request-response round trip. The terminal builds a message that the acquirer and card network route to the issuer, and the issuer combines account and cryptographic checks into a single approval or decline code that travels back the same path. Money moves later, during clearing and settlement.

What happens after a card tap?

The tap starts a message that crosses four parties and returns inside a second. The point-of-sale (POS) terminal hands an authorization request to the acquirer, which passes it through a payment switch and the card network to the issuer processor. Every hop adds network time and queueing time to the budget.

Contactless checkouts complete in roughly 500 to 1,500 milliseconds end to end, with the issuer's own decision occupying a small slice of that window. The rest is radio handshake and wire time you don't control.

What matters here is the split between authorization and money movement. Approval only places a hold. Capture and settlement run in batches hours or days later, which is why an authorization can be reversed without any funds ever leaving an account. Design the hot path for the hold.

The terminal builds the request

The terminal assembles a structured message before anything leaves the store. It packages the primary account number or token and the Europay, Mastercard, and Visa (EMV) chip data the card produced during the tap.

Most of that structure comes from ISO 8583, where an authorization request carries the message type indicator 0100 followed by a bitmap and a variable set of data elements. The bitmap tells the receiver which fields are present, which is why parsers can validate a message without reading it end to end.

The chip contributes the piece that can't be faked downstream. During the first Generate Application Cryptogram command, the card returns an Authorization Request Cryptogram (ARQC) when it wants an online decision. Because the terminal can't produce that value itself, everything after this point inherits the card's own signature over the transaction data, and any field the terminal mangles will surface later as a cryptogram mismatch rather than a routing error.

The switch finds the issuer

The switch's job is routing, and it does structural validation on the way. It checks that the message parses and that the bank identification number maps to a known route. Then it forwards to the correct network or directly to the issuer processor when the card is on-us.

Card networks apply their own controls at this layer. Visa's authorization platform is engineered for a peak of roughly 65,000 transaction messages per second, which tells you the routing tier is built to add microseconds, not milliseconds, per message.

That volume ceiling has a design consequence most teams miss. A switch that spends real time on business logic becomes the bottleneck for every issuer behind it, so routing decisions belong in memory-resident lookup tables while anything requiring a database read belongs at the issuer. If your switch is querying a relational store to decide where a message goes, you've already lost your latency budget before the issuer sees the request.

The response retraces the route

The issuer answers with a response message that follows the request path in reverse. In ISO 8583 terms, the 0100 request becomes a 0110 response that carries a response code and the issuer's response cryptogram.

The response code is the whole outcome. Code 00 approves. Codes in the 100-series decline without card pickup, such as 101 for expired card and 102 for suspected fraud. Code 91 signals the issuer is unavailable, which is a different problem from a decline and gets handled differently.

Terminals treat these codes as a contract, so an issuer that returns a vague code strips the acquirer and merchant of any useful recovery path. A soft decline coded precisely enough to retry recovers revenue. "Do not honor" does not, because nobody downstream can tell whether the account was frozen or the limit was hit.

The issuer combines several checks

Infographic illustrating funds availability check process with labeled steps, rounded cards, and curved arrows in a corporate fintech style.

The issuer runs its checks, then applies a fixed precedence order to whatever those checks return. Precedence is what keeps the outcome deterministic when two subsystems disagree.

The precedence order carries real liability weight. Under the Mastercard Switch Rules, an issuer is liable for every transaction authorized through Stand-In Processing as long as the network correctly applied the issuer's own parameters. The decision logic you encode is the decision you own financially.

That reframes how you think about the ordering. Hard blocks like closed accounts and failed cryptograms have to short-circuit everything else, because approving past them creates an unrecoverable loss rather than a recoverable one. Soft signals like an elevated fraud score belong last, where they can be overridden by an explicit allow rule. Teams that treat these as a flat list of conditions end up with contradictory approvals that only appear during reconciliation, weeks after the money is gone.

Start building your financial platform?

Speak with EGS engineers about open banking, payment infrastructure, cloud systems, and enterprise software.

Get in Touch →

Funds and limits permit spending

The funds check answers whether this specific amount is spendable right now, which is a narrower question than what the balance says. Available funds means the posted ledger balance minus outstanding authorization holds, minus any pending debits, adjusted for credit line and overdraft arrangements.

Holds are the awkward part because they expire on the network's schedule rather than yours. Visa's rules give most card-present and merchant-initiated transactions five calendar days from authorization to clearing, while customer-initiated card-not-present transactions get ten.

Card controls sit alongside the balance check and run just as fast: account status and velocity counters. Velocity is the one that costs you. Counting transactions across a rolling window requires either a write on every authorization or a probabilistic structure, and if you put that counter in your primary ledger database, every tap now carries a write to your most contended table. Keep the counters in memory and reconcile them against the ledger asynchronously.

Fraud models estimate transaction risk

Scoring runs inside the authorization window or it doesn't run at all. A model evaluates the transaction against the account's behavioral history and against network-wide patterns, returns a score, and rules translate that score into an action within the same call.

Visa's Advanced Authorization service evaluates up to 400 unique attributes per transaction while scoring in flight. Feature computation at that scale only fits the budget when the features are precomputed and cached against the account.

Rules and models do different jobs and remain separate. Models produce a probability. Rules encode policy you can defend to a regulator or an internal audit. The cost of getting the balance wrong runs one direction: Datos Insights projected global e-commerce losses to false declines reaching $231.35 billion in 2025. A threshold tuned purely to minimize fraud basis points is destroying more revenue than it saves, and the loss never appears on a statement.

Cryptography verifies transaction integrity

The cryptogram check proves the transaction data is authentic. The issuer host recomputes the ARQC using the card's derived key and the transaction fields it received, then compares. A match confirms the chip generated this specific transaction and that nothing in the message was altered in transit.

Hardware security modules (HSMs) hold the keys and do the math. A Thales payShield 10K handles up to 10,000 cryptographic operations per second and is certified to FIPS 140-2 Level 3 and PCI HSM v3, so key material never exists in application memory.

If verification succeeds, the HSM generates an Authorization Response Cryptogram (ARPC) using the ARQC and the response code as inputs, which the card validates on its return. Treat your HSM pool as a hard capacity constraint rather than a service you call freely, because the command rate is fixed by hardware and can't be autoscaled the way a stateless service can. Size the pool against peak transactions per second with headroom, and measure HSM queue depth as a first-class latency signal.

Sub-100ms requires a disciplined hot path

Sub-100ms describes the issuer's internal decision budget. Nothing about a fast decision engine shortens the NFC handshake or the transatlantic hop. What it buys you is margin against the network's timeout.

That margin matters because the outer bound is generous and the inner one is not. Networks invoke Stand-In when an issuer misses a response threshold set between 8 and 15 seconds, which sounds like an enormous amount of room until a dependency stalls under peak load.

Here's the reasoning behind the tight internal target. Building to a 100ms budget forces every synchronous dependency to justify itself, which means when a downstream cache degrades or an HSM queue backs up, you absorb the damage inside your own timeout instead of handing the decision to a network rulebook you configured months ago. The budget is a forcing function for dependency discipline. It's not a performance vanity metric, and it never comes at the cost of ledger correctness.

Memory serves frequently needed data

Hot-path reads belong in memory. The account state an authorization needs fits comfortably in RAM for even large portfolios, and keeping it there removes disk I/O from the critical path.

The gap is not subtle. Benchmarked read latency runs about 0.095 ms for Redis against 0.65 ms for PostgreSQL, roughly seven times faster before you account for connection overhead and query planning.

Freshness is where this gets dangerous. A stale card status flag approves a transaction on a card the customer reported stolen four minutes ago. So mutable state needs explicit invalidation on every write path, and balances in particular need a reconciliation loop against the authoritative ledger rather than a time-to-live and a shrug. Read-through caching with generous expiry works for merchant category tables.

Start building your financial platform?

Speak with EGS engineers about open banking, payment infrastructure, cloud systems, and enterprise software.

Get in Touch →

Independent checks run concurrently

Checks that don't depend on each other run at the same time. Fraud scoring and cryptogram verification have no data dependencies between them, so running them in parallel makes the critical path equal to the slowest single check rather than the sum.

Parallelism reshapes the tail as well as the median. Jeffrey Dean and Luiz André Barroso showed in The Tail at Scale that when a request fans out across many backends, the slowest component dominates the response time, which is why they proposed hedged and tied requests as mitigations.

Fan-out only helps if the combining logic is predictable. Define a fixed precedence and a hard deadline that cancels outstanding work. The subtle failure is a cancelled fraud call that already incremented a velocity counter, which leaves your state ahead of your decision. Cancellation has to be safe at every point where a check writes anything.

Horizontal scaling preserves consistency

Authorization scales horizontally when the decision service holds no session state. Any node handles any request, and connection pools to the switch and HSM stay persistent rather than reconnecting.

Partitioning is what keeps correctness intact under that model. EGS describes modern platforms scaling through stateless services behind load balancers, with microservices isolating authorization from ledger and settlement work while idempotency keys prevent double charges.

The concurrency problem is specific and worth naming. Two authorizations against the same account arriving milliseconds apart on different nodes can both read the same available balance and both place a hold. Partitioning by account identifier routes them to the same node and serializes them naturally, which is cheaper than distributed locking and faster than optimistic retry. Pair that with a unique constraint on the retrieval reference number so a network retry of the same message returns the stored response instead of creating a second hold.

Failures require explicit payment outcomes

When the issuer never responds or responds after the acquirer stopped waiting, the system is in an ambiguous state where the authorization exists on one side of the wire and not the other. Recovery logic has to resolve that ambiguity explicitly.

Card networks separate these outcomes at the code level. Response code 91 means the issuer is unavailable and the external ledger couldn't be reached, while code 96 signals a system malfunction or a breach of scheme response time limits. Visa converts both to N0 when it forces single-request Stand-In.

Treating ambiguity as a decline is the mistake that costs money twice. You lose the sale, and you leave a hold sitting on the account with no matching clearing record, which surfaces as a customer complaint and a reconciliation break. Build the uncertain state into your transaction model as a real status with its own resolution path.

Timeouts can trigger stand-in decisions

Stand-In lets someone else decide when you can't. The card network or issuer processor responds on the issuer's behalf using parameters the issuer configured in advance as it evaluates card status and spending ceilings without access to the live balance.

The two flavors differ sharply in risk. Basic Stand-In declines everything when no response arrives within the threshold, which protects integrity and avoids scheme fines, while Enhanced Stand-In approves transactions that meet the criteria you defined and declines the rest.

Because Stand-In runs blind to your ledger, the parameters have to assume the worst plausible account state. Set accumulative limits low enough that a full day of Stand-In on your largest portfolio produces a loss you'd accept rather than a loss you'd escalate. And log every Stand-In advice the network sends when you come back online, because those transactions are already authorized and your balances are wrong until you post them.

Reversals repair uncertain authorizations

A reversal cancels an authorization. ISO 8583 defines an acquirer reversal request as message type 0400, sent when the acquirer needs to undo a transaction the issuer already approved, most commonly because the response never reached the terminal.

Three mechanisms have to work together for this to hold up:

  • A stable transaction identifier carried on the original request and the reversal, so both can be matched to one logical transaction.

  • Idempotent handling keyed on that identifier, so a repeated reversal releases the hold once and returns the same result on every subsequent attempt.

  • A late-response rule that discards or reverses an approval arriving after the acquirer has already timed out and reversed.

Without a reversal, the hold sits until the network's own expiry runs out, which can take days. Your customer sees pending funds for a purchase that never happened, and your support queue absorbs the difference. That's why the reversal path deserves the same latency and reliability engineering as the authorization path itself.

Tail latency reveals system health

Averages hide the failures that matter. Measure each hop separately from terminal to issuer, then track p95 and p99 on every one. A 40ms mean with a 900ms p99 means one transaction in a hundred is a customer complaint.

Instrument correctness alongside speed. Track timeout rate and approval rate broken out by decline code so you can see when a latency regression starts converting into declines. The observability set for idempotent payment APIs includes duplicate request count and in-progress conflict count alongside gateway timeout rate, which is the right instinct: correctness metrics and latency metrics belong on the same dashboard.

Load testing has to reproduce peak concurrency and failure modes together. A system that holds p99 at three times normal volume can still collapse when a single HSM node drops out at that volume, because the queue that was invisible at p99 becomes the whole story.

EGS can modernize authorization infrastructure

If your authorization path is losing time to synchronous dependencies you can't fully account for, the fix starts with measuring each hop and rebuilding the ones that don't justify their cost. That work is easier with a team that has already built the components.

Energize Global Services (EGS) is a technology company founded in 2007 with offices in Boston and Yerevan that works on banking systems and POS terminal solutions. The engineering work spans payment switches and custom banking infrastructure built for low-latency transaction processing.

Book a call with the EGS engineering team and bring your current latency numbers and your timeout rate. That's the conversation that produces a concrete plan rather than a general assessment.

Start building your financial platform?

Speak with EGS engineers about open banking, payment infrastructure, cloud systems, and enterprise software.

Get in Touch →

A partial authorization approves less than the requested amount when the account has limited available funds. The terminal receives the approved amount and, where supported, asks for another payment method for the remainder. The merchant can't treat the original amount as fully approved.

An offline approval is made by the card and terminal without a live issuer response. EMV rules and limits stored on the chip govern the decision, so the terminal has no current view of the account balance. Issuers restrict offline use because the transaction reaches the account later.

Hotels and fuel stations often authorize an estimated amount before they know the final charge. A hotel may cover an expected stay and incidental charges, while a fuel dispenser may reserve a preset amount before pumping begins. The merchant captures the final amount and releases any unused portion.

A merchant can't freely raise an approved amount after authorization. For tips, extended stays, or added purchases, card-network rules require an incremental authorization or allow a limited adjustment under defined conditions. The issuer then evaluates the added amount against the account's current available funds and controls.

Contact your bank after the merchant confirms a cancellation or completion and the hold remains beyond the release period it provides. Ask the merchant first to send a reversal if the purchase didn't occur. Your bank can explain its hold policy, but it usually needs the merchant's message to remove the hold early.

Schedule a Meeting

Book a time that works best for you

You Might Also Like

Discover more insights and articles

A single hand holds a credit card partially inserted into a sleek payment device, set against a minimal background.

What Systems Decide Whether a Payment Is Approved or Declined

Four systems can stop a card payment, from the merchant's gateway and fraud tools through the acquirer and the card network to the issuing bank. The issuer makes the final authorization call after checking the account balance and the fraud score, then applying scheme rules. Any earlier system can block the transaction before the issuer ever sees it.

A bank analyst monitors transaction flow on multiple screens in a modern office, showcasing a high-tech, professional environment.

Liquidity Management in Real-Time Payments Systems

Liquidity management in real-time payments is the practice of keeping enough funds or credit immediately available in settlement accounts to clear every outbound payment instantly, around the clock. Because instant payments settle transaction by transaction with no netting window, banks prefund or continuously replenish balances instead of squaring positions once a day.

A modern corporate meeting room with banking professionals discussing KPIs on a large dashboard display, surrounded by laptops and reports.

Key Challenges Banks Face When Adopting Instant Payments

Banks adopting instant payments face five connected gaps, from batch-based cores that can't post in seconds to liquidity management that stops at 5 p.m. Each gap has to close before launch.

A professional reviews a SEPA Instant Credit Transfer flow diagram on a large monitor in a modern corporate engineering office.

How European Instant Payments Infrastructure Operates

European instant payments run on SEPA Instant Credit Transfer (SCT Inst), the European Payments Council scheme for euro credit transfers cleared and settled continuously in seconds. Payer banks send ISO 20022 instructions through either RT1 or TIPS, and the beneficiary bank credits the account and returns a status inside ten seconds every day of the year.