The order exists and we have reserved the inventory. The payment provider now takes four seconds to respond. The simplest implementation keeps the HTTP request open:
browser → order API → payment provider → order API → browser
Keeping the request open can be the right design. Adding a queue does not automatically make payment more reliable. The decision depends on what the caller must know before we respond and whether the product can show a pending state honestly.
Synchronous work gives an immediate answer
If the API waits, it can return a clear result: payment succeeded, payment was declined, or the attempt could not be completed within the deadline.
This is easy for the caller to understand, but the request now inherits the payment provider’s latency and failures. During those four seconds, the system keeps several resources open:
- an incoming request and response;
- application memory and tracing state;
- an outgoing socket;
- possibly a concurrency slot or database connection if the code holds one accidentally;
- load-balancer and client timeout budget.
The full request deadline has to be long enough for the work, and every internal timeout has to fit inside that deadline. A ten-second provider timeout cannot help when the gateway stops waiting after five seconds.
A timeout creates uncertainty
Suppose the provider charges the card, but the response packet is lost. The order API times out. Did payment fail?
The order API only knows that it did not receive a final answer. If it retries without an idempotency key, the customer may be charged twice. If it marks the payment as failed, that may disagree with the provider’s state.
A timeout only tells us that this caller did not receive an answer in time:
declined → known business result
succeeded → known business result
timeout → outcome unknown from this observer
The workflow therefore needs a way to reconcile the unknown result. It can query the provider, wait for a webhook, or retry later with the same idempotency identity.
Async work returns a pending state
An asynchronous design accepts the request, records durable intent, and returns before payment completes:
HTTP/1.1 202 Accepted
Location: /orders/ord_123
{"id":"ord_123","status":"payment_pending"}
A worker later processes the payment. The client polls, subscribes to updates, or receives a notification when the state changes.
This shortens the original request and lets workers control how many payments they attempt at once. A queue can also absorb a temporary burst. In return, the system now has more states and more components to operate, and the interface has to explain what payment_pending means.
Once we do this, the API contract has changed. The client now has to understand and display a pending order.
Do not lose work between the database and queue
A dangerous implementation commits the order, then publishes a message:
1. commit order with payment_pending
2. publish ProcessPayment
If the process crashes between those steps, the order remains pending with no message to advance it.
The transactional outbox records the state change and an event in the same database transaction:
begin;
insert into orders (..., status) values (..., 'payment_pending');
insert into outbox (topic, aggregate_id, payload)
values ('process-payment', $1, $2);
commit;
A publisher reads unsent outbox rows and delivers them to the broker. The outbox closes the gap where an order can commit without its message. Publication may still happen more than once, so the consumer has to handle duplicates.
At-least-once delivery requires idempotent consumers
Queues commonly redeliver when a worker crashes after performing work but before acknowledging the message. The payment worker may therefore see the same command twice.
Use a stable operation key at every retry boundary:
- order API assigns the payment attempt ID;
- the queue message carries it;
- the worker records processed attempts;
- the provider receives it as its idempotency key when supported.
Durable uniqueness should enforce idempotency. An in-memory “seen” set disappears when the worker restarts and does not coordinate several workers.
Retry only with a budget
Retries can recover from a brief failure, but they add more traffic when the dependency remains unavailable. A retry policy needs:
- a maximum attempt count or absolute deadline;
- exponential backoff;
- jitter so many workers do not retry together;
- classification of retryable and permanent errors;
- idempotency for the attempted operation;
- a destination for work that exhausts its budget.
A circuit breaker can temporarily stop new calls when a dependency is clearly failing. It protects capacity and provides faster failure, but it cannot decide the order’s final business state. The state machine still owns that.
Model payment as a state machine
A single paid boolean cannot represent all the outcomes we have already seen. A payment state machine may look like this:
payment_pending
├── authorized
├── declined
├── unknown → reconciliation
└── cancelled
authorized
├── captured
├── voided
└── capture_failed
Transitions should be validated. A late “authorized” webhook should not overwrite an already refunded payment without deliberate rules. Store provider event IDs so duplicate webhooks are harmless.
The order and payment should be allowed to have separate states. If the order is cancelled after the charge succeeds, the workflow may need to issue a refund. It cannot erase the fact that the charge happened.
Decide with product requirements
Keep payment synchronous when:
- users need the final answer immediately;
- latency is predictably within the request budget;
- the traffic level and dependency limits support it;
- uncertainty and retry behavior are handled.
Prefer an asynchronous workflow when:
- the work is slow or bursty;
- the caller can accept a pending state;
- work must survive process failure;
- concurrency needs to be controlled independently;
- retries and reconciliation are first-class requirements.
Hybrid designs are common. The API can try payment synchronously for a short budget, then continue asynchronously if the result remains unknown. That improves the fast path without lying about the slow path.
I would decide between these designs by asking what durable state we can promise when the API responds. That answer tells us whether payment should finish on the request path or continue in the background.
In the next episode, the full request takes 1.8 seconds and we need to find where the system spent that time.
