In a request-driven system the service that did something calls

everyone who needs to know, so it waits on all of them and fails

when any of them fails. In an event-driven system it publishes a

fact (payment.captured) and moves on; the ledger, the warehouse

and the email service each subscribe and react on their own time.

You trade that coupling for eventual consistency, duplicate

deliveries and harder debugging, so keep anything the caller needs

before it can answer as a direct call.

A payment gets captured. Five things need to happen because of it: the ledger books it, the warehouse reserves the stock, the customer gets a receipt, analytics counts it, and later someone adds a loyalty points service. The question this post is about is simple: whose job is it to make all of that happen?

The obvious answer is the checkout service's, since it's the one that knows the payment happened:

It reads well and it works. Now look at what it costs.

Latency is the sum. The ledger takes 40ms, the warehouse 60ms, the email provider 300ms on a good day, analytics 50ms. The customer waits 450ms for a receipt email they'll read later.

Availability is the product. If each of those four services is up 99.9% of the time, checkout only succeeds when all four are: 0.999⁴ ≈ 99.6%. That's the difference between about 9 hours of downtime a year and about 35. Analytics being down now means customers can't pay.

Every new consumer is a change to checkout. The loyalty team can't ship until the payments team adds a fifth call, reviews it, and deploys it. Checkout has become the place where every other team's reaction to a payment lives.

The root of all three is the same: the service that did the thing is also responsible for telling everyone about it, one call at a time.

Flip the responsibility. Checkout's job is to capture the payment and say so. Whoever cares listens:

Each consumer subscribes on its own:

The message goes through a broker, something like or , which stores it so each subscriber can read it on its own schedule. (RabbitMQ keeps a message until it's taken; Kafka keeps everything for a retention window, days by default, whether anyone has read it or not.)

Re-run the three costs. Checkout now waits on one write to the broker,

not four services, and depends on one thing being up (the broker)

instead of four. Analytics going down means analytics falls behind,

and catches up from the broker when it comes back, as long as it's

back within the retention window; payments keep working. And the

loyalty team subscribes to payment.captured and ships without asking

the payments team for anything.

The word "message" covers two different things, and mixing them up is the most common way an event-driven design goes wrong.

The tell for a command wearing an event's name is an event named after

what should happen next: SendReceiptEmail. That's checkout telling

the email service what to do, with the coupling still there, just

routed through a broker. Name the fact (PaymentCaptured) and let the

email service decide that a receipt is its reaction.

Brokers deliver in two shapes, and you usually need both.

A queue hands each message to exactly one of the workers reading it. Run three copies of the email service on one queue and each receipt goes to one copy, whichever is free, not to all three. That's how a consumer scales.

A topic hands each message to every subscriber. The ledger, the

warehouse and the email service all get their own copy of

payment.captured. That's how producers stay unaware of consumers.

Most brokers combine them. In Kafka, a topic delivers to every consumer group, and within a group each message goes to one member. So each service is a group: every service sees every payment, and within a service the work is split across its instances.

Nothing above is free. Here's what you sign up for.

The API returns 201 before the ledger has booked anything. If the

customer's next click is "view my balance" and that page reads the

ledger, it may not show the payment yet. This is eventual consistency:

every consumer gets there, just not at the moment the producer

answers. Design screens around it, for example by showing the payment

from checkout's own record rather than the ledger's.

Assume every broker delivers at least once: it's the usual default, and the only setting that doesn't risk losing messages. A consumer that crashes after booking the ledger entry but before acknowledging the message gets it again. Every consumer has to be idempotent: record the event's id alongside its work, in the same transaction, and skip ids it has seen. It's the same idea as an on an API.

Kafka keeps order within a partition, not across a topic, and never

across two topics. If payment.refunded can overtake

payment.captured, the ledger will try to reverse something it hasn't

booked. So the samples above, with one topic per event type, can't

promise that order. When order matters, publish every payment event to

one payments topic with the payment id as the partition key: all

events for one payment then land on one partition, in order.

When a receipt doesn't arrive, there's no single call chain to read. The request ended at checkout; the failure happened seconds later in a different service. Put a trace id in every event's metadata and have every consumer log it, so one search pulls up the whole story. It's what are for.

An event with a field the consumer can't parse fails every time. Retry it a few times with , then move it to a dead-letter queue: a side queue someone inspects and replays, so one bad message doesn't block every message behind it.

Fan-out works when the reactions are independent. The ledger doesn't care whether the email went out. But some flows are a sequence with consequences. A refund has to reverse the ledger entry, then return the money through the provider, then restock, and if the provider refuses, the ledger reversal has to be undone.

You can build that as choreography: each service reacts to the

previous one's event (LedgerReversed → provider refunds →

RefundIssued → warehouse restocks). No one owns the whole flow,

which is flexible and also the problem: to answer "where is refund 123

stuck?" you have to reconstruct it from three services' logs.

Or as orchestration: one refund service sends commands in order and tracks the state of each refund, including what to undo when a step fails. More central, easier to see. Both are forms of a : a sequence of local steps, each with a compensating step to undo it. They differ in whether one service owns the sequence or it emerges from the reactions.

The rule of thumb: independent reactions to a fact are choreography. A multi-step business process that can fail halfway belongs to an orchestrator, which still talks to the rest of the system through events and commands.

One call was left out of the handler on purpose: the fraud check. It runs before the capture, and it stays a direct call:

Checkout can't answer the customer until it knows the answer, so an event doesn't help. There's nothing to do while waiting. The split that decides most cases:

That last row matters. If checkout, the ledger and the email sender are one app with one database, owned by one team, a broker adds latency, infrastructure and every item on the bill above, to decouple code that a function call already separates well enough. Events earn their place when separate services, owned by separate teams, need to react to the same thing without waiting on each other.

Back to the payment. Checkout captures it, says so, and answers the customer in the time it takes to write one message. The ledger, the warehouse, the receipt and the loyalty points all still happen, each in its own service, on its own schedule, and none of them can take payments down.