Writing
Event-Driven Architecture with Kafka: Should Systems React to Events Instead of Calling Each Other Directly?
Event-driven architecture changes how systems coordinate, how failures show up and who carries the operational load. Kafka can support it, but the decision needs a concrete business benefit and an honest look at the cost.
By Ka Lun Chan · Software architecture · Architecture / Leadership
When one system should call another directly
Event-driven architecture changes three things at once: how systems coordinate work, how failures show up, and who carries the operational load. I’ve worked with Kafka-based microservices, alongside SaaS and government applications on AWS with CI/CD and distributed teams, and those three changes are what I look at before anything about the technology. My first lessons in all three came long before Kafka, running an open-source VoIP platform where a lost message was a dropped call.
A direct call is the right default when the caller needs an answer now. A checkout page has to know whether the payment was authorized. A login form has to know whether the password was right. The request flow is easy to follow: one call, one response, one place to look when it fails.
The cost is dependency. The caller’s latency includes the callee’s latency. If the callee is down, the caller is partly down too. Chain four services together synchronously and a slowdown in the last one backs up all the others, which is how one slow dependency turns into a wider outage.
When a system should publish an event instead
Take a hypothetical online store. A customer places an order. Several things should follow: reporting needs updated sales numbers, the customer should get a confirmation email, and fulfillment should start preparing the shipment.
With direct calls, the orders service would call each of those in turn and wait on all of them. With events, it saves the order and publishes one fact: OrderPlaced. Reporting, notifications and fulfillment each react on their own. A new consumer, say a fraud check, can be added later without changing the orders service.
- Orders service to orders topic
- orders topic to Reporting
- orders topic to Notifications
- orders topic to Fulfillment
The wording of that event matters. OrderPlaced is a fact about something that already happened, owned by the service that knows it. SendNotification would be a command, a request for specific work, which makes the sender responsible for knowing who does the work and whether it got done. A fact leaves each consumer to decide what to do, while a command keeps the sender tied to the work it asked for.
Where Kafka fits
Apache Kafka is a durable event-streaming platform. Producers write events to topics, Kafka stores them for a configured retention period, and any number of independent consumer groups read them at their own pace and can replay them later.
In the example, OrderPlaced goes to an orders topic. The topic is split into partitions so the load can spread across brokers and consumers. Reporting, notifications and fulfillment are separate consumer groups, so each one reads every event and tracks its own position. If the reporting consumer has a bug, the team can fix it and replay the retained events to rebuild the numbers.
That retained, replayable log is what Kafka adds. Plenty of asynchronous work doesn’t need it. Sending one welcome email after a signup is a background job, and a simple queue handles it well.
What complexity moves into the background
When the orders service gets an acknowledgment from Kafka, the event is stored. None of the downstream work has happened yet. Everything after that is eventually consistent, and the team has to manage what that means in production:
- Consumer lag: how far behind each consumer is, and who gets alerted when it grows.
- Duplicates: delivery is usually at least once, so consumers will sometimes see the same event twice.
- Ordering: Kafka orders events within a partition, so related events need the same partition key.
- Schema evolution: changing an event without breaking consumers needs compatibility rules.
- Retries and failed events: a malformed message can stall a consumer until something moves it aside.
Each of those needs an owner, a dashboard and a runbook before the first production incident.
Keeping business records and events consistent
The orders service has to do two things: save the order to its database and publish the event. Those are two different systems. If the database write succeeds and the publish fails, the order exists and nobody downstream knows. That’s the dual-write problem.
A transactional outbox is one common fix. The service writes the order and an outbox row in the same database transaction. A separate relay reads the outbox and publishes to Kafka, retrying until it succeeds. On the consuming side, idempotent consumers record which events they’ve handled, so a duplicate does nothing. I used that pattern on an event-driven microservices platform, and it’s the part I’d design first.
Kafka’s exactly-once features help, within a defined scope. Idempotent producers and transactions can make a read-process-write flow inside Kafka behave as if it ran once. They don’t reach outside Kafka, so they can’t make an email, a payment or a write to another database happen exactly once. That still takes idempotency keys and careful design in the consumer.
Events can still create tightly coupled systems
Asynchronous doesn’t automatically mean independent. A shared event schema that every team edits becomes a shared database by another name. A consumer that depends on a field nobody documented breaks when the producer changes it. A business process that needs five consumers to finish in a particular order is a distributed transaction with no coordinator. Each of these brings back the coupling events were supposed to remove, and it’s harder to see. I go deeper on that trap in how I decide whether a microservice should be a microservice.
What Kafka costs to own
The infrastructure bill is the visible part. The total cost of ownership includes:
- Broker capacity, or a managed service’s charges
- Storage for replication and retention, and network transfer, including traffic between availability zones
- Compute for consumers, plus schema registries and connectors if you use them
- Monitoring, security, upgrades and incident response
- Engineering time on schemas, replay, recovery and debugging
- Local development and test environments that behave enough like production
- Opportunity cost: hours spent maintaining infrastructure instead of improving the product
Several factors drive those costs. Throughput, message size, retention and replication drive storage and network charges directly. Partition count and consumer patterns mostly create capacity and operational overhead: more partitions mean more to balance and monitor, and every new consumer group reads the full stream again. Pricing models differ between providers, so check current pricing before you commit.
A managed service, such as Amazon MSK or Confluent Cloud, takes on running and patching the brokers, and some offerings handle scaling for you. Your team still owns topics, schemas, consumers, lag, recovery and the bill. Self-managed Kafka gives you more control over configuration and cost, and it needs people who understand the cluster and an on-call rotation that covers it.
Kafka can also help avoid costs: the same integration rebuilt for each new system, polling jobs that ask “anything new?” all day, user requests blocked on slow downstream work, and data that can’t be reprocessed after a bug. Those savings depend on the design, so I count them only when the system has those problems today.
Direct calls, queues or event streaming
| Approach | Immediate response | Replay | Multiple consumers | Operational effort | Cost |
|---|---|---|---|---|---|
| Direct API call | Yes. The caller waits for the answer. | No. A missed call is gone unless the caller retries. | The caller knows and calls each one. | Lowest, until dependency chains get long. | Mostly the services themselves, plus timeouts and retries. |
| Background job or simple queue | No. Work happens after the response. | Limited. Messages usually disappear once handled. | Usually one worker type per queue. | Moderate: workers, retries and a dead-letter queue. | Often low, especially as a managed queue. |
| Kafka event streaming | No. Consumers react on their own schedule. | Yes, within the retention period. | Many independent consumer groups read the same events. | Highest: schemas, lag, partitions, recovery. | Cluster or managed service, storage, network and engineering time. |
Most systems I’d design use more than one. Direct calls cover the moments where a user is waiting for an answer. Events carry the reactions that can happen afterward, and background jobs take single tasks that don’t need a retained history.
Event-driven architecture also isn’t event sourcing. Event sourcing stores every change to an entity as an event and rebuilds state from that history, which is a much bigger commitment. You can use Kafka to publish events while your database stays the source of truth.
A leadership checklist
- What business problem requires asynchronous coordination?
- Do we need retained history, replay, or multiple independent consumers?
- How much delay can users and downstream systems tolerate?
- Who owns event contracts, consumers and production recovery?
- How will we detect missing, duplicated or delayed processing?
- Could a simpler queue meet the requirement?
- What will this cost to operate over the next year?
If the team can’t answer the fourth and fifth questions, the design isn’t ready, however clean the diagram looks. It’s the same capacity question I raise in your team size should influence your architecture.
Short answers
When should I use Kafka instead of direct API calls?
Use a direct call when the caller needs an immediate answer. Use Kafka when several independent consumers need to react to the same events, when you need retained history or replay, and when the team can operate the resulting lag, duplicates and schema changes. For a single background task, a simpler queue is often enough.
What is the transactional outbox pattern?
The service writes its business change and an outbox record in the same database transaction. A separate relay publishes outbox records to Kafka and retries until it succeeds, so a saved change is never left without its event.
Does Kafka exactly-once mean my side effects happen once?
No. Kafka’s exactly-once features cover reading from and writing to Kafka within a transaction. Sending an email, charging a payment or writing to another database still needs idempotency keys and careful consumer design.
Is event-driven architecture the same as event sourcing?
No. Event-driven architecture uses events to let systems react to each other. Event sourcing stores every state change as an event and rebuilds state from that history. You can use Kafka for events while a regular database stays the source of truth.
So, should systems react to events?
Sometimes. Kafka is a tool for a specific kind of problem, and using it says nothing about how mature an architecture is.
Systems should react to events when asynchronous coordination brings a benefit worth having, and the organization can support the failure behavior and operating cost that come with it.
Where a user needs an answer now, call directly. Where the work is a reaction, publish the fact, and make sure someone owns what happens next.