Section650

Writing

TopicArchitecture
Reading11 min

Software Architecture Through the Years: What Different Systems Taught Me

Telecom, publishing, monoliths, microservices and Kafka each taught me the same thing from a different side: the right architecture depends on the problem, and every choice leaves your team something to operate.

By Ka Lun Chan · Architecture / Leadership / Engineering judgment

The right architecture depends on the problem

I’ve been building software long enough to watch architecture change several times. As a co-founder, I built an open-source VoIP and telecommunications platform, billing system included, from the ground up. I was CTO of a media publishing platform, where WordPress, Ruby on Rails, editorial workflows and SEO all had to coexist. I’ve built traditional monolithic applications, and later I worked with distributed systems and microservices.

I expected the lesson to be that architecture kept improving. What I learned instead is that each kind of system was right for some problems and wrong for others.

The right software architecture is the one that fits the problem, the team and the stage of the business. Monoliths, microservices and event-driven systems each solve some problems and create others, so the useful question is which of those tradeoffs you can afford.

When software and the network were the same problem

Some of my earliest architecture lessons came from telecommunications and VoIP. I ran operations for a voice service and later co-founded a global communications platform.

We built that platform from the rack up. I racked servers, ran the cabling, configured the servers, routers and switches, and wrote the software on top, including our own billing system. Seeing every layer taught me early that software architecture doesn’t stop at the application boundary.

A single call travels through a lot of systems before anyone hears a voice on the other end.

Fig. 662-1One phone call, many places to fail Every hop is part of the system
  • Phone to Access network
  • Access network to SIP infrastructure
  • SIP infrastructure to Application logic
  • SIP infrastructure to Carrier
  • Carrier to PSTN
  • PSTN to Destination

When a call failed, the application wasn’t necessarily the problem. It could be DNS, routing, NAT, a firewall, SIP signaling, a codec negotiation, the carrier, latency or packet loss. Or it could be the application. You couldn’t open one repository and understand the whole system.

Understand the whole path, including the parts outside your component.

That lesson has stayed with me. Today the path runs through Kubernetes, AWS, APIs, queues and databases instead of SIP gateways and carriers, but debugging a distributed system still means understanding what happens between the components. The operations side of that story, from Expect scripts in a NOC to AI agents, is in how DevOps changed during my career.

Publishing taught me a different kind of architecture

From the outside, publishing looks like writing an article and putting it on a website. At scale there’s a lot more to it.

We had different systems serving different purposes, including WordPress and Ruby on Rails applications. Content moved through editorial workflows and reached readers through websites, search engines, social platforms and other channels. Architecture decisions reached well past how the application behaved. They affected:

  • editorial workflows and publishing speed
  • search visibility, URL structure and page performance
  • advertising, analytics and content discovery
  • integrations and migrations
  • years of existing content

A developer could make a technically reasonable decision and create a business problem without noticing. Changing URLs could hurt SEO. Changing how pages rendered could change what search engines saw. Replacing a CMS could disrupt the editors. One migration could break thousands of old links. Slow pages lost readers and search visibility, and eventually revenue.

The business runs through the software, so architecture shapes how the business operates.

I wrote more about this side of the job in what being a CTO in publishing taught me about SEO, AEO and GEO.

The monolith wasn’t the enemy

Like many engineers, I spent a lot of time on monolithic applications, and I think monoliths have an unfair reputation.

A well-designed monolith can be extremely productive. You have one application and one deployment, often with one database, so transactions are straightforward and debugging and local development are easier. You can trace a request through the code without jumping across five services, three queues and two observability platforms. For a small team or a new product, that simplicity is worth a great deal.

The problems start when the boundaries inside the monolith disappear. Everything depends on everything else. A small change requires understanding half the application. Deployments get riskier, tests get slower, and one team’s changes break another team’s work. Parts of the system that need to scale differently can’t.

Eventually the organization’s boundaries and the application’s boundaries stop matching, and the architecture starts pushing back.

Then came microservices

Microservices offered an attractive answer: break the system into independently deployable services. Teams can own individual services. Services can scale on their own. Different workloads can use different technologies, deployments get smaller, and failures can be isolated.

Sometimes you get all of that. You also get new work, because splitting a system up moves its complexity onto the network.

Microservices trade complexity inside one application for complexity across a network. A function call becomes a network request, a database transaction may become an asynchronous workflow, and an exception may become a timeout.

Now the team has to think about service discovery, API contracts, authentication and authorization between services, retries, timeouts, queues, event ordering, idempotency, distributed tracing, schema evolution, deployment coordination and eventual consistency.

Networks fail, which made my telecommunications experience relevant again. I go deeper on when a split is worth it in how I decide whether a microservice should actually be a microservice.

Kafka changed how I thought about communication

Another shift for me was moving from mostly synchronous calls, where service A calls B and B calls C, toward event-driven systems.

Fig. 662-2Event-driven: producers don’t need to know their consumers Looser coupling, harder debugging
  • Service A to Kafka topic
  • Kafka topic to Service B
  • Kafka topic to Service C
  • Service C to Service D

Producers don’t need to know every consumer, consumers process events independently, and services become less tightly coupled. In exchange, a new set of questions shows up:

  • What happens if a message is processed twice?
  • What happens if processing fails halfway?
  • What happens when events arrive out of order?
  • How do we evolve schemas, and how do we replay events?
  • How do we find which service produced bad data?
  • How do we debug a business transaction that crossed six services asynchronously?

The diagram gets cleaner, but operating the system often gets harder. The event-driven platform case study shows how we answered some of those questions with a transactional outbox and idempotent consumers.

Architecture and organizations are connected

One of the biggest things I’ve learned as an engineering leader is that software architecture usually reflects the organization that builds it.

A six-person engineering team rarely needs twenty microservices. You could build them, but then six engineers have to operate twenty services: twenty pipelines, twenty sets of logs, twenty things to monitor and twenty things that can fail.

As teams grow, the calculation changes. If thirty engineers keep stepping on each other inside one application, clearer service boundaries can create real independence. That’s why I no longer treat architecture as a purely technical decision. I ask:

  • How many engineers do we have?
  • How often does this system change?
  • Which teams own which capabilities?
  • Where are the actual scaling problems?
  • What are our reliability requirements?
  • How hard will this be to operate at 2 AM?

And one question that doesn’t get asked enough: what is the simplest architecture that solves the problem we actually have?

I’ve become more skeptical of architecture fashion

Technology has trends like every other industry. Everything needed to be service-oriented. Then everything needed to be microservices. Then Kubernetes, then serverless. Now everything seems to need AI.

I’ve learned to be careful when an architecture conversation starts with a technology instead of a problem. “We should use microservices.” “We should use Kafka.” “We should use Kubernetes.” My answer to each is the same: why? Until someone answers that, they’re technology preferences, and nobody has made an architecture decision yet.

The architecture conversation should start with the constraints. What are we building, and who uses it? How much traffic is there? How quickly will it change? What needs to scale, and what needs to stay reliable? What does the team understand, and what can it realistically operate? Then choose the technology.

Sometimes the best architecture is boring

After working with increasingly complicated systems, I appreciate boring architecture more than I used to:

  • A Django or Rails application on PostgreSQL
  • Redis when it’s actually needed
  • A queue where asynchronous processing provides real value
  • A few well-defined services where there are genuine domain or scaling boundaries
  • Good monitoring, deployment automation and backups
  • Clear ownership

That stack won’t produce many conference talks, but it can run a very successful business.

If an architecture mostly makes the engineering team feel sophisticated, it’s the wrong one.

The full circle

VoIP taught me about distributed systems before I regularly used the term, and publishing showed me that architecture affects much more than engineering. Monoliths taught me the value of simplicity, microservices taught me that boundaries matter, and event-driven systems taught me that decoupling comes with an operational cost. Cloud infrastructure made distributed architectures dramatically easier to build.

After years of racking our own hardware, moving the platform out of our data center into cloud regions on three continents was one of the most fun projects I’ve worked on. We moved to lower our cost structure and to put the platform closer to users, with each user routed to the nearest region.

Engineering leadership added the lesson I use most: every architectural decision creates a future operational obligation. Someone will have to deploy it, monitor it, debug it and eventually change it. And someone will probably get paged when it breaks.

Architecture is about tradeoffs

Earlier in my career, I sometimes thought there was a best architecture. After working across enough systems, I don’t anymore. There are architectures that suit particular problems, teams and stages of a business.

A monolith isn’t automatically bad. Microservices aren’t automatically scalable. Kafka doesn’t automatically make a system usefully event-driven. Kubernetes doesn’t automatically make an application reliable. Another layer of abstraction doesn’t automatically improve anything.

Sometimes the right decision is to break a system apart, and sometimes it’s to keep it together. Sometimes asynchronous processing is the answer, and sometimes a database transaction is perfectly fine. Good architecture is mostly about understanding which of those you’re facing.

Short answers

Are monoliths bad?

No. A well-designed monolith is often the most productive choice for a small team or a new product: one deployment, simple transactions, easy debugging and local development. Monoliths become a problem when the boundaries inside them disappear and every change touches everything else.

What do microservices actually cost?

Microservices move complexity from inside one application onto the network. Teams take on service discovery, API contracts, retries, timeouts, idempotency, distributed tracing, schema evolution, deployment coordination and eventual consistency, plus a pipeline, logs and monitoring for every service.

What are the tradeoffs of event-driven architecture with Kafka?

Kafka lets producers publish events without knowing every consumer, which loosens coupling. In exchange, teams have to handle duplicate and out-of-order messages, failed processing, schema evolution, replays and debugging transactions that cross several services asynchronously.

How does team size affect software architecture?

Architecture has an organizational cost. A six-person team running twenty microservices has to operate twenty pipelines and monitor twenty things that can fail. When thirty engineers keep blocking each other in one application, clearer service boundaries can give them real independence.

How should you choose a software architecture?

Start with the constraints, not the technology: what you’re building, who uses it, traffic, rate of change, what must scale, reliability requirements and what the team can realistically operate. Then pick the least complicated system that solves the problem and can evolve with the business.

The question that survived every trend

The technology has changed enormously over my career. The question I keep coming back to has gotten simpler.

What is the least complicated system we can build that reliably solves the problem, can evolve with the business, and that our team can actually operate?

Talk through an architecture decision