Section650

Writing

TopicArchitecture
Reading7 min

The Hidden Cost of Software Architecture: Why Scaling Too Early Is Expensive

Architecture diagrams have no price column, and they should. Every component runs somewhere, is watched by someone and is paid for every month, and scaling before the demand exists is one of the most expensive mistakes a growing company makes.

By Ka Lun Chan · Software architecture · Architecture / Engineering judgment / Leadership

Every architecture decision has a bill

Architecture diagrams do not have a price column, and they should. Every box on the whiteboard is a thing that runs somewhere, is monitored by someone, fails on its own schedule and is paid for every month. Most teams count the engineering cost of building a system. Fewer count the cost of operating it, and almost nobody counts it at the design review, which is the only place it is cheap to change.

Software architecture drives operating cost through the number of components that must run and be watched, the way they communicate, how much of the work a vendor does for you, how efficiently the database is used, and how much capacity is bought ahead of demand. Scaling too early spends on the last of those before there is any demand to justify it, and it is one of the most expensive mistakes a growing company can make.

I have paid that bill from several seats: as a founder watching a data center invoice land before the revenue did, as a CTO who re-platformed a media business onto AWS and brought the AWS bill down by about 30% along the way, and as an advisor to teams that had built for a million users and were serving a few thousand. The patterns repeat.

Monolith or microservices: count the boxes

Microservices buy flexibility: independent deployment, separate scaling, a boundary a team can own. They also multiply the things that cost money. Each service is its own compute, its own logs and metrics, its own deployment pipeline, its own set of failure modes, and its own share of someone’s attention. Traffic between services is not free either, in latency or, across availability zones, in actual dollars. Ten services serving the traffic of one modest monolith will usually cost more to run and considerably more to operate, and the operating cost is paid in engineers.

The platform I co-founded ran as a modular monolith for most of its life and scaled to hundreds of thousands of users that way. The boundaries were in the code, which cost nothing to cross, and we split things out only where a real reason appeared. I wrote the longer argument in what it costs to own a service for five years. The short version: a boundary should pay for itself, and the rent is due monthly.

Synchronous or event-driven: Kafka is also a team

Event-driven architecture decouples producers from consumers and lets each side scale on its own. I have built on Kafka and I would do it again for the right workload. The right workload is the part that gets skipped. A Kafka cluster is brokers, storage, replication, a schema registry, consumers that have to be idempotent, and an operations burden that lands on whoever is nearest when a partition rebalances at the wrong moment. A managed Kafka service removes some of that and replaces it with a line on the invoice that grows with throughput and retention.

For a system with real fan-out, high write volume or several consumers that must not block each other, the decoupling is worth every dollar. For a service that calls two others in order, a queue or a plain HTTP call does the job for a fraction of the cost and with far less to understand at three in the morning. The tradeoffs of Kafka have their own post, and the event-driven platform case study shows where it earned its place.

Managed or self-hosted: pay in money or in people

Managed services charge a premium over the raw compute underneath them. In exchange, somebody else patches, backs up, fails over and carries the pager. For a small team that premium is almost always cheaper than the engineer-hours it replaces, and the engineer-hours are the scarcer resource. On client work I default to managed databases, managed container orchestration and serverless for occasional work, and the team gets to spend its time on the product.

The math flips at scale. A large, steady workload can turn a managed premium into the cost of a full-time engineer, and at that point self-hosting with a competent team can be the cheaper option, provided the team actually exists. The mistake in both directions is the same one: deciding once and never revisiting. The line moves as the workload grows and as the team changes, and I put a reminder on the calendar to check which side of it we are on.

The database is where money leaks

If I had to pick one place where infrastructure cost hides, it is the database. A query that scans a table instead of using an index runs fine at ten thousand rows and needs a bigger instance at ten million. An index nobody uses still costs writes and storage. An application that opens a new connection per request will push a database to its connection limit long before it is busy, and the usual response is to buy a larger database rather than add a connection pool. An N+1 query pattern turns one page view into a hundred round trips and shows up on the bill as a database that is mysteriously always at 80%.

None of that appears on a whiteboard, and all of it is architecture: decisions about how the system uses its most expensive, least elastic resource. On the publishing platform, query fixes, connection handling and caching were a large part of what let the database be the size it should have been instead of the size fear had made it.

Scaling for users you don’t have

The most expensive version of all of this is building for a million users when the product has a few thousand. Multi-region before there is a second region of customers. A service mesh for four services. Kafka for a signup form. Autoscaling groups that never scale because the baseline was set for a launch that is still a slide. Each of these costs money every month, and worse, costs engineering attention that a young product needs for finding out what customers want.

The argument for it is always the same: we will need it eventually, and it is harder to add later. The second half is sometimes true and usually not. What is hard to change later is the data model and the boundaries. What is easy to add later is capacity, because capacity in the cloud is a purchase you can make in an afternoon. So I spend the early effort on clean boundaries and a data model I can live with, and I buy the capacity when the measured demand asks for it. I learned that on a platform that served three continents from one data center for longer than fashion would have allowed, and moved when the latency and the cost structure, not the diagram, said it was time. The capacity side of that story is in what carrier networks taught me about scaling cloud infrastructure.

How I decide

Five inputs settle most of these questions. The size of the team, because every component is carried by people and a small team can carry only so many. The maturity of the product, because a system still looking for its market needs flexibility more than throughput. The actual traffic, measured, not the number in the deck. The growth we have evidence for, as opposed to the growth we hope for. And the business constraints: budget, compliance, what the client can own after hand-over.

Fig. 683-1 Three decisions, and which way I lean by default
DecisionMy defaultWhen I go the other way
Build or buyBuy when the capability is not what customers choose you for, a vendor does it well, and the integration is thin.Build when it is the product, when the vendor’s pricing scales faster than your revenue, or when the integration would be most of the work anyway.
Managed or self-hostedManaged for a small team: the premium is cheaper than the engineer-hours to run it, and the on-call is someone else’s.Self-host when the workload is large and steady enough that the premium is a salary, and you have the people to carry the pager.
Scale now or scale laterLater, by default: keep the boundaries clean so scaling is possible, and spend the money when measured demand, not a forecast, asks for it.Now, only for the one resource that fills first and has a real lead time, such as a data model that is painful to change once it holds customer data.

The common thread is to be honest about which resource is scarce. Early on it is engineering attention and cash, so the right architecture is the one that spends the least of both while keeping the doors open. Later it may be compute, or the database, or the team’s ability to understand the system, and the right architecture changes with it. The job of a technical leader is to keep asking which resource is scarce this year, and to make sure the design review has a price column.

Short answers

How does software architecture affect operating costs?

Through the number of components that must run and be monitored, how they communicate, how much a vendor does for you, how efficiently the database is used, and how much capacity is bought ahead of demand. Each extra service, cluster or cross-zone call is a recurring cost, paid in dollars and in engineer attention.

Why is scaling too early expensive?

Because capacity in the cloud is easy to add later and costs money every month until it is needed. Multi-region, service meshes or Kafka for a product with a few thousand users spend cash and engineering attention a young product needs elsewhere. What is hard to change later is the data model and the boundaries, so that is where early effort belongs.

When are managed services more expensive than self-hosting?

For a small team the managed premium is almost always cheaper than the engineer-hours it replaces. The math flips for large, steady workloads where the premium approaches an engineer’s salary and the team exists to carry the pager. The line moves as the workload and the team change, so the decision needs revisiting.

Put a price column on the whiteboard

Every box in an architecture diagram runs somewhere, is watched by someone and is paid for every month.

Spend early effort on boundaries and the data model, which are hard to change. Buy capacity when measured demand asks for it, which is easy.

That is how I keep technical decisions connected to what the business pays to run them.

Talk through an architecture decision