Writing
From Carrier Networks to the Cloud: What 20+ Years Taught Me About Infrastructure Costs
I have paid for infrastructure from the carrier network to the company that owns the invoice. Each layer taught me something different about cost, and together they are why I never evaluate a technical decision on whether it works alone.
By Ka Lun Chan · Software architecture · Leadership / Architecture / Founder story
Seven seats, one bill
I have looked at infrastructure cost from more seats than most people get to sit in: as a network engineer at a carrier, as the person racking the servers, as a telecom engineer pricing a phone call, as a software developer and architect, as the infrastructure and DevOps engineer, as an engineering leader with a headcount budget, and as a founder whose own company paid the invoice. The view from each seat is different. The bill is the same bill.
Infrastructure cost is never one number. It is capital spent ahead of demand, recurring consumption, carrier and vendor charges, the engineers who run what was built, and the complexity that taxes every change afterwards. Having paid each of those from a different seat is why I do not evaluate a technical decision on whether it works. I evaluate it on what it costs to operate as the business grows.
This post is about what each layer taught me about money. The forecasting method is in capacity planning, from carrier networks to the cloud, and the architecture decisions have their own post on the hidden cost of software architecture, linked further down. I will not repeat either here.
Carrier networks: capacity was capital
At a national carrier, capacity was something you bought. Routers, line cards, switch ports, circuits between sites. Adding it meant a purchase order, a shipping date, a circuit provisioning window, a configuration change, and coordination across the teams that owned each of those steps. You could not try something and roll it back by lunch. A decision to add capacity tied up capital for years, and a decision not to showed up as congestion that customers could feel.
That is where my understanding of infrastructure economics was formed. Every link had a utilization graph and every graph had a date on it after which the hardware would not arrive in time. Running the network security and IP services side, with Cisco access control lists and later a managed firewall and VPN service, added a second lesson: security services ran on the same shared infrastructure, so a capacity decision on a link was also a security and a reliability decision. Nothing was ever just one kind of cost.
Data centers: what a server really costs
At the company I co-founded, I racked the servers, routers and switches myself, cabled them, configured the network, got the circuits turned up and installed Linux on every box before I wrote the software that ran on it. Later I ran the server farm behind more than eight million VoIP minutes a month at a voice services company. Physical infrastructure teaches you the price of things in a way a console never does, because you have to carry it.
The costs I dealt with directly were the hardware itself, the rack space and connectivity in the colocation facility, and the operational support: the technician visits, the spare parts, the redundancy you buy so one failed drive does not become an outage. The wider list that anyone running a data center learns to count includes the equipment lifecycle, meaning the day the box is out of support arrives whether you planned for it or not; power and cooling, which scale with the hardware whether it is busy or idle; and utilization, because a server at ten percent load costs the same as one at eighty. The number that mattered most was the last one. In a data center you pay for peak capacity every day of the year, and most days are not the peak.
I would not trade that experience for anything. When someone proposes a deployment today, I can picture the cabinet, the cables and the invoice, and I ask the questions that experience taught before the architecture is drawn.
VoIP: the cheapest route is not always the right one
Telecom is where capacity, performance and cost became one problem for me, because a phone call touches all three at once. On the platform we built with Asterisk, OpenSIPS, MediaProxy and Python business logic through AGI, every call consumed signaling capacity, media bandwidth, proxy capacity and server channels, and every call paid a termination charge to whichever carrier completed it.
Least-cost routing chooses, per call, the cheapest carrier that can complete it. The naive version picks the lowest rate and moves on. We improved ours with call statistics: answer rates, call durations, post-dial delay and quality by route, fed back into the routing decision. The reason is that the cheapest route is often cheap for a reason. It fails to connect, it connects late, or it connects with audio nobody wants to pay for twice. A customer who hangs up and redials because the quality was bad costs you the minutes, the retry, and sometimes the customer. A route that is slightly more expensive and reliably good is the higher-margin route once you count what bad quality costs. That is a business decision with an engineering implementation, and I have written about it as routing as a business decision.
The other improvement was choosing the media path by the caller’s location, estimated from the source IP address. Audio that hairpins through a single site crosses an ocean twice; audio steered to the geographically appropriate media path arrives sooner, sounds better, and uses less of the bandwidth we paid for. One routing change moved quality and cost in the same direction, which is the kind of change I look for first in any system.
Cloud: capital became consumption
Moving from the data center to the cloud changed the shape of the cost without changing its nature. Capacity stopped being capital bought ahead of demand and became consumption billed behind it. That removes the dark hardware problem and introduces a subtler one: waste is the default state of a cloud account, because nothing stops an over-sized instance, an unused volume or a forgotten environment from running forever.
As CTO of a media publishing platform I led a re-platform onto AWS, and the infrastructure improvements that came with it contributed to AWS costs about 30% lower. Right-sizing instances that had been bought for a peak that never came back. Scaling the web tier on real traffic. Fixing queries and connection handling so the database could be the size it should have been. Caching pages that did not change between requests. Finding resources nobody owned. It is the data center discipline applied to an invoice, and none of it is clever. On client work through Yippify the toolkit is newer, ECS and Fargate, Lambda, RDS, Redis, S3 and CloudFront, with Terraform and GitHub Actions so every change is reviewed before it costs money, and the questions are the ones I asked about cabinets.
Architecture: every box has a price
Software architecture is an infrastructure cost decision in disguise. Each service is compute, monitoring and attention. Kafka is a cluster and a team. A managed database trades a premium for someone else’s pager. A query without an index is a bigger instance next quarter. Buying a capability is often cheaper than building and maintaining one, unless the capability is the product. I lay those tradeoffs out in the hidden cost of software architecture, and the principle underneath them is simple. Good architecture is the right solution for the business, and the newest technology is rarely what makes it right.
People and process cost too
Infrastructure optimization does not stop at the servers. As an engineering leader I plan team capacity the way I planned links: measured, with lead times, with headroom. Hiring adds delivery capacity after months, not on the day the offer is signed, and only if the work can be divided. Technical debt is interest paid on every change. Operational overhead grows with every service added and every process introduced. Incidents consume the week after as well as the night of. Deployment friction is a tax on every release.
So adding engineers does not automatically add throughput, and adding microservices, Kubernetes or another process does not automatically make a system better. Each one is a tradeoff with a cost on both sides, and the job is to weigh it from the engineering side and the business side at the same time. The pattern of a team whose load exceeds its capacity is the one I describe in what sprint carryover is telling you.
The founder’s questions
The seat that changed me most was the one where the company was mine. As an engineer you ask whether a solution works. As a founder, and now running Yippify, the questions that follow are the ones that decide whether the business survives the solution. Can we afford to operate it? What will it cost at ten times the usage? Who maintains it when the person who built it is busy? Is the extra complexity justified by a problem we have today? What happens if demand jumps? And the one that catches the most spending: are we solving a real business problem, or an interesting one?
Those questions shaped how I run my own company. Budgets are small and real, so infrastructure gets chosen for the product’s actual maturity and measured demand, not for the company it might become. Software subscriptions get reviewed like headcount, because each one is a recurring cost with a quiet growth curve. Build-or-buy is decided by whether the capability is what clients pay us for. Complexity has to earn its place, because every piece of it is maintained by a small team that also has to ship. Technical quality is balanced against business need, which sounds like a compromise and is really the whole job.
I have carried those questions back into every engagement. A client does not need a system that would impress a conference. They need one that does the job, that their team can own after hand-over, and that costs what they expected it to cost next year.
Every layer, on one page
Here is the whole stack as I have worked through it, with the capacity challenge and the cost consideration at each layer. The point of the figure is the breadth. Many engineering leaders know one or two of these layers deeply. I have paid the bill at all of them.
Carrier networks & security
- Experience
- Cisco routing and ACLs, NetScreen managed security, operations across 2,000+ collocations.
- Capacity challenge
- Links and ports fill months before replacement hardware can arrive.
- Total cost of ownership
- Capital spent ahead of demand: too much sits dark, too little congests.
What I built
Integration, expansion and capacity planning across the footprint; process, training and escalation for security and IP services.
Technologies
Cisco routing, ACLs, ATM, IP, VoIP, SNMP and Cacti, Expect and Korn shell, an R forecasting program.
Lesson carried forward
Measure first, order ahead of the lead time, and write the runbook for the fourth person.
Case study: capacity across 2,000+ collocationsPhysical infrastructure & data centers
- Experience
- Racked and cabled the servers, configured the network, ran Linux in production.
- Capacity challenge
- Every expansion is hardware, circuits, power and a technician on site.
- Total cost of ownership
- Hardware lifecycle, colocation, connectivity, power and cooling, redundancy.
What I built
The first data center for my own company, from empty rack to production; later the server farm behind 8M+ VoIP minutes a month.
Technologies
Dell servers, Cisco routers and switches, Linux, circuits and colocation.
Lesson carried forward
You pay for peak capacity every day of the year, and the cheapest capacity is the capacity you never consume.
The build-out, with photosTelecommunications & VoIP
- Experience
- Asterisk, OpenSIPS, MediaProxy and Python AGI; a server farm behind 8M+ minutes a month.
- Capacity challenge
- Signaling, media, proxies and servers fill at different rates.
- Total cost of ownership
- Carrier termination and bandwidth, with routing that trades quality against price.
What I built
A VoIP platform with least-cost and geolocation routing, in-house billing, and mobile apps on Nokia, iPhone and Android.
Technologies
Asterisk, OpenSIPS, MediaProxy, SIP, Python AGI, C++ on Symbian.
Lesson carried forward
Quality, security and cost are one decision; a control that drops calls is an outage.
What VoIP taught me about distributed systemsSoftware development & architecture
- Experience
- Multi-tier platforms, a modular monolith to 400k+ users, event-driven systems, APIs.
- Capacity challenge
- Boundaries that scale without multiplying what has to be run.
- Total cost of ownership
- Every service, cluster and query is a recurring line on the bill.
What I built
The platform behind 400k+ users, a publishing platform re-built from WordPress to Rails, event-driven services on Kafka, APIs with authorization on the record.
Technologies
Python, Django, Rails, React, Next.js, Node.js, PostgreSQL, Redis, Kafka.
Lesson carried forward
A boundary should pay for itself, and the rent is due monthly.
Should it be a microservice?Cloud infrastructure & DevOps
- Experience
- AWS with ECS and Fargate, RDS, Lambda, Terraform and CI/CD; about 30% lower AWS costs on a re-platform.
- Capacity challenge
- Capacity in minutes, and waste as the default state of the account.
- Total cost of ownership
- Right-sizing, demand-based scaling, managed versus self-hosted.
What I built
One data center to cloud regions on three continents; a publishing re-platform onto AWS; client platforms provisioned with Terraform and shipped through pipelines with security and accessibility gates.
Technologies
AWS, ECS and Fargate, Lambda, RDS, S3, CloudFront, Terraform, Docker, GitHub Actions.
Lesson carried forward
Capacity became a bill. Measure first and plan from data, the same as on the carrier network.
Capacity planning, from carrier networks to the cloudEngineering leadership
- Experience
- Teams of 30+, CTO and VP of Operations roles, hiring and delivery planning.
- Capacity challenge
- Engineering capacity fills up too, with a longer lead time than compute.
- Total cost of ownership
- Headcount, technical debt and operational overhead, planned as a budget.
What I built
Engineering organizations at a startup through acquisition, at a media publisher, and on client teams; engineers mentored until they could replace me.
Technologies
Written decisions, small pull requests, frequent deploys, security and accessibility as release gates, honest delivery status.
Lesson carried forward
The measure of a lead is the team that works without them.
How I run engineeringStartup founder & Yippify
- Experience
- Co-founded a platform and led it to acquisition; runs Yippify today.
- Capacity challenge
- Every technical choice is paid for by the business making it.
- Total cost of ownership
- Can we afford to run it, who maintains it, and is the complexity justified?
What I built
A communications company from the first line of code to acquisition; Yippify; FedPath, VeloWise, SurgeIQ and Useful Little Tools.
Technologies
Whatever the business needed that week, from product and marketing to cabling and billing.
Lesson carried forward
Working code is the entry fee. Whether the business can afford to run it is the decision.
Being every department at once
What visitors, clients and hiring managers should take from it is this. I understand how technology works, and I understand how technology decisions affect performance, reliability, scalability, operating cost and business outcomes, because I have been responsible for each of those from the physical network up to the company paying to run it.
Short answers
How do infrastructure costs differ between carrier networks, data centers and the cloud?
Carrier and data center capacity is capital bought ahead of demand, with long lead times, lifecycle, power, cooling and rack space, and you pay for peak capacity every day. Cloud capacity is consumption billed behind demand, available in minutes, where waste is the default state unless someone right-sizes, scales on real traffic and removes what nobody owns.
Why is the cheapest VoIP route not always the best one?
Because cheap routes are often cheap for a reason: they fail to connect, connect late or carry poor audio. A customer who redials costs the minutes, the retry and sometimes the customer. Feeding call statistics such as answer rates, durations and post-dial delay back into least-cost routing picks the route with the better margin once quality is counted.
How does being a founder change how an engineer evaluates technology?
An engineer asks whether a solution works. A founder also asks whether the business can afford to operate it, what it costs at ten times the usage, who maintains it, whether the complexity is justified by a problem that exists today, and whether it solves a real business problem rather than an interesting one.
From the network layer to the business layer
More than two decades across carrier networks, data centers, VoIP, software architecture, cloud infrastructure, engineering teams and my own businesses, making technology work at scale.
Performance, reliability, capacity and cost are one decision, and I have owned all four at every layer.
That is the judgment I bring to an engineering organization or a client engagement: whether it works, what it takes to run, and what it costs to grow.