Section650

Writing

TopicEngineering judgment
Reading9 min

Different Industries, Same Engineering Problems: From Carrier Networks to VoIP and Digital Publishing

What building network capacity forecasting tools, optimizing VoIP call routing, and analyzing website traffic taught me about engineering: the products changed, the data changed, and the method stayed the same.

By Ka Lun Chan · Software architecture · Engineering judgment / Architecture / Leadership

Three systems that didn’t look alike

Early in my career I worked on carrier networks: routers, switches and the circuits between them. Later I built VoIP systems that routed telephone calls while trying to keep cost and call quality in balance. Then I ran the technology for a digital publishing business, where the raw material was website traffic, search rankings and advertising revenue. One of those moves packets, one connects calls, and one attracts readers and earns from their attention. Nobody would put them in the same category.

Carrier capacity planning, VoIP call routing and publishing analytics are the same engineering problem with different data. Each one collects operational measurements, finds the pattern in them, forecasts or spots an opportunity, makes a decision with money attached, measures what happened, and adjusts. The products changed, and so did the definition of success. The method did not.

I did not see that at the time. Each job felt like learning a new field from the beginning. It took the third one to notice that the lessons from the first two had been quietly doing most of the work.

Carrier networks: when to buy capacity

On a carrier network, capacity could not be added by clicking a button. A link approaching its limit meant understanding the traffic, forecasting demand, planning the upgrade and coordinating with carriers and equipment vendors, on lead times measured in months. Upgrading early spent money on capacity that sat idle. Upgrading late put congestion in front of customers. The question that mattered was when a link would fill, and the answer had to arrive before the lead time ran out.

I worked on an R program that took network statistics collected over SNMP, the protocol that lets you read interface counters and other operational numbers off network equipment, and used them to analyze utilization on each connection and forecast when it would need more capacity. SNMP gave us visibility. The program’s job was to turn that visibility into a date.

A hypothetical link shows why the obvious number was the wrong one. Say the connection carries 1 Gbps and its average utilization is 400 Mbps. That sounds comfortable. But if traffic reaches 850 Mbps every evening, and those peaks are higher each month, the average is hiding the problem. What we needed was the trend in the peaks, the growth rate, and the headroom left at the busy hour, which is what the R analysis produced: for each link, an estimate of when it would cross the threshold we had decided on in advance.

Fig. 686-1 Why the average hid the problem: one hypothetical day on a 1 Gbps linkAverage 400 Mbps looks comfortable. The evening peak is at 85% of capacity, and it was growing every month.
Bars show utilization by hour, low overnight and rising through the day. The daily average is 400 megabits per second. The peak at 8 pm is 850 megabits per second, 85 percent of the 1,000 megabit capacity.0500100012 am6 am12 pm6 pm11 pmMbpsCapacity 1,000 MbpsAverage 400 Mbps8 pm: 850, 85%

The goal was an answer to a business question, when to invest in more capacity, with reliability on one side of it and spending on the other. That taught me something I have leaned on ever since. More data does not produce better decisions. Understanding what the data means, and what action it calls for, does.

VoIP: where to send a call

VoIP replaced the bandwidth question with a routing one. The platform ran on Asterisk, OpenSIPS, MediaProxy, Linux and software we wrote, and the questions were familiar anyway. How much traffic can the system carry? Where are the bottlenecks? Which routes perform? Where should a call go? How do we keep quality acceptable without spending more than we need to? And how do we know whether any of those decisions is working?

Least-cost routing decides which carrier completes a call. Carriers charge different rates per destination, and across a large volume of minutes a small difference per minute becomes a real number on the bill, so the naive version picks the cheapest route and moves on. The cheapest route is frequently cheap for a reason. It fails to connect, connects slowly, or connects with one-way audio and delay, and a customer who stops using the service has made the cheap route very expensive.

So we fed call statistics back into the routing decision. The measurements available to a VoIP operator include the answer-seizure ratio, the share of call attempts that are answered; the average call duration, which drops when quality is bad because people hang up; completion and failure rates; the carrier’s rate for the destination; call volume; route availability; and latency and media quality where it can be measured. Each one answers a different question. One carrier is cheap with a poor answer rate. Another completes everything and charges for it. A third is excellent to one country and useless to the next. With history per route and per destination, the routing decision could balance cost against quality instead of chasing the lowest rate. We also used the caller’s source IP address to route media through the geographically appropriate path, which improved the audio and reduced what we paid to carry it.

Different data, same shape. Instead of deciding when to upgrade a link, we were deciding where to send a call, and the decision was once again a cost on one side and the customer’s experience on the other. I have written about the rest of that platform in what open-source VoIP taught me about distributed systems.

Publishing: which page to improve

Digital publishing looked like a different planet. Content management systems on WordPress, Rails and platforms we built, search engines, advertising networks, and a practically unlimited number of opportunities to attract a reader. The business model was different too: traffic earned revenue through advertising and affiliate programs, and a good article could keep bringing search visitors for years. The engineering questions were the ones I had been asking for a decade. Where is the traffic coming from? Which pages perform, and which underperform? What pattern explains the difference? Where should the next unit of effort go, and how do we improve results without raising costs?

The analytics covered organic search traffic, rankings, page views, traffic sources, engagement, advertising revenue, revenue per thousand page views, affiliate conversions, and performance per piece of content. A hypothetical pair of articles shows why the obvious number was wrong again. Article A draws 100,000 page views a month and earns $500. Article B draws 20,000 and earns $600. On traffic, A wins. On revenue per thousand views, B earns six times as much, from a fifth of the audience. The interesting questions follow from that: why does B monetize better, is its audience different, does the topic carry commercial intent that advertisers pay for, and could we build more useful content around it? The goal was the right audience, at a value that supported the business, and the largest audience was not always that.

SEO was the same loop with a different lever. A page already drawing visitors but ranking below competitors for the searches it answered was an opportunity: improve the content, fill the gaps, fix the technical problems, strengthen the internal links, and measure. Rankings were never guaranteed and the competition never stood still, but the process was the one from the network: collect, identify, improve, measure, repeat. The business side of that era is in what digital publishing taught me.

Three industries, one approach

Fig. 686-2 One method in every industry: measure, analyze, forecast, optimize, validateThe last arrow is the one teams skip. Validation feeds the next measurement, and that is where the model gets corrected.
Five steps arranged in a circle: measure what is happening, analyze the pattern, forecast what happens next, optimize what to change, validate whether it worked. An arrow from validate back to measure closes the loop.MeasureWhat is happening?AnalyzeWhat is the pattern?ForecastWhat will happen next?OptimizeWhat do we change?ValidateDid it work?Then adjustand go around again
Fig. 686-3 Three industries, one method: measure, analyze, predict, optimize, validate
Carrier networksVoIPDigital publishing
Primary dataSNMP interface statisticsCall and carrier statisticsWeb and search analytics
Main challengeCapacity and congestionCost against call qualityTraffic against monetization
Optimization goalReliable capacity at a reasonable costGood calls at a sustainable marginValuable traffic and revenue
Key decisionWhen to upgrade a linkWhich carrier or route to useWhich content to improve or create
Business impactInfrastructure spend and reliabilityTermination cost and customer experienceAdvertising revenue and profitability

Put side by side, the differences are all in the nouns. The verbs are the same in every column: measure, analyze, predict, optimize, validate. That sequence is the thing I actually carried from job to job, and it is why the third industry was easier to learn than the first.

The hard part was never collecting the data

Collecting metrics is easy compared with deciding from them. A monitoring system gathers thousands of measurements. A VoIP platform writes a record for every call. A website tracks every interaction a visitor has. More dashboards do not make better decisions, and in each of these industries the most available metric was a trap. Average utilization hid the peaks that caused outages. The lowest carrier rate hid the calls that never connected. Page views hid the articles that lost money. Optimizing the wrong metric made the business worse while the chart went up.

So I try to understand the business objective before I decide what to measure. Are we cutting cost, improving reliability, growing revenue, improving the customer’s experience, or making room for growth? The answer decides which metrics matter and how to read them. A number without that context is decoration.

Reality doesn’t follow the model

History helps predict the future and never guarantees it. Traffic patterns change. Calling behaviour changes. Search algorithms change. A model that was right last year is quietly wrong this year if the conditions underneath it moved, and nobody sends a notice. Forecasting is therefore a loop. On the network that meant comparing predicted utilization against what the links actually carried. In VoIP it meant checking that a routing choice was still delivering acceptable quality at the expected cost. In publishing it meant measuring whether the content changes produced the visibility, engagement or revenue we expected, or just the work.

A useful system also tells you afterwards whether its recommendation was right, and that second half is the part teams skip.

Cloud and AI: same questions, new tools

Cloud infrastructure runs the same loop with better instruments. We watch CPU, memory, database performance, request rates, latency and spend; we forecast demand, find the bottleneck, decide when to scale, and ask whether an architectural decision is worth what it costs. The tools moved from SNMP counters and an R script to cloud monitoring, SQL and Python. The questions did not move at all: do we have enough capacity, are we using it well, what happens when demand jumps, what will it cost, and what should we change. I describe that side in capacity planning from carrier networks to the cloud. The carrier years gave me the foundation long before the cloud was part of my daily work.

AI adds a layer on top. Assistants can read logs, investigate an anomaly, summarize a trend and propose optimizations faster than a person can, and I use them for exactly that. The distinction that still matters is between finding a pattern and making a good operational decision. An assistant will recommend cutting capacity because average utilization is low, when the capacity exists for the peak, or for the day a region fails. It will recommend the cheaper carrier without the answer rate in front of it. It will recommend three hundred articles on popular keywords without asking whether any of them helps a reader. The tools are getting better quickly. The business and operational context that turns a pattern into a decision is still supplied by someone who has been wrong about it before.

What it taught me about leading engineers

These experiences shaped how I run engineering teams more than any management book did. I do not believe engineering should be disconnected from business outcomes, and the three industries gave me a concrete test: when we build something, what problem does it solve, and how will we know whether it worked? A technically impressive system is not automatically a successful one. A network upgrade that adds capacity nobody needs is sound and wasteful. A routing algorithm that minimizes cost while ruining call quality is a failed optimization. A publishing platform that draws traffic and spends more than it earns is not a business.

Performance, reliability, scalability, complexity and cost are connected, and understanding how they connect is what lets an engineering team talk to the rest of the company. Instead of “we need more servers”, we can say what demand we are forecasting, what risk we are managing and what the investment buys. Instead of “we improved performance”, we can say what it did to customer experience or operating cost. That is the difference between implementing technology and understanding its effect on the business, and it is the thing I look for when I hire and the thing I try to teach. The way I frame it with my own teams is in how I run engineering.

Short answers

How do you forecast network capacity from SNMP data?

Collect interface utilization over SNMP, then analyze the trend in peak-hour traffic rather than the average, because a link averaging 40% can already be hitting 85% every evening. Project the growth in the peaks against a headroom threshold decided in advance, and the result is a date by which each link needs more capacity, ahead of the hardware lead time.

Why is the cheapest VoIP route not always the best?

Because cheap routes often fail to connect, connect slowly or carry poor audio. Feeding call statistics such as the answer-seizure ratio, average call duration, completion rate and media quality per route and destination back into least-cost routing lets the system balance cost against quality instead of chasing the lowest rate.

What do carrier networks, VoIP and digital publishing have in common for an engineer?

The same loop: measure, analyze, predict, optimize, validate. The data differs, SNMP counters, call records or web analytics, and so does the decision, when to upgrade a link, where to send a call, which page to improve. In each one the most available metric hid the real problem, and the decision had a cost on one side and a customer on the other.

Different products. Same problems.

Carrier capacity, call routing, website traffic, cloud infrastructure: every one of them was limited resources, changing demand, performance that mattered, costs that mattered, and decisions made with incomplete information.

Measure, analyze, predict, optimize, validate. Then check whether the decision was right, and adjust.

After more than two decades, using the information we have to make a system work better and create more value for the business is still one of my favorite problems to solve.

Tell me what you’re measuring