Section650

Writing

TopicArchitecture
Reading7 min

What Open Source VoIP Taught Me About Distributed Systems Before We Called Them Distributed Systems

Partial failure, routing, retries and records that had to reconcile were everyday VoIP problems. Working with Asterisk, OpenSIPS, Python AGI and billing taught me the questions I still ask before choosing any architecture.

By Ka Lun Chan · Architecture

Distributed systems, before I used the term

Distributed systems were a well-established field long before I worked on VoIP. The title describes my own path: I was dealing with the problems every day before the phrase was part of my working vocabulary.

As a co-founder, I built a communications platform on open-source VoIP. The pieces I worked with most were Asterisk, OpenSIPS, Python scripts running through the Asterisk Gateway Interface (AGI), a billing system, and routing informed by where callers were. Years later, when I started hearing microservices debates, a lot of them sounded familiar.

Many of the problems engineers now discuss around microservices were everyday concerns in VoIP: coordinating separate components, routing requests, handling partial failures, tracking state across steps and keeping business records accurate. The components themselves weren’t microservices, but the questions were the same.

A phone call crossed several system boundaries

Each component in a VoIP platform has a different job, and each one fails differently.

OpenSIPS handles SIP signaling: the messages that set up, change and end a call. It decides where a call request should go and passes it along. The audio doesn’t flow through it. The media path, carried over RTP, is a separate stream that can take a different route, which is why a call can be set up perfectly and still have one-way audio.

Asterisk handles the call itself. Its dialplan decides what happens step by step, and it can terminate and process media for features like prompts and recordings. When the dialplan needs a business decision, it can run an AGI script. With plain AGI, Asterisk starts the script as a local process and talks to it over standard input and output. That script is where Python could check an account, look up a rate or choose a route from backend data.

Billing sits after all of that. It turns usage into money: call records, rating, balances and reconciliation.

A hypothetical call, and where it goes wrong

Here’s an illustrative flow. It isn’t the topology of any system I ran, and real deployments differ a lot.

Fig. 666-1Hypothetical flow: one prepaid call across four components Where the billing record is written
  • SIP phone to OpenSIPS
  • OpenSIPS to Asterisk
  • Asterisk to Python AGI script
  • Asterisk to Call detail record
  • Call detail record to Rating & balance

A prepaid customer places a call. OpenSIPS receives the SIP request and forwards it to an Asterisk server. The dialplan runs a Python AGI script, which checks the customer’s balance and picks a carrier route. The call connects and lasts twelve minutes. At hangup, Asterisk writes a call detail record, and a billing job rates the call and deducts it from the balance.

Now suppose the billing job deducts the charge, and the database connection drops before the job hears that the write succeeded. The job assumes it failed and tries again. Without protection, the customer pays twice for one call.

A second problem sits earlier in the flow. If the AGI script can’t reach the balance database, should the call go through? Allowing it risks usage nobody can bill. Blocking it turns away a paying customer because of an internal timeout. That’s a policy decision, and someone on the business side should make it before the outage, not during it.

Why “the system is up” was never the whole answer

In that example, every component could report itself healthy. OpenSIPS routed the call, Asterisk handled it and the customer had a clear conversation. The failure was in a record.

That’s what partial failure looks like day to day. One component can be fine while another times out. A call can connect while the billing update fails. A completed call and a correctly billed call are separate outcomes, and a platform has to get both right. A status page that only asked whether each server was up would have missed the problem customers cared about most.

Routing was a business decision too

Choosing where a call goes involves several inputs that pull against each other: carrier availability, cost per minute, call quality, regulatory policy and the caller’s location.

Location needs care. A telephone number’s prefix tells you where the number was issued, not where the caller is standing; mobile and VoIP numbers travel with their owners. IP geolocation is an estimate, and sometimes a badly wrong one. And physical distance doesn’t guarantee a better network path. A nearby data center can sit behind a congested link while a farther one performs better. Location was a useful signal for routing, weighed with the others, and routing needs a sensible fallback for when it’s wrong.

Retries needed context

“Just retry it” means very different things depending on what you retry.

  • Retrying a rate lookup is usually harmless. Reading the same data twice costs a little time.
  • Trying another carrier before the call is answered can rescue it. Starting a second call after the first one connected rings someone twice.
  • Replaying a billing event can charge a customer twice for one call.

Idempotency means an operation can run more than once with the same result as running it once. In billing, that usually means keying each charge to a unique identifier for the call, so a replayed event finds the charge already recorded and does nothing.

SIP gives every dialog a Call-ID, and Asterisk assigns its own unique IDs to channels. Using an identifier like that as the key turns a dangerous retry into a safe one.

State and records needed reconciliation

A call passes through several stages: setup, an active session, a finished call detail record, and finally a charge on an account. Each stage lives in a different place and can fail on its own.

A call detail record (CDR) records that a call happened and how long it lasted. A billing system also needs the rate that applied, the balance before and after, adjustments and refunds, and a way to match its records against what carriers bill you. Reconciliation is the regular job of comparing those sources and explaining every difference. It only works if the same call can be followed through every stage by a shared identifier.

Debugging meant following the whole path

When a customer reported a problem, the answer was rarely in one log. You matched the SIP messages from signaling to the channel in Asterisk, the AGI script’s log lines, and the rows in billing, and the identifiers had to line up for any of that to work.

Distributed tracing today does something similar with far better tools: one trace ID carried across service calls, with timing for each step. The two aren’t identical, but the habit transfers directly. I wrote about the same lesson from the network side in what different systems taught me about software architecture.

Every boundary came with work

Separating signaling, call handling, business logic and billing made each piece easier to reason about. It also added configuration to keep consistent, monitoring for each component, version compatibility between them, and failure handling at every handoff.

That trade is the same one teams face when they count services today. Each boundary should solve a real problem, and the team has to be able to carry it, which is the argument in your team size should influence your architecture.

The tools improved; the questions stayed

Containers, orchestration and observability tools make this work far easier than it was. They help a team deploy consistently, restart what crashes and see across components. I’ve written about that change in how DevOps changed during my career.

They don’t answer the design questions. Someone still has to decide who owns each piece of state, what happens when a step fails halfway, which operations are safe to retry, and how the system recovers and proves its records are right.

Short answers

What is partial failure in a distributed system?

Partial failure is when some components work while others fail or time out. In VoIP, a call can connect and sound fine while the billing update for it fails, so “the system is up” doesn’t mean every outcome the business needs happened.

What is idempotency in billing?

Idempotency means running an operation more than once has the same result as running it once. In billing, each charge is keyed to a unique identifier for the call, so a retried or replayed event finds the charge already recorded instead of charging the customer twice.

What is the difference between SIP signaling and media?

SIP signaling carries the messages that set up, change and end a call. The media, the audio itself, usually travels separately over RTP and can take a different path, which is why a call can be set up correctly and still have audio problems.

Boundaries and failure first, tools second

VoIP taught me to look at a system through what crosses its boundaries and what happens when one of them fails.

Before choosing tools, understand the boundaries, the failure behavior, who owns each piece, and how the system recovers.

That order still shapes my architecture decisions. The tools are much better now, and I use them gladly, once those four questions have answers.

Talk through an architecture decision