Writing
From Cisco Routers and Carrier Networks to Cloud Infrastructure and AI Agents
The tools changed from Cisco routers and physical servers to cloud regions and AI agents. The questions I ask about a system mostly did not, and most of them I learned at a console port.
By Ka Lun Chan · Still Building After 20 Years · Architecture / Learning / Leadership
Before the cloud, the network was the computer
This is the first story in a short series I am calling Still Building After 20 Years. It is the infrastructure one, and it starts where my engineering career did: on a carrier network, at a console prompt, with a Cisco router on the other end of a serial cable.
The network spanned more than 2,000 collocations, carrying ATM, IP and VoIP traffic for a national broadband provider. The work was routers and switches, access control lists, circuits that went down at inconvenient hours, Linux and Unix boxes doing the monitoring, and a lot of scripts in Expect and KornShell that typed into consoles so a person did not have to. There was no console in a browser and no API to call for more capacity. If a link was full, someone ordered hardware, and someone else drove to a building to install it.
What that kind of work teaches is how to think in paths. When a customer could not reach something, the answer was somewhere between them and it, and the job was to walk the path one hop at a time until you found the hop that disagreed with the others. I have used that habit on every system since, including ones with no routers in them.
Security as an operations job
My security education happened on the same network. The ACLs on the routers were my first security policy, and when the carrier launched a managed firewall and VPN service on NetScreen, I was one of the first three people supporting it in a brand-new Security Operations Center. Later I flew to the East Coast to train the Technical Assistance Center that would support it after us. I wrote the full version of that story in network security long before DevSecOps, so here is only what it changed in me.
I learned that security work is mostly operations. Change requests, troubleshooting a tunnel that is up in one direction, reading a rule set in order with the implicit deny at the bottom, writing it down so the next shift can do it. And I learned that a security control that blocks legitimate traffic is an outage, which is a lesson I find myself repeating to people configuring cloud permissions twenty-something years later.
Teaching the TAC taught me something else: if I could not explain how to troubleshoot it, I did not understand it as well as I thought.
VoIP: where the network met the business
Voice was where infrastructure turned into a business for me. At a voice services company and then at the communications platform I co-founded, the stack was Asterisk for calls, OpenSIPS for SIP signaling, MediaProxy for users behind NAT, Python business logic through AGI, and a billing system we wrote ourselves. I racked the servers for that platform in a data center, cabled them, and later moved it into cloud regions on three continents.
Two pieces of that work belong in this story. Least-cost routing picked a carrier for every call, and we improved it with call statistics so the cheapest route also had to be a good one. And we used the caller’s source IP to send the audio through a media path closer to them, which made calls sound better and cost less to carry at the same time. Both were networking problems with a price tag on them, which is the most useful kind of problem I know. The details are in different industries, same engineering problems.
Up the stack, with the network still in my head
From there my work moved upward. PHP and LAMP backends, including the one behind a hybrid mobile app for a startup. Ruby on Rails for a publishing platform I re-built off WordPress. Python and Django, React and Next.js, PostgreSQL. AWS, Docker, Terraform and CI/CD pipelines. Kafka and event-driven services where the decoupling earned its cost.
I did not leave the network behind. I carried it up. A VPC with public and private subnets is a zone design. A security group is a stateful ACL. A service mesh is routing with opinions. A distributed system is a set of hops that each fail differently, and debugging one is walking the path. When a junior engineer tells me the cloud is magic, I understand why it looks that way, and I also know there is a router somewhere that disagrees.
The network background helped most in the boring places: latency budgets, DNS, timeouts, what happens when one dependency is slow rather than down. Those are the failures that never show up in a demo and always show up in production.
Security belongs to every layer
Having worked at every layer from the wire up, I think of security as a property of each layer, owned by whoever works there, more than as a team down the hall. Network controls decide what can talk to what. Infrastructure controls decide what runs and who can log in. Application code decides, on every request, whether this caller may touch this record. Secrets management decides how many places a key can leak from. Cloud access controls decide how much damage one leaked key can do. Production operations decide how fast anyone notices.
A weakness at any one of those layers is a weakness in the system, and the people who can fix it are usually the people who built it. That is why I want engineers to own the security of what they ship, with specialists to help, rather than the reverse. The habits are the ones the router taught me: deny by default, allow on purpose, log what you deny, and test every rule with the traffic it is supposed to permit.
AI agents are a network problem too
Lately I have been experimenting with AI agents: Hermes Agent on a VPS I administer, Claude Code and OpenAI’s Codex in the terminal, and OmniRoute as a gateway in front of several model providers. The newest part of my career turns out to be the most familiar.
An agent calls a model over the network, with a credential, through a gateway, with a timeout. It calls tools, which are other services with their own permissions. It keeps state somewhere and runs on a schedule on a machine that has to stay up. Every one of those is a thing I learned to worry about on a carrier network. The model is new. The authentication, the authorization, the latency, the retries, the logs, the cost of every request and the boundary around what the agent may touch are not. I wrote about the gateway and the scheduling separately; the short version is that an agent is a distributed system with a very good text parser in the middle.
What carried over
| Where I learned it | Where I use it now |
|---|---|
| A router ACL: deny by default, allow what must pass | IAM policies, security groups, an agent’s tool allow-list |
| Trace the path hop by hop when a circuit fails | Trace a request through gateway, service, queue and model provider |
| Capacity ordered ahead of the hardware lead time | Token budgets and rate limits planned before the bill arrives |
| A runbook so the TAC could support what the SOC built | Instructions and data an unattended agent can run from |
| Least-cost routing weighed against call quality | Model routing weighed against answer quality |
The tools changed almost completely. Routers became APIs, data centers became regions, scripts became pipelines, and now some of the scripts write themselves. The questions did not change much. What can talk to what? Who is allowed? What happens when this hop fails? Who will notice? What does it cost? I learned to ask those at a console port, and I have not found a system yet where they stopped being the right questions.
The next story in the series is about why, after all of that and several leadership roles, I still write code.
Short answers
How does network engineering experience help with cloud architecture?
Cloud concepts map onto network ones: a VPC with public and private subnets is a zone design, a security group is a stateful ACL, and debugging a distributed system is walking a request path hop by hop. Latency, DNS, timeouts and slow dependencies are the failures network engineers learn first.
Are AI agents different from earlier infrastructure?
The model is new; the rest is familiar. An agent calls a model over the network with a credential, through a gateway, with a timeout, calls tools with their own permissions, keeps state and runs on a schedule. Authentication, authorization, reliability, observability, cost and security boundaries all still apply.
The tools changed. The fundamentals mostly did not.
From Cisco routers and physical servers to cloud infrastructure and AI agents, the work kept moving up the stack.
Walk the path, deny by default, plan capacity before the lead time, write it down for the next shift, and know what every request costs.
Those habits came from the network, and they are still how I engineer systems that are reliable, secure and maintainable.