Section650

Writing

TopicArchitecture
Reading11 min

I Was Working on Network Security Long Before DevSecOps Was a Thing

How working in carrier networks, helping launch a managed security service, and training a technical support team shaped the way I approach software and infrastructure security today.

By Ka Lun Chan · Software architecture · Architecture / Leadership / Learning

Before cloud security, there were routers and ACLs

The words today are Zero Trust, DevSecOps, identity and access management, infrastructure as code, automated scanning. My first security work had none of those words in it. It was an access control list on a Cisco router in a carrier network, and the question it answered was older than any of them: which traffic is allowed to reach which thing.

An access control list (ACL) is an ordered set of permit and deny rules that a router checks against each packet: source address, destination address, protocol and port. It is applied to an interface in one direction, inbound or outbound, the first matching rule wins, and anything that matches nothing is dropped. In a carrier network, ACLs protected the network itself: who could reach a router’s management interface, which neighbours were trusted, what a customer’s traffic was allowed to do once it entered our side.

What I learned quickly is that network security was never about blocking traffic. The network existed to carry traffic. Every rule had two ways to fail. It could let through something it should have stopped, which you might not find out about for a long time. Or it could block something legitimate, which you found out about in minutes, because a customer’s application stopped working and the ticket came to us. A rule that drops a customer’s real traffic is an outage with a security label on it.

So the actual skill was understanding which systems needed to talk to each other, writing rules that allowed exactly that, and knowing where to look when something that should connect didn’t. An ACL was a frequent suspect in connectivity troubleshooting, and reading one carefully, in order, with the implicit deny at the bottom in mind, became a habit I never lost.

When the carrier started selling security

Then the carrier introduced a managed security service built on NetScreen firewall and VPN products, and security stopped being something we did to protect our own network and became something we sold.

NetScreen Technologies was founded in 1997 and built firewalls and VPN gateways as purpose-built hardware, with security processing in custom silicon rather than on a general-purpose server. The devices ran ScreenOS, which organized a network into zones and expressed security as policies between them. They were fast, they were reliable, and they were a good fit for a carrier that wanted to put a managed box at the customer’s edge. Juniper Networks bought the company in 2004 for about four billion dollars, and ScreenOS lived on in Juniper’s product line for years afterward.

A managed security service changed what the carrier was responsible for. Before, a customer bought connectivity and we made sure the circuit was up. Now the customer bought a firewall we configured, a VPN between their offices that ran across our network, a security policy we maintained on their behalf, and monitoring to tell us when any of it broke. Every change request was a security decision made for someone else. Every outage on that box was both a security event and a connectivity outage. The operational side of that, the process for taking a change, verifying it, applying it and recording it, had to be built alongside the service.

One of the first three people in the Security Operations Center

I was one of the first three people supporting that service in the Security Operations Center. I want to be precise about what that was. It was not a title and I was not running it. It was a chair in a new room, next to two other people, with a service that had customers and not yet a runbook.

Being that early meant learning several things at once. The product, from the documentation and from devices in front of us. The customers’ environments, because a firewall policy only makes sense when you know what sits behind it and what it needs to reach. And the process, because none existed yet: how a change request came in, how we confirmed it was really the customer asking, how we applied it without taking them offline, where we wrote down what we did, and what we escalated to the vendor.

In a team of three, everything ends up with you at some point. That is exhausting and it is also the fastest way I know to learn a system end to end. The habit I took from it was writing things down as we learned them, so the fourth person would not have to learn them the way we did. Looking back, that was the first operational documentation I ever cared about, and it was because I could see exactly who would need it.

Firewalls, VPNs and the whole path

The technology is worth a short explanation, because the ideas have not changed much. A stateful firewall remembers the connections it has allowed, so it can let the reply traffic back in without a separate rule for it. Interfaces belong to zones, with names like Trust, Untrust and DMZ, and a policy says which traffic may cross from one zone to another: from where, to where, which service, and whether to permit, deny and log it. Network address translation lets a private network sit behind one public address, or maps a public address to one inside machine. An IPsec VPN builds an encrypted tunnel between two sites, and both ends have to agree on every parameter before it comes up. The firewall is also a router, so it has routes, and the routes can be wrong. And it keeps logs and a session table, which is where most troubleshooting starts.

The lesson underneath all of it was that a security device sits on a path, and you cannot troubleshoot it in isolation. A customer calls to say an application cannot connect. The cause might be routing on either side, an ACL upstream, the firewall policy, the NAT rule, a VPN tunnel that is up in one phase and not the other, DNS, or the application’s own configuration, which nobody mentioned because it was changed yesterday by someone else. The firewall was the suspect in every ticket. It was the cause less often than people assumed. Finding out which meant tracing the path from the client to the server and ruling things out in order.

That is still how I debug a production problem that touches several systems. It is the same instinct behind the way I followed a VoIP call through signaling, media and billing years later.

Flying east to train the TAC

As the service grew, I travelled to the East Coast to train the Technical Assistance Center, the team that would take customer calls and troubleshoot the service after us. It is one of the experiences I remember most clearly from those years, and the reason has little to do with firewalls.

Up to that point I had been learning the technology. Now I had to explain it to engineers who would be on the phone with customers without me in the room. That is a different skill. The questions they asked found every gap in my understanding and every gap in our documentation. Things I did by feel had to become steps someone else could follow at three in the morning: what to check first, how to read a log, what a healthy tunnel looks like, when to stop and escalate. I had to turn how I troubleshoot into something repeatable, and watching other engineers get confident with it was more satisfying than fixing the problem myself.

I did not have the words for it then, but that trip was the start of how I think about mentoring engineers: understanding a system is one skill, and being able to teach someone else to troubleshoot it is another. A support organization that depends on the two or three people who were there at the beginning is a single point of failure, and I have spent most of my career since trying not to be one.

Security is more than a firewall

The security work did not end when I left carrier networks. It changed shape with every job. On the voice platforms it was AAA: every call was authenticated, authorized and accounted for, so the platform always knew who was calling, what they were allowed to call, and what it cost. In application work it became authentication, single sign-on, OAuth, tokens, roles and the check on the individual record. On mobile it became the same questions asked of a client we did not control: where the credentials live on the device, what travels on the wire, and what the server re-checks. In cloud work it became IAM policies, network configuration and secrets, written as code. In government work it became federal security requirements, access control and auditability as part of the scope. And in leadership it became deciding who owns security on a team and when in the delivery cycle it gets attention.

Fig. 679-1 My security engineering journey, by stageStages, not a schedule. Undated on purpose.
  1. Carrier network operations

    Cisco router access control lists: which traffic may reach which system, and which rule just took a customer offline.

  2. A managed security service

    The carrier starts selling managed firewalls and VPNs on NetScreen. Security becomes a product with customers.

  3. Security Operations Center

    One of the first three people supporting the new service. Learning the product, the customers and the process at the same time.

  4. Training the TAC

    A trip to the East Coast to teach the Technical Assistance Center how to support and troubleshoot it.

  5. Voice platforms and AAA

    Authentication, authorization and accounting for every call: who is calling, what they may call, and what it cost.

  6. Application security

    Authentication, SSO, OAuth, JWT, role-based access and object-level authorization in Django, Rails and Node APIs.

  7. Mobile application security

    Nokia, iPhone and Android apps as untrusted clients: credentials on the device, tokens on the wire, and every permission checked again on the server.

  8. Cloud infrastructure security

    IAM, network controls, secrets and Terraform on AWS. The ACL came back as a security group and a policy document.

  9. Government systems

    Federal security requirements: access control, auditability and secure deployment as part of the scope, with FedRAMP-aligned environments.

  10. Engineering leadership

    Security as a release gate and a named owner, pulled to the front of delivery instead of the week before launch.

That is not a strict chronology, and I have gone back and forth between layers more than once. The point is that security was never a separate specialty I visited. It was part of every layer I worked in, and each layer taught me something the next one needed.

What carried into software, cloud and government work

An ACL is an authorization policy whose subjects are IP addresses. Role-based access control is an ACL whose subjects are roles. Object-level authorization is the ACL on a single record. The habits transfer directly: deny by default, allow explicitly, mind the order of the rules, log what you deny, and test every rule with the traffic it is supposed to permit, not only the traffic it is supposed to block.

When I later built APIs in Django REST Framework, Rails and Node.js, the authorization bugs I cared about most were the ones a firewall engineer would recognize. A valid token proves who is calling. It says nothing about whether they may touch this record. I wrote about that distinction in API architecture lessons from real systems, and the reason I check it first is that I spent years watching what happens when a rule permits more than its author intended.

Mobile made that lesson sharper. From the VoIP apps we wrote for Nokia phones through the iOS and Android apps I have shipped since, the app in someone’s pocket is a client we do not control. Anything it holds can be read, anything it sends can be replayed, and anything it is trusted to decide will eventually be decided wrong. So the credentials live in the platform’s secure storage, the tokens are short-lived, the secrets stay out of the bundle, and every authorization decision is made again on the server, where the app cannot argue with it.

Telecom taught me that security and reliability are the same discipline wearing different badges. On a VoIP platform, registration, routing and billing all have to agree on who is calling and what they may do, and a security control that drops legitimate calls is an outage. The platform I co-founded needed routing, billing and access control to be right at the same time, and a failure in any of them looked, to the customer, like the same failure.

Cloud infrastructure felt familiar faster than I expected. A security group is a stateful ACL. An IAM policy is an ACL over API calls. A VPC with public and private subnets is a zone design. Terraform is the configuration I used to type into a router, now in version control with a reviewer. The differences are real too: the perimeter is mostly gone, identity carries the weight the network edge used to carry, everything is an API, and change moves through a pipeline instead of a maintenance window. The questions underneath are the ones from the ACL days. Who needs to reach what, and what happens to legitimate traffic when the rule is wrong.

Government work added a discipline I had not had before. Systems built for federal programs come with security requirements that are written down, including FedRAMP-aligned environments, and the requirements have to be visible in how the system is designed, deployed and operated: access control, auditability, secure configuration and a deployment path that can be shown to someone. My role there was engineering, designing and building systems to meet those requirements. I have not obtained an authorization myself or worked as a security assessor, and I would not want this story to suggest otherwise.

Security at every layer of a company

Good security is not the property of a security team. It is the sum of decisions made at every layer, by the people who work there, most of whom do not have security in their title. The figure below walks the stack from the network up to the organization. For each layer it lists what typically goes wrong and the controls that usually answer it, which is general practice, and where I have worked in that layer myself, which is first-hand. I have kept the two apart on purpose.

Fig. 679-2 Security across the technology stackSelect a layer. First-hand experience is marked; the rest is general practice.

Network

What goes wrong here

Traffic reaching systems it has no business reaching, flat networks where one compromised host sees everything, and management interfaces exposed to the world.

Typical controls General practice

  • Access control lists and firewall policies
  • Segmentation into zones
  • VPNs for traffic that crosses untrusted networks
  • Logging what was denied

Where I’ve worked in it First-hand

Cisco router ACLs in a carrier network, then NetScreen firewall and VPN services in the Security Operations Center.

The carrier network role

Two things stand out when you read it as a whole. The technical controls get most of the attention and the organizational ones decide whether the technical ones exist. A team with no named owner for security does not have a firewall problem. It has a priority problem, and the firewall configuration is where it shows up.

What I learned

Security starts with understanding how systems work. I could not write a correct ACL until I understood which systems needed to talk, and I cannot write a correct authorization check today without understanding who the users are and what they are trying to do. Most security failures I have seen were failures of understanding before they were failures of tooling.

Security and availability have to work together. A control that breaks legitimate use will be removed, quietly, by someone under pressure, and then you have neither. The rule that survives is the one that was tested with real traffic.

Operational knowledge matters as much as the tool. A firewall is only as good as the people who configure it, read its logs and troubleshoot it at the wrong hour. The same is true of IAM, and of every scanner in a pipeline.

Documentation and training are engineering work. The trip to train the TAC taught me that a system nobody else can support is not finished, and I have treated runbooks and hand-over as part of done ever since.

Security belongs in the engineering decision, not after it. The cheapest place to find a security problem is the design review. The most expensive is the week before launch, which is where it ends up when nobody owns it earlier.

The tools changed completely. The fundamentals did not. Router ACLs became security groups and IAM policies. Physical firewalls became policies as code. Zones became VPCs and tenants. Reading a rule in order, with the default deny in mind, and asking who it will accidentally block, is the same job it was on a Cisco router in a carrier network.

Short answers

Where did Ka Lun Chan’s security experience start?

In carrier network operations, writing access control lists on Cisco routers. When the carrier launched a managed firewall and VPN service on NetScreen products, he was one of the first three people supporting it in the Security Operations Center, and he later travelled to the East Coast to train the Technical Assistance Center to support it.

How does network security experience apply to application and cloud security?

The same habits carry over: deny by default, allow explicitly, mind the order of rules, log what you deny, and test each rule with the traffic it should permit. An ACL is an authorization policy over IP addresses; role-based access control is one over roles; object-level authorization is the rule on a single record; a security group is a stateful ACL and an IAM policy is an ACL over API calls.

What was NetScreen?

NetScreen Technologies, founded in 1997, built firewall and VPN appliances with security processing in custom hardware, running ScreenOS with zone-based policies and IPsec VPNs. Juniper Networks acquired it in 2004 for about four billion dollars, and ScreenOS continued in Juniper products for years.

What does security at every layer of a company mean?

That security is the sum of decisions at the network, infrastructure, application, mobile, data, cloud, software delivery, operations and leadership layers, made mostly by people without security in their title. Technical controls get the attention; organizational ones, such as a named owner and vulnerabilities fixed as high-priority work, decide whether the technical controls exist.

I didn’t discover security when it got a name

My security experience started with Cisco ACLs in a carrier network and a chair in a brand-new Security Operations Center, and it never stopped.

Security is part of every layer I have worked in, from the router to the application, the cloud account and the roadmap conversation.

Those early years are why I treat security as an engineering decision, test every rule with the traffic it is supposed to allow, and make sure the system can be supported by people who weren’t there at the beginning.

Security engineering, role by role