Section650

Writing

TopicLearning
Reading13 min

From Expect Scripts to AI Agents: How DevOps Changed During My Career

I started out writing Expect scripts that pretended to be a human operator. Now AI agents can reason about the work an operator does. The tools changed almost completely. The engineering problem underneath them barely moved.

By Ka Lun Chan · Learning / AI / Architecture

Where I started: a Network Operations Center

Early in my career, I worked in a Network Operations Center (NOC) for a national broadband network. We kept a carrier network running across more than 2,000 collocations, from DS1 circuits up to OC links, with routers, switches, Unix systems and fault management tools.

Deployments looked very different from what engineers see today. There was no Terraform, no Kubernetes, no GitHub Actions and none of the cloud abstractions most teams now take for granted. We automated deployments and routine operational work with Expect scripts and KornShell (ksh) scripts.

Automation only went so far. Some changes needed a person standing in front of the equipment. We worked with technicians in the field who drove to a site, plugged into a console port, and made the change or ran the troubleshooting with us on the phone.

Looking back, much of what I believe about DevOps started there. The tools have changed almost completely since then, and the engineering problem underneath them much less.

Expect scripts automated a person at a terminal

A lot of the equipment we managed wasn’t designed to be automated. It had no API. Often the only interface was an interactive command line built for a person typing at a keyboard. So we automated the person.

That’s what Expect is for. A script played the part of an operator, one prompt at a time:

  1. Connect to the device.
  2. Wait for a prompt.
  3. Send a command.
  4. Check the response.
  5. Handle the next prompt, or the error you didn’t want.
  6. Repeat.

A simplified version looked something like this:

Expect (simplified)
spawn telnet $host
expect "Username:"
send "$user\r"
expect "Password:"
send "$pass\r"
expect {
  "#"     { send "show interfaces status\r" }
  timeout { puts "no prompt from $host"; exit 1 }
}
expect "#"
send "exit\r"

The happy path was the easy part. The work was in everything else: a prompt that never came, a device that answered with a slightly different banner after a firmware upgrade, a session that dropped halfway through a change. Every one of those had to be handled, or the script would carry on typing into a terminal that was no longer in the state it assumed.

What the NOC taught me about operations

Those years taught me most of the operational habits I still rely on. None of them were called DevOps at the time.

Repeatability
If a change was worth doing twice, it was worth scripting, so the hundredth site got the same change as the first.
Failure handling
A script that only handles success will eventually make an outage worse.
Deployment safety
Change one site, check it, then widen the rollout. A bad change pushed everywhere at once is very hard to undo.
Rollback
Before making a change, know how to put things back, and save the configuration you’re about to replace.
Logging
Capture every session. At three in the morning, the log is how you find out what the script really did.
Knowing the state
Automation acts on what it believes about a system. When that belief is wrong, it does the wrong thing confidently.
Dependencies
A device rarely fails alone. You have to know what sits upstream and downstream before you touch it.
People and equipment
Some steps were software and some were a technician in a cabinet. Coordinating the two was part of the deployment.

Troubleshooting under pressure belongs on that list too. When a circuit went down, what mattered was how quickly we could find the fault and how sure we were that the fix wouldn’t break something else.

Then DevOps arrived

Shell scripts never went away. What changed is that, piece by piece, the industry agreed on standard tools for the jobs those scripts used to do:

  • Jenkins gave teams standardized CI/CD pipelines.
  • Docker gave us a consistent way to package an application and everything it needs.
  • Terraform turned infrastructure configuration into version-controlled code.
  • Kubernetes standardized a huge amount of application orchestration.
  • GitHub Actions moved CI/CD into the same repository as the code.
  • Cloud platforms put APIs in front of infrastructure that once took a truck roll and a console cable.

Infrastructure that used to be configured by hand, device by device, could now be described in a file and applied with one command:

Terminal
terraform apply

Deployment scripts that used to live on somebody’s server became a pipeline that starts when code is pushed:

Pipeline
git push
  → CI
  → test
  → build
  → container image
  → deploy

And a process that crashed at night no longer needed someone to log in and restart it. Kubernetes notices the failed container and starts a new one.

I’ve used a lot of this in my own work. On a public-sector modernization, new services ran as containers on AWS, provisioned with Terraform and shipped through a pipeline with accessibility and security checks built in. When I compare that with a NOC change window, the difference is enormous.

Sometimes it still feels like writing deployment scripts

Having worked in both eras, I’ll admit that some days modern DevOps feels like I’m still writing deployment scripts. We finally agreed on better ways to write them.

  • YAML replaced some of the shell scripts.
  • Terraform replaced some of the infrastructure scripts.
  • Dockerfiles standardized how applications are packaged.
  • Kubernetes manifests describe how applications should run.
  • CI/CD pipelines describe how code moves from a commit to production.

The syntax changed, the abstractions got much better, and infrastructure became programmable. The idea underneath is one I recognize from the NOC.

Describe what you want the system to do, automate it, watch what happens, and handle the failures.

What standardization fixed

That doesn’t make modern DevOps “just scripts.” Standardization was a major engineering improvement because it changed where operational knowledge lives.

In the NOC, knowing how to deploy something could live in a shell script on one server, in a runbook that was mostly right, or in the head of the engineer who had done it last. When that engineer was on vacation, the knowledge was too.

Modern DevOps moved much of that knowledge into declarative, version-controlled files. That brought a long list of things we used to work hard for:

  • Environments you can rebuild from scratch, the same way every time
  • Code review and version history for infrastructure changes
  • Automated tests before anything reaches production
  • Standard environments from a laptop to production
  • Easier rollback, because the previous version is a commit away
  • Better observability
  • Much less dependence on tribal knowledge

DevOps changed from hand-written scripts that automated people at terminals to standardized, version-controlled tools: CI/CD pipelines, containers, infrastructure as code, Kubernetes and cloud APIs. Over that time, the industry turned operational knowledge into structured, repeatable instructions that machines can read and people can review.

Removing single points of failure in knowledge is a big part of how I lead engineering teams, and infrastructure as code is one of the best tools we’ve had for it. It also turned out to be what AI needed: operations described in files that a model can read.

Can AI automate DevOps?

Increasingly, yes. AI can already write Terraform, Kubernetes manifests and CI/CD pipelines, read logs and stack traces, investigate failed deployments and propose fixes. It works best taking on investigation and repetitive work, with an engineer still accountable for what reaches production.

The list of things AI can help with is already long:

  • Generating Terraform, Kubernetes manifests and pipeline definitions
  • Reading logs, interpreting stack traces and correlating errors across systems
  • Investigating failed deployments and alerts
  • Explaining unfamiliar infrastructure
  • Spotting suspicious configuration
  • Writing monitoring queries and runbooks
  • Summarizing incidents and proposing remediation

An engineer can now say something like this to an agent:

“Our Django service started returning 500s after the last deployment. Find out what changed.”

To answer it, the agent can work through the same trail I would: Git history, deployment history, application logs, infrastructure configuration, metrics, traces and recent changes. Then it comes back with a likely root cause. I’ve been building with LLMs, agents and the Model Context Protocol, and that kind of investigation is where I see them help most right now.

From a 3 a.m. page to a pull request

Today, an incident usually goes like this: an alert fires, a person wakes up, opens a dashboard, reads logs, connects with SSH or kubectl, works out what’s wrong, changes the configuration and deploys the fix.

I expect more incidents to start going differently:

Fig. 663-1AI-assisted incident response, with a person at the gate Human approval before production
  • Alert to AI agent
  • AI agent to Root-cause hypothesis
  • Root-cause hypothesis to Pull request
  • Pull request to Tests & policy checks
  • Tests & policy checks to Human approval
  • Human approval to Deploy
  • Deploy to Verify

The agent does the first hour of investigation before anyone is awake. It brings back a hypothesis, the evidence for it, and a proposed change as a pull request. The change runs through the same tests and policy checks as any other. A person reads it, approves it, and watches the deploy verify.

That shape matters to me more than how clever the agent is. Every step that touches production still goes through a pipeline, a review and a record.

Don’t hand AI the production keys

AI-assisted DevOps and AI changing production on its own are very different things.

An AI system that misunderstands an application is inconvenient. One that misunderstands IAM, networking, a production database, Terraform state, a Kubernetes cluster or a destructive command can cause a serious outage or a security incident. In the NOC, a bad script could take down a site. An agent with broad credentials can do the same thing at cloud speed.

So the guardrails I’d put in place look a lot like the ones we already use for people:

Least privilege
Give the agent only the access the task needs, scoped to one environment.
Read-only by default
Investigation needs to read logs and metrics. It doesn’t need write access.
Pull requests, not direct changes
Infrastructure changes go through the repository, where they get reviewed, tested and recorded.
Human approval for production
A named person approves every production change and owns the result.
Tests and policy checks
The same automated gates apply whether a person or an agent wrote the change.
Audit trails
Every action an agent takes is logged, so you can reconstruct what it did and why.
Rollback
Every change has a tested way back.
Blast-radius limits
Separate environments, staged rollouts and limits on how much one change can touch.
Secrets stay out
Production secrets never go into a prompt or a context window unless there’s no other way.

These are the controls we already built for people, applied to a new kind of operator. When I designed an LLM document pipeline with a human in the loop, the confidence threshold for sending work to a person was a product decision. Production access deserves at least that much care.

For now, let AI take the investigation and the repetitive work, and keep the accountability with engineers.

The full circle: from Expect to AI agents

Two decades ago, I wrote Expect scripts because the machines in front of me weren’t designed to be automated. We taught software to act like a human operator: wait for this prompt, type this, look for that response, then type the next command.

Then infrastructure got APIs. Then it became code. Then it became declarative, so you described the state you wanted and the tooling worked out how to get there. Each step made the machine’s interface more structured, and each one moved automation further from imitating a person at a keyboard.

AI agents operate one level above that. They read the structured interfaces we spent years building and reason about the outcome someone asked for.

I started out writing software that pretended to be a human operator. Now we’re building AI systems that can reason about the work an operator does.

Laid out end to end, the interfaces I’ve worked through look like this:

  1. Console
  2. Shell scripts
  3. Expect and ksh
  4. Configuration management
  5. CI/CD
  6. Containers
  7. Infrastructure as code
  8. Kubernetes
  9. Cloud APIs
  10. AI-assisted operations
  11. Agentic operations, maybe

The objective at the bottom of that list is close to the one at the top: make systems reliable, repeatable, observable and recoverable, with less manual work each time.

Tools change. The system underneath them doesn’t change much.

Watching these tools come and go has made me careful about getting attached to any of them. Jenkins, Terraform, Kubernetes, GitHub Actions and whichever AI operations platform comes next are all ways of expressing intent to a system.

What lasts is understanding the system underneath the tool:

  • Networking, DNS and certificates
  • Linux, processes and databases
  • Security and access control
  • Deployment strategies and rollback
  • Distributed systems and their failure modes
  • Observability
  • Risk

AI is unlikely to replace DevOps engineers outright. It will write more of the Terraform, investigate more of the logs and propose more of the routine changes. Someone still has to judge whether a proposal makes sense, and know what happens when it doesn’t.

It’s the same argument I made in you don’t need to remember everything to be a good software engineer. I don’t need to remember Expect syntax anymore, and soon I may not need to remember Terraform syntax either. I still need to know what a routing change does to the traffic behind it. That came from the NOC, and I still use it.

Architecture has followed a similar path, which I wrote about in what different systems taught me about software architecture. Both stories end up in the same place: every abstraction is useful until it fails, and then someone has to understand what’s underneath.

Short answers

Can AI automate DevOps?

Increasingly, yes. AI can write Terraform, Kubernetes manifests and CI/CD pipelines, read logs and stack traces, investigate failed deployments and propose fixes as pull requests. It works best taking on investigation and repetitive work, with read-only access by default and a person approving every production change.

Will AI replace DevOps engineers?

Unlikely, at least not outright. AI will write more of the infrastructure code and do more of the first investigation during incidents. Someone still has to understand networking, security, databases and failure modes well enough to judge whether a proposed change is safe, and to own the result.

How has DevOps changed?

DevOps moved from hand-written scripts that automated people at terminals to standardized, version-controlled tools: CI/CD pipelines, containers, infrastructure as code, Kubernetes and cloud APIs. Operational knowledge that once lived in scripts, runbooks and people’s heads is now mostly in files that can be reviewed, tested and rolled back.

What did engineers use before modern DevOps?

Shell scripts such as ksh, Expect scripts that automated interactive command lines, runbooks, and people. Expect was common because a lot of network equipment had no API, so scripts waited for a prompt, sent a command and checked the response, the way an operator would. Some changes still needed a technician on site with a console cable.

What is the future of DevOps?

More AI-assisted operations: agents that investigate alerts across logs, metrics, traces and recent commits, then propose a fix as a pull request that goes through tests, policy checks and human approval. The guardrails, such as least privilege, audit trails, rollback and blast-radius limits, matter as much as the agents.

Still automating the same thing

Compared with sitting in a NOC writing Expect and ksh scripts, modern DevOps can feel incredibly advanced. Some days it also feels strangely familiar.

I’m still learning the AI part, the same way I learned every layer before it. But the job hasn’t moved far from where I started:

Make this system do what I want, reliably, without a person doing every step by hand.

AI may be the next abstraction in that very long journey. I’d like it to be one we can review, test and roll back.

Talk through your DevOps or AI operations plans