Writing
AI Agents Are Easy to Demonstrate. Making Them Reliable in Production Is Hard.
Getting an agent to work in an afternoon is delightful. Getting it to run every morning without me, safely and within budget, is the same engineering I have done for twenty years, with a model in the middle.
By Ka Lun Chan · Still Building After 20 Years · AI / Architecture / Leadership
The afternoon it worked
This is the third story in Still Building After 20 Years, and the most current one. For the past while I have been experimenting with AI agents: Claude Code and OpenAI’s Codex in the terminal, ChatGPT in the browser, Hermes Agent on a VPS I administer, OmniRoute as a gateway across several model providers, and scheduled jobs that run without me. I am building a job-search agent on that setup and planning a content research assistant for Useful Little Tools.
Getting an agent to work the first time is delightful. You describe the task, connect a tool or two, run it, and an afternoon later something that would have taken a week of glue code is answering questions and calling APIs on your behalf. The first time my agent came back with a sensible answer I felt the way I felt the first time a script I wrote configured a router while I watched. Then I asked the question I have asked about every system I have run: what happens tomorrow, at six in the morning, when I am not watching?
An AI agent demo shows that a model can do the task once with a person watching. A production agent has to do it every day without one: survive failed calls and missed schedules, never do the same irreversible thing twice, stay inside a budget, touch only what it is allowed to, and produce results someone can trust and check. The model is the easy part. Everything around it is distributed systems engineering.
What production asks that a demo never does
A demo has an operator. I typed the prompt, I saw the answer, I would have caught a wrong one, and I would never have let it email anyone. Take me out and every one of those duties has to move into the system. Instructions have to be self-contained because there is no conversation to remember. Data has to be on disk because there is no me to paste it in. Errors have to be reported because nobody is reading the screen. And the decision about whether the output is good enough has to be designed, not felt.
None of my agents is in production yet, and I am saying so plainly because the gap between “it worked” and “it runs” is exactly what this post is about.
Reliability is distributed systems again
When I built a VoIP platform, a single call crossed signaling, media, business logic and billing, and each of those failed independently. I wrote about the lessons in what open-source VoIP taught me about distributed systems. An agent is the same shape with a model in the middle.
The model API will time out, rate-limit you and occasionally return something malformed. Every call needs a timeout and a bounded retry with exponential backoff, and the retries need a ceiling, because an agent that retries forever is an agent that spends forever. Tools will fail too, and the agent needs to know which failures are worth retrying and which mean stop.
Idempotency matters more than people expect. If a run dies halfway and the scheduler starts it again, will it send the report twice, apply the label twice, or, worse, submit the application twice? Every action with a side effect needs a key so a repeat changes nothing, the same discipline that kept a replayed billing event from charging a caller twice.
Schedulers miss runs. A VPS reboots, a cloud schedule starts late, a desktop task skips because the laptop was asleep. A missed run has to be detected, by something other than me noticing on Friday that Monday’s report never came. State has to live somewhere durable, so the next run knows what the last one did. And observability is not optional: what ran, how long it took, what it called, what it cost, and whether the result passed whatever check I defined. A green status on a run means the process exited. It does not mean the task succeeded, and I have learned to read the transcript before I believe the light.
Cost is an architecture decision
Every model call has a price, and an unattended agent makes a lot of them. Token consumption is set by design choices: how much context you send, how many steps the agent takes, how many items reach the model at all. That is why I put deterministic filtering in front of the model in my job-search design, so the expensive step sees only what could possibly matter.
Different steps deserve different models. Extracting fields from a short posting is cheap work a fast, inexpensive model does well. Judging fit and drafting text is where a stronger model pays for itself. A gateway like OmniRoute makes switching providers and falling back easier, and adds a container to run and a key store to secure. Free tiers are useful for experiments and unreliable for anything you depend on. Local models trade a bill for hardware and maintenance. And sometimes the right amount of AI is zero. FedPath scores opportunities with transparent rules and keeps solicitation documents on its own server rather than sending them to a model, because users need to see why a contract scored the way it did. That was a deliberate architectural decision, and I would make it again.
Scope what the agent can touch
An agent with tools is a program that acts with your credentials, takes instructions from text, and reads text from the open internet. That last part should worry anyone who has run production systems, because untrusted input can carry instructions. The defences are the old ones, applied with more care. Least-privilege credentials created for the agent alone, so revoking one breaks nothing else. Secrets in a secrets store, never in a prompt or a skill file. A tool allow-list instead of a shell with root. A human approval in front of anything irreversible: sending, publishing, paying, deleting, applying. And an audit trail of what the agent did, so that “why did it do that?” has an answer.
I learned deny-by-default on carrier routers long before anyone said “agent”. It is the same rule, and it matters more when the thing you are configuring can decide for itself what to try next.
How I would decide it is ready
As an engineering leader, the question I care about is whether I would let the agent run unattended with my name on the result. Impressive is a lower bar. These are the questions I would ask a team before saying yes.
| Question | What a good answer looks like |
|---|---|
| Who owns it? | A named person who gets the alert, not “the AI” |
| What does done look like? | A measurable result someone checks, beyond the run exiting cleanly |
| What happens when the model or a tool fails? | Timeouts, bounded retries with backoff, and a failed run that says so |
| What if it runs twice? | Idempotent writes, keyed so a repeat changes nothing |
| What if it does not run at all? | A missed run is detected, not discovered a week later |
| What can it touch? | Least-privilege credentials and tools; irreversible actions behind approval |
| What does a run cost? | Tokens and dollars logged per run, with a budget that stops it |
| Would a script do? | Deterministic code for everything that does not need judgment |
Behind those sit the usual production commitments: a reliability target that someone agreed to, monitoring that pages a person, an incident process that includes “turn the agent off”, and a cost budget reviewed like any other line item. If a team cannot answer them, the agent is still a demo, however good the demo is.
Short answers
What is the difference between an AI agent demo and a production agent?
A demo shows a model can do the task once with a person watching. A production agent does it every day without one: it survives failed calls and missed schedules, never repeats an irreversible action, stays within a budget, touches only what it is allowed to, and produces results someone can check.
How do you make an AI agent reliable?
Treat it as a distributed system: timeouts and bounded retries with exponential backoff on every model and tool call, idempotent side effects keyed so a repeat changes nothing, detection of missed runs, durable state between runs, and observability of what ran, what it cost and whether the result passed a check.
When should you not use AI in an agent?
For anything that does not need judgment. Fetching, filtering, de-duplicating and scoring by transparent rules are better as deterministic code you can test. FedPath, for example, scores contracts with visible rules and keeps solicitation documents off outside AI services.
New capabilities, old problems, higher stakes
AI agents bring capabilities I could not have built a few years ago. They do not remove a single traditional software engineering problem.
Retries, idempotency, state, observability, cost and least privilege matter more when the system can decide what to do next.
That is the end of the series, for now. The start was a router console, and the habits I learned there are still the ones keeping the newest systems honest.