Writing
Scheduling AI Work: What I Learned Comparing Hermes, ChatGPT and Claude
The moment an AI task runs on a timer instead of in a conversation, every assumption about working with an assistant stops holding. A content research assistant for Useful Little Tools is my test case, and five ways to schedule it are the comparison.
By Ka Lun Chan · Learning and tools · Learning / AI
What changes when nobody is watching
I have spent the last few weeks learning self-hosted agents by putting Hermes Agent on a Hostinger VPS, reaching it from VS Code and designing a job-search agent to run on it. Somewhere in that I realized the thing I was learning was not Hermes. It was scheduling. The moment an AI task runs on a timer instead of in a conversation, almost every assumption I had about working with an assistant stops holding.
In a conversation, I supply the context, catch the mistakes, and decide what happens next, turn by turn. On a schedule, nobody does. The task has to carry its own instructions, find its own context, fail in a way someone will notice, and never take an action I would have stopped. Scheduling turns an assistant into a system, and systems need an owner for execution, context, credentials, uptime, logs and recovery.
This post compares the five ways I could schedule one concrete job, and what each one makes me responsible for. I am still learning; nothing here reports a finished automation.
The example: a content research assistant
The job I keep using as a test case is a research assistant for Useful Little Tools, my site of free calculators. Once a week it should read the articles already on the site, research the questions people ask around each tool, drop any idea that duplicates something published or already suggested, save three ranked ideas with a one-line reason each, and send me a short brief to review. That is where it stops. Drafting and publishing stay separate, done by me when I choose, because an assistant that researches is useful and an assistant that publishes unreviewed is a liability.
It is a good example because every step is checkable, it needs real data from my own site, it needs memory of past suggestions, it has an obvious failure mode (suggesting last month's idea again), and it has a clean human gate at the end.
Five ways to schedule the same job
Hermes, self-hosted. Hermes has its own cron. I write the job as a skill, attach it to a scheduled run, and the agent executes it on the VPS with whatever tools, memory and files live there, delivering the result to Telegram. The model is whatever I configured: Claude through an Anthropic API key, a Claude Code login where the provider permits it, or a bundled credit. One distinction that confused me at first: Hermes using Claude as its model is a normal API call from Hermes's own agent loop. Hermes running Claude Code as a coding tool is different, a shell command the agent can run like any other, which starts a separate agent with its own permissions and billing. The first is how the research assistant would think; the second is something I would use only for code.
ChatGPT scheduled tasks. In ChatGPT you ask for a recurring task and OpenAI runs it on its side, on a cadence, and notifies you. The help centre is specific about limits: three active tasks on Free and Go, five on Plus, ten on Business and Edu, fifteen on Pro and Enterprise; free accounts get a once-a-day window rather than an exact time; and a task created inside a project cannot read that project's files. It is the least work to set up and the least able to reach my own data.
Claude Cowork scheduled tasks. On a paid Claude plan, Cowork saves a prompt and runs it on a schedule as its own session, with your connectors, skills and the files saved to your account, and an optional folder to work in. The help centre says tasks run remotely now, so the laptop can be closed, except for tasks that need local files or apps, which run only locally. Cadences are hourly, daily, weekdays, weekly or on demand. This fits the research assistant well if the inputs are connectors and documents I can put in the account.
Claude Code routines. Claude Code's cloud scheduling is a routine: a saved prompt, one or more GitHub repositories, connectors and a trigger, run on Anthropic's infrastructure with no permission prompts, at most once an hour. Every run clones the repo fresh, so it sees no local files, and each run is a session I can open afterwards. The docs make a point I have learned the hard way elsewhere: a green status means the session exited cleanly, not that the task succeeded. Since the Useful Little Tools site lives in a repository, a routine could read the articles directly.
Claude Code from an external scheduler. The last option is the oldest one: cron on the VPS runs Claude Code in non-interactive mode with a prompt, the way I used to run any script. Everything is mine: the schedule, the environment, the logs, the retries. Claude Code also has Desktop scheduled tasks, which run locally only while the app is open and the computer is awake, and /loop, which lives inside one session and expires after seven days, so neither suits a job that should run when I am asleep.
| Hermes (self-hosted) | ChatGPT scheduled tasks | Claude Cowork scheduled tasks | Claude Code routines | Claude Code via cron | |
|---|---|---|---|---|---|
| Runs where | Your VPS, inside the agent | OpenAI’s cloud | Anthropic’s cloud (locally only if it needs local files) | Anthropic’s cloud, on a fresh clone of a repo | Wherever cron runs: VPS, CI, laptop |
| Needs your machine | Yes, the VPS | No | No | No | Yes, whatever hosts cron |
| Credentials | Yours, on the VPS | OpenAI’s, plus connectors you link | Your Claude connectors and account files | Your GitHub access and connectors | Whatever the environment holds: a login or an API key |
| Context on each run | Its memory, skills and files on the box | The saved task prompt | The saved instructions, connectors, chosen folder | The saved prompt plus the repo | The prompt you pass, plus files on disk |
| Minimum interval / caps | Cron, any interval | 3 to 15 active tasks by plan; free once a day | Hourly presets | 1 hour; 100 scheduled runs an hour | Cron, any interval |
| Billing | Model API tokens, or a login the provider allows | Included in the plan | Plan usage (billing not stated) | Plan usage; overage via usage credits or runs rejected | API tokens if a key is set, else plan usage |
| Logs and failure | Your container logs; you build recovery | Task history in the app | Each run is a session you review | Each run is a session; green is not success | Your cron log and exit codes |
| Best for | Learning, control, custom tools | Simple reminders and checks | Recurring work on your own documents and connectors | Repo-centred, unattended work | Pipelines you already own |
Who owns what
Execution. With Hermes or cron, I own it: the VPS has to be up, patched and reachable, and if it is down at 6 a.m. the job does not run. With the three managed options the vendor owns it, which is the whole convenience.
Context. A scheduled run starts cold. Hermes has the most durable context because its memory, skills and data sit on the same disk. Cowork has what is in my account. A routine has what is in the repo. ChatGPT has the saved prompt. Cron has whatever I wrote to disk. In every case, the context the job needs has to be put somewhere the job can read it, before the run.
Credentials. Self-hosting means my keys on my server, which is both the control I wanted and a thing I now have to protect. The managed options hold their own credentials and borrow mine through connectors, which is less for me to secure and more I have to trust. A routine acts as me on GitHub and in every connector I leave attached, so the docs are right to say remove the ones the job does not need.
Logs and recovery. Hermes and cron give me raw logs and nothing else; retries, alerts and the "it silently did nothing for a week" problem are mine to build. The managed options keep run histories, and Cowork and routines leave a full session to read. None of them, managed or not, knows whether the research brief was any good. Only I do, which is why the human review step is not optional.
Subscription login versus API billing
Logging in with a subscription does not guarantee that a scheduled run is covered by it. ChatGPT tasks are part of the plan, within the task caps. Cowork tasks run under your plan, and the help article does not state separate billing. Claude Code routines draw down subscription usage like interactive sessions, and when the limit is hit, accounts with usage credits turned on continue at metered overage while accounts without them have further runs rejected until the window resets. A headless Claude Code run is billed to the API whenever an API key or auth token is set in its environment, because those take precedence over a claude.ai login. And Hermes calling Claude bills per token to the Anthropic key, or, through the Claude Code login route, needs a Max plan with extra-usage credits.
The practical consequence is that a scheduled job can quietly cost more than the interactive version of the same work, in two ways. It can run more often than I would have asked by hand. And it can land on a different billing path than I assumed, most often because a key sat in an environment variable on the box running cron. Before I schedule anything I now write down which account pays, and I check the usage page after the first week.
Script or model?
Most of the research assistant is a script. Fetching the list of published articles is a script. Fetching the stored ideas is a script. Checking whether a candidate idea matches a published title or an earlier suggestion is a script, and a far more reliable one than asking a model whether two things are "basically the same". Sending a message is a script. I would write every one of those deterministically, test them, and run them before any model is involved, for the same reason I put deterministic filtering before the model in the job-search agent.
The model earns its place in two steps. Reading what people ask about a tool, across messy sources, and turning it into candidate questions is interpretation, which a script does badly. Ranking three ideas and explaining why each is worth writing is judgment I would like a draft of, as long as I make the final call. Everything else is plumbing, and plumbing that is deterministic is plumbing I can trust at 6 a.m.
Durable instructions, not a remembered chat
The mistake I made first was designing the assistant in a VS Code conversation, getting it to work once, and assuming the scheduled version would know what we had agreed. It would not. A scheduled run has no memory of my chat. Whatever I want it to know has to be durable: the instructions in a skill file or a routine prompt that is self-contained and says what success looks like; the data, the published articles and the suggestion history, in a file or a database the job can open; the credentials in the one place the job reads them; and the rules, such as "never publish, only brief", written into the instructions rather than remembered from a session.
That is why Hermes's skills, Cowork's saved instructions, routine prompts and a cron job's script all converge on the same shape. The scheduled process is a small program with a model inside it, and programs need their inputs on disk. The chat was the design session. The artifact of the design session has to be files.
Where I would start
If the job needs only the vendor's own tools and a reminder, ChatGPT scheduled tasks or Cowork scheduled tasks are the least work and the least to maintain, and the caps tell you when you have outgrown them. If the job lives in a repository, a Claude Code routine gives you unattended runs with a reviewable session each time, at the cost of a GitHub-shaped workflow and no local files. If you want to learn how these systems work, or you need custom tools, your own data on your own disk and control of every step, self-host with Hermes or plain cron, and accept that uptime, logs, retries and security are now yours.
Whichever you pick: write the instructions as if for a stranger, put the data where the run can read it, decide which account pays, keep every irreversible action behind a human approval, and read the first few runs in full. The convenience of managed scheduling is real. So is the control of running your own, and the maintenance that comes with it. For the research assistant I am starting on Hermes, because the point of this experiment is to learn what the managed versions are hiding from me, and I will report what it costs.
Short answers
What changes when an AI task runs on a schedule instead of in a chat?
Nobody is there to supply context, catch mistakes or decide the next step. The task has to carry self-contained instructions, read its data from files or a database, hold its credentials somewhere it can reach, fail in a way someone notices, and never take an action a person would have stopped. Scheduling turns an assistant into a system with an owner for execution, context, credentials, uptime, logs and recovery.
Does a ChatGPT or Claude subscription cover scheduled runs?
ChatGPT scheduled tasks are part of the plan within active-task caps (3 on Free and Go, 5 on Plus, 10 on Business and Edu, 15 on Pro and Enterprise). Claude Cowork tasks run under a paid plan. Claude Code routines draw down subscription usage, and once the limit is hit, runs continue on metered overage only if usage credits are on; otherwise they are rejected until the window resets. A headless Claude Code run bills the API whenever an API key or auth token is in its environment.
When is a script enough and when does the model add value?
Fetching articles, loading stored ideas, checking for duplicates and sending a message are deterministic and should be scripts you can test. The model earns its place where interpretation or judgment is needed: turning messy reader questions into candidate ideas, and ranking them with reasons for a person to review.
What is the difference between Hermes using Claude as a model and Hermes running Claude Code?
Using Claude as the model is an ordinary API call from Hermes’s own agent loop, billed to the configured key or login. Running Claude Code is a shell command the agent can execute like any other, which starts a separate coding agent with its own permissions and billing. The first is how a research assistant thinks; the second is for code.
A schedule turns an assistant into a system
The difference between asking an AI once and running it every week is not the model. It is everything around the model that now needs an owner.
Durable instructions, data on disk, a named account that pays, deterministic plumbing, and a human gate before anything irreversible.
Next in the series: the research assistant's first scheduled runs, what it suggested, and what the first week cost.