Section650

Writing

TopicLearning
Reading10 min

Scheduling AI Work: What I Learned Comparing Hermes, ChatGPT and Claude

The moment an AI task runs on a timer instead of in a conversation, every assumption about working with an assistant stops holding. A content research assistant for Useful Little Tools is my test case, and five ways to schedule it are the comparison.

By Ka Lun Chan · Learning and tools · Learning / AI

What changes when nobody is watching

I have spent the last few weeks learning self-hosted agents by putting Hermes Agent on a Hostinger VPS, reaching it from VS Code and designing a job-search agent to run on it. Somewhere in that I realized the thing I was learning was not Hermes. It was scheduling. The moment an AI task runs on a timer instead of in a conversation, almost every assumption I had about working with an assistant stops holding.

In a conversation, I supply the context, catch the mistakes, and decide what happens next, turn by turn. On a schedule, nobody does. The task has to carry its own instructions, find its own context, fail in a way someone will notice, and never take an action I would have stopped. Scheduling turns an assistant into a system, and systems need an owner for execution, context, credentials, uptime, logs and recovery.

This post compares the five ways I could schedule one concrete job, and what each one makes me responsible for. I am still learning; nothing here reports a finished automation.

The example: a content research assistant

The job I keep using as a test case is a research assistant for Useful Little Tools, my site of free calculators. Once a week it should read the articles already on the site, research the questions people ask around each tool, drop any idea that duplicates something published or already suggested, save three ranked ideas with a one-line reason each, and send me a short brief to review. That is where it stops. Drafting and publishing stay separate, done by me when I choose, because an assistant that researches is useful and an assistant that publishes unreviewed is a liability.

It is a good example because every step is checkable, it needs real data from my own site, it needs memory of past suggestions, it has an obvious failure mode (suggesting last month's idea again), and it has a clean human gate at the end.

Five ways to schedule the same job

Hermes, self-hosted. Hermes has its own cron. I write the job as a skill, attach it to a scheduled run, and the agent executes it on the VPS with whatever tools, memory and files live there, delivering the result to Telegram. The model is whatever I configured: Claude through an Anthropic API key, a Claude Code login where the provider permits it, or a bundled credit. One distinction that confused me at first: Hermes using Claude as its model is a normal API call from Hermes's own agent loop. Hermes running Claude Code as a coding tool is different, a shell command the agent can run like any other, which starts a separate agent with its own permissions and billing. The first is how the research assistant would think; the second is something I would use only for code.

ChatGPT scheduled tasks. In ChatGPT you ask for a recurring task and OpenAI runs it on its side, on a cadence, and notifies you. The help centre is specific about limits: three active tasks on Free and Go, five on Plus, ten on Business and Edu, fifteen on Pro and Enterprise; free accounts get a once-a-day window rather than an exact time; and a task created inside a project cannot read that project's files. It is the least work to set up and the least able to reach my own data.

Claude Cowork scheduled tasks. On a paid Claude plan, Cowork saves a prompt and runs it on a schedule as its own session, with your connectors, skills and the files saved to your account, and an optional folder to work in. The help centre says tasks run remotely now, so the laptop can be closed, except for tasks that need local files or apps, which run only locally. Cadences are hourly, daily, weekdays, weekly or on demand. This fits the research assistant well if the inputs are connectors and documents I can put in the account.

Claude Code routines. Claude Code's cloud scheduling is a routine: a saved prompt, one or more GitHub repositories, connectors and a trigger, run on Anthropic's infrastructure with no permission prompts, at most once an hour. Every run clones the repo fresh, so it sees no local files, and each run is a session I can open afterwards. The docs make a point I have learned the hard way elsewhere: a green status means the session exited cleanly, not that the task succeeded. Since the Useful Little Tools site lives in a repository, a routine could read the articles directly.

Claude Code from an external scheduler. The last option is the oldest one: cron on the VPS runs Claude Code in non-interactive mode with a prompt, the way I used to run any script. Everything is mine: the schedule, the environment, the logs, the retries. Claude Code also has Desktop scheduled tasks, which run locally only while the app is open and the computer is awake, and /loop, which lives inside one session and expires after seven days, so neither suits a job that should run when I am asleep.

Fig. 693-1 Five ways to run the same AI job on a schedule, as each vendor documents it today
Hermes (self-hosted)ChatGPT scheduled tasksClaude Cowork scheduled tasksClaude Code routinesClaude Code via cron
Runs whereYour VPS, inside the agentOpenAI’s cloudAnthropic’s cloud (locally only if it needs local files)Anthropic’s cloud, on a fresh clone of a repoWherever cron runs: VPS, CI, laptop
Needs your machineYes, the VPSNoNoNoYes, whatever hosts cron
CredentialsYours, on the VPSOpenAI’s, plus connectors you linkYour Claude connectors and account filesYour GitHub access and connectorsWhatever the environment holds: a login or an API key
Context on each runIts memory, skills and files on the boxThe saved task promptThe saved instructions, connectors, chosen folderThe saved prompt plus the repoThe prompt you pass, plus files on disk
Minimum interval / capsCron, any interval3 to 15 active tasks by plan; free once a dayHourly presets1 hour; 100 scheduled runs an hourCron, any interval
BillingModel API tokens, or a login the provider allowsIncluded in the planPlan usage (billing not stated)Plan usage; overage via usage credits or runs rejectedAPI tokens if a key is set, else plan usage
Logs and failureYour container logs; you build recoveryTask history in the appEach run is a session you reviewEach run is a session; green is not successYour cron log and exit codes
Best forLearning, control, custom toolsSimple reminders and checksRecurring work on your own documents and connectorsRepo-centred, unattended workPipelines you already own

Who owns what

Execution. With Hermes or cron, I own it: the VPS has to be up, patched and reachable, and if it is down at 6 a.m. the job does not run. With the three managed options the vendor owns it, which is the whole convenience.

Context. A scheduled run starts cold. Hermes has the most durable context because its memory, skills and data sit on the same disk. Cowork has what is in my account. A routine has what is in the repo. ChatGPT has the saved prompt. Cron has whatever I wrote to disk. In every case, the context the job needs has to be put somewhere the job can read it, before the run.

Credentials. Self-hosting means my keys on my server, which is both the control I wanted and a thing I now have to protect. The managed options hold their own credentials and borrow mine through connectors, which is less for me to secure and more I have to trust. A routine acts as me on GitHub and in every connector I leave attached, so the docs are right to say remove the ones the job does not need.

Logs and recovery. Hermes and cron give me raw logs and nothing else; retries, alerts and the "it silently did nothing for a week" problem are mine to build. The managed options keep run histories, and Cowork and routines leave a full session to read. None of them, managed or not, knows whether the research brief was any good. Only I do, which is why the human review step is not optional.

Subscription login versus API billing

Logging in with a subscription does not guarantee that a scheduled run is covered by it. ChatGPT tasks are part of the plan, within the task caps. Cowork tasks run under your plan, and the help article does not state separate billing. Claude Code routines draw down subscription usage like interactive sessions, and when the limit is hit, accounts with usage credits turned on continue at metered overage while accounts without them have further runs rejected until the window resets. A headless Claude Code run is billed to the API whenever an API key or auth token is set in its environment, because those take precedence over a claude.ai login. And Hermes calling Claude bills per token to the Anthropic key, or, through the Claude Code login route, needs a Max plan with extra-usage credits.

The practical consequence is that a scheduled job can quietly cost more than the interactive version of the same work, in two ways. It can run more often than I would have asked by hand. And it can land on a different billing path than I assumed, most often because a key sat in an environment variable on the box running cron. Before I schedule anything I now write down which account pays, and I check the usage page after the first week.

Script or model?

Most of the research assistant is a script. Fetching the list of published articles is a script. Fetching the stored ideas is a script. Checking whether a candidate idea matches a published title or an earlier suggestion is a script, and a far more reliable one than asking a model whether two things are "basically the same". Sending a message is a script. I would write every one of those deterministically, test them, and run them before any model is involved, for the same reason I put deterministic filtering before the model in the job-search agent.

The model earns its place in two steps. Reading what people ask about a tool, across messy sources, and turning it into candidate questions is interpretation, which a script does badly. Ranking three ideas and explaining why each is worth writing is judgment I would like a draft of, as long as I make the final call. Everything else is plumbing, and plumbing that is deterministic is plumbing I can trust at 6 a.m.

Durable instructions, not a remembered chat

The mistake I made first was designing the assistant in a VS Code conversation, getting it to work once, and assuming the scheduled version would know what we had agreed. It would not. A scheduled run has no memory of my chat. Whatever I want it to know has to be durable: the instructions in a skill file or a routine prompt that is self-contained and says what success looks like; the data, the published articles and the suggestion history, in a file or a database the job can open; the credentials in the one place the job reads them; and the rules, such as "never publish, only brief", written into the instructions rather than remembered from a session.

That is why Hermes's skills, Cowork's saved instructions, routine prompts and a cron job's script all converge on the same shape. The scheduled process is a small program with a model inside it, and programs need their inputs on disk. The chat was the design session. The artifact of the design session has to be files.

Where I would start

If the job needs only the vendor's own tools and a reminder, ChatGPT scheduled tasks or Cowork scheduled tasks are the least work and the least to maintain, and the caps tell you when you have outgrown them. If the job lives in a repository, a Claude Code routine gives you unattended runs with a reviewable session each time, at the cost of a GitHub-shaped workflow and no local files. If you want to learn how these systems work, or you need custom tools, your own data on your own disk and control of every step, self-host with Hermes or plain cron, and accept that uptime, logs, retries and security are now yours.

Whichever you pick: write the instructions as if for a stranger, put the data where the run can read it, decide which account pays, keep every irreversible action behind a human approval, and read the first few runs in full. The convenience of managed scheduling is real. So is the control of running your own, and the maintenance that comes with it. For the research assistant I am starting on Hermes, because the point of this experiment is to learn what the managed versions are hiding from me, and I will report what it costs.

Short answers

What changes when an AI task runs on a schedule instead of in a chat?

Nobody is there to supply context, catch mistakes or decide the next step. The task has to carry self-contained instructions, read its data from files or a database, hold its credentials somewhere it can reach, fail in a way someone notices, and never take an action a person would have stopped. Scheduling turns an assistant into a system with an owner for execution, context, credentials, uptime, logs and recovery.

Does a ChatGPT or Claude subscription cover scheduled runs?

ChatGPT scheduled tasks are part of the plan within active-task caps (3 on Free and Go, 5 on Plus, 10 on Business and Edu, 15 on Pro and Enterprise). Claude Cowork tasks run under a paid plan. Claude Code routines draw down subscription usage, and once the limit is hit, runs continue on metered overage only if usage credits are on; otherwise they are rejected until the window resets. A headless Claude Code run bills the API whenever an API key or auth token is in its environment.

When is a script enough and when does the model add value?

Fetching articles, loading stored ideas, checking for duplicates and sending a message are deterministic and should be scripts you can test. The model earns its place where interpretation or judgment is needed: turning messy reader questions into candidate ideas, and ranking them with reasons for a person to review.

What is the difference between Hermes using Claude as a model and Hermes running Claude Code?

Using Claude as the model is an ordinary API call from Hermes’s own agent loop, billed to the configured key or login. Running Claude Code is a shell command the agent can execute like any other, which starts a separate coding agent with its own permissions and billing. The first is how a research assistant thinks; the second is for code.

A schedule turns an assistant into a system

The difference between asking an AI once and running it every week is not the model. It is everything around the model that now needs an owner.

Durable instructions, data on disk, a named account that pays, deterministic plumbing, and a human gate before anything irreversible.

Next in the series: the research assistant's first scheduled runs, what it suggested, and what the first week cost.

More from the learning path