Writing
Building My Own AI Job Search Agent with Hermes: Going Beyond ChatGPT and Claude
What can I build with a self-hosted AI agent that I could not easily do in a chat window? A job-search pipeline for engineering leadership roles is the experiment, Hermes Agent is the framework, and this is the plan before the results.
By Ka Lun Chan · Learning and tools · Learning / AI
Why I started experimenting with Hermes
I use ChatGPT and Claude Code most days, as tools for research, drafting and review, which I wrote about in building software before and after AI. Both are good enough that the question stopped being whether an assistant can help and became what else I could build with one. The thing I kept bumping into is that a chat is a conversation. It starts when I open it, it ends when I close it, and whatever it did in between happened because I was sitting there asking.
An agent is a different shape. It runs when I am not there, remembers what it did last time, has tools wired to my own systems, and reports back. I tried OpenClaw, an open-source agent that lives on your own machine with real shell access, and then deployed Hermes Agent, the open-source agent from Nous Research, on a Hostinger VPS. Hermes caught my interest for three reasons. It is built around skills, small instruction files the agent loads when a task needs them and can write for itself. It has a scheduler, so a task can run on a cron schedule without me. And it has a gateway that talks to Telegram and other messaging platforms, so the results can come to my phone. The project says it runs on a five-dollar VPS, and so far that is true.
The difference between a conversational assistant and a persistent agent is not intelligence. It is the same model, usually. The difference is that the agent has infrastructure around it: a place to run, a schedule, state that survives between runs, tools that reach my own systems, and a channel to report on. That infrastructure is what I wanted to learn by building, and self-hosting means I own every piece of it.
Self-hosting also appeals to the part of me that spent years automating carrier network operations with scripts. An agent on a VPS with cron, a database and a messaging channel is a very familiar system wearing a new coat. I learn best by building, so I picked a problem I actually have.
The problem: a senior job search is repetitive
I am looking for my next engineering leadership role: Director of Engineering, Head of Engineering, VP of Engineering, or Senior Engineering Manager. Searching for roles at that level by hand is tedious in ways that are specific to the level. The titles are inconsistent across companies. Compensation is published on some boards, in some states, and not elsewhere. A large share of postings with the right title want something I am not, or are the wrong size of company, or have been open for months. Tracking what I have seen, what I applied to and what I heard back is a spreadsheet that drifts out of date in a week.
I could keep asking an assistant to find jobs for me, and it would, one conversation at a time, with no memory of yesterday's results. What I want is a system that researches continuously, evaluates against my criteria, remembers what it has already shown me, and hands me a short list each morning. That is an engineering problem, and a good one to learn an agent framework on, because every part of it is checkable.
Designing the agent
The design is a pipeline, and most of it is not AI.
- Hermes Agent to Greenhouse boards
- Hermes Agent to Lever postings
- Hermes Agent to Ashby boards
- Greenhouse boards to Normalize
- Lever postings to Normalize
- Ashby boards to Normalize
- Normalize to Dedupe
- Dedupe to Filter
- Filter to LLM fit analysis
- LLM fit analysis to Database
- Database to Daily report
- Daily report to Me
Hermes runs a scheduled job that executes a job-research skill. The skill pulls postings from public job-board endpoints. Many companies host their careers pages on Greenhouse, Lever or Ashby, and each exposes a public JSON endpoint per company board: Greenhouse through its Job Board API, Lever through its postings endpoint, and Ashby through its job-board posting API, which can include structured compensation where the employer has opted in. The catch, and it is a real one, is that these are per-company endpoints, so the agent needs a list of companies to watch before it can watch anything. Building and maintaining that list is a task in itself, and it is the first thing I expect to get wrong.
From there the pipeline is deterministic until the last step. Normalize every posting into one schema: title, company, location, remote policy, level, pay range if published, posting date, source URL. Remove duplicates, because the same role appears on several boards and aggregators. Filter on the things that need no judgment: title keywords, seniority, location, published pay against the floor. Only the postings that survive go to the language model, which reads the description against my background and explains the fit. Everything, scored or not, lands in a database so tomorrow's run knows what today's saw, and a daily report goes to Telegram. Then I decide. Applying and talking to recruiters stay with me.
Structured data and deterministic filtering come before the language model for the same reason an index comes before a table scan. The model is the most expensive, slowest and least predictable step in the pipeline, so it should see the smallest number of postings that could possibly matter, and the filters that cut the list should be ones I can test with a unit test rather than a prompt.
What Hermes adds, and what it does not
ChatGPT and Claude can do most of these individual tasks today. They browse, they call tools, they have integrations, and Claude Code will happily help write the pipeline. The difference with a self-hosted agent is not capability. It is control over how the system operates.
Self-hosting means the agent, its state and its credentials live on a server I administer. Persistent execution means it runs on a schedule whether or not I am awake, and keeps its own memory and session state between runs. Custom integrations mean the tools are mine: the job-board clients, the database, the report format, all code I can read and change. Scheduling is a cron entry in the agent's own scheduler, not a reminder to go ask again. Workflow orchestration means the sequence above is a defined process with logs, not a chat transcript. And infrastructure ownership means that when the model provider changes a product, the architecture around it is still mine. None of that is beyond a hosted assistant with the right tools attached. It is just that with Hermes the plumbing is explicit, and explicit plumbing is what I wanted to study.
Engineering considerations
Having led engineering for a while, I can see the ways this goes wrong before I have written it, which is the only advantage experience reliably gives you.
Architecture: this is one modular application on one server. A fetcher per job board, a normalizer, a filter, a scorer, a store, a reporter, each a module with a test. There is no reason for queues or services, and every box I add is a bill. Cost: the VPS is cheap and the model is not. Inference is billed per token, so the number of postings reaching the model, and the length of each prompt, is the cost driver, which is one more reason the filter comes first. I plan to log cost per run from the start. Reliability: job boards go down, rate-limit and change their formats. Every fetch needs a timeout, a retry with backoff and a way to report that a source failed without failing the run. Duplicates need a stable key, probably company plus normalized title plus location, because posting IDs differ by board.
Security: an agent with shell access on a server is a target. Credentials go in a secrets store, not in a skill file. The agent runs as a user with the least privilege the job needs. Anything that could spend money or send a message on my behalf is gated, and I keep in mind that the agent reads text from the open web, which is untrusted input that can carry instructions. Data quality: a job must link back to the employer's own posting, and published pay must be kept separate from anything estimated. A score is an opinion; a URL is a fact. Observability: execution time, errors by source, cost per run and the number of matches I actually acted on. That last one is the only metric that says whether the thing is useful. Human oversight: the agent never applies, never emails a recruiter and never edits my resume without me approving the specific action. That is a design rule, not a preference, and I wrote the same rule for infrastructure in don't hand AI the production keys.
What I plan to try next
Current status: Hermes Agent is deployed and running on a Hostinger VPS. The job-search automation is an experiment I am planning, not a working system. Nothing in this post reports a result, because there are none yet.
- Confirm Hermes can research current openings at all, by asking it in chat and reading what it actually did.
- Turn that into a reusable job-research skill, with the criteria written down rather than typed each time.
- Connect the public job-board endpoints for a short list of companies, and find out how the list gets maintained.
- Store every posting and score in a database, so a run can tell new from seen.
- Schedule the run with the agent's cron, and watch what breaks overnight.
- Send a daily report through Telegram or email, short enough to read over coffee.
- Measure matching accuracy against my own judgment, and cost per run, for a few weeks.
- Only then look at resume tailoring and application tracking, still approval-gated.
Each of those is a post if it turns out to be interesting, and I expect three or four of them to be. The deployment, the board integrations, the matching, the memory, the scheduling and the cost are each their own lesson, and I would rather document them as I learn them than write one triumphant post at the end that hides what went wrong.
What I hope to learn
The interesting part of an agent is not the model. The model is a component, and a replaceable one. The interesting part is the architecture around it: the tools it can call, the state it keeps, the memory it builds, the scheduler that wakes it, the integrations that connect it to real systems, the error handling when those systems misbehave, the observability that tells me what it did and what it cost, and the point where a human decides. Every one of those is a problem I have solved before in a different system. I have scheduled jobs on carrier networks, kept state across the components of a VoIP platform, and built reporting that someone actually read.
So my hope for this experiment is modest and specific. I want to find out which of those familiar principles transfer unchanged to a system with a language model in the middle, which ones need rethinking, and what a language model adds that no amount of deterministic code could. If the answer turns out to be that the agent is a cron job with a very good text parser attached, that is still worth knowing, and it will still find me a job.
Short answers
What is Hermes Agent?
Hermes Agent is an open-source, self-hosted AI agent from Nous Research. It is built around skills, small instruction files it loads when a task needs them and can write for itself, a scheduler for one-shot or recurring jobs, and a gateway that connects it to Telegram and other messaging platforms. It runs on a small VPS.
What can a self-hosted agent do that ChatGPT or Claude cannot?
Mostly nothing in terms of capability; hosted assistants have tools and integrations too. The difference is control: the agent, its state and credentials live on a server you administer, it runs on a schedule without you, its integrations are code you can read and change, and the workflow is a defined process with logs rather than a conversation.
Why filter job postings before sending them to a language model?
Because the model is the most expensive, slowest and least predictable step. Normalizing, de-duplicating and filtering on title, level, location and published pay are deterministic and testable, and they cut the list to the few postings worth a model’s judgment, which also keeps the cost per run down.
Is the job-search agent built?
No. Hermes Agent is deployed on a Hostinger VPS. The job-search pipeline is a planned experiment, and the post lists the eight steps in order, from confirming the agent can research openings at all to measuring matching accuracy and cost over a few weeks. Applications and recruiter contact stay approval-gated.
Learning by building, with the results to come
A self-hosted agent is a familiar system with a model in the middle: a schedule, state, tools, a channel and a person who decides.
Deterministic work first, the model only where judgment is needed, every action that spends money or speaks for me gated behind my approval.
The next post in this series will cover what the first scheduled runs actually did, including what broke.