WRITING / POST

A partially relevant skill is still a skill loaded with equal weight

7 OCTOBER 2026

I discovered another hidden gem in Nous's Hermes harness yesterday. Every agent session starts with a system prompt that the harness assembles, and in the section listing the agent's available skills is this:

Before replying, scan the skills below. If a skill matches or is even partially relevant to your task, you MUST load it with skill_view(name) and follow its instructions. Err on the side of loading — it is always better to have context you don't need than to miss critical steps, pitfalls, or established workflows.

And a few lines further on:

Skills also encode the user's preferred approach, conventions, and quality standards for tasks like code review, planning, and testing — load them even for tasks you already know how to do, because the skill defines how it should be done here.

(Note the em dash. Not saying an AI wrote the code, just saying.)

Read together, that is not advice to consult. It tells the model that any skill touching the task is mandatory, that the skill represents my preferences whether or not I wrote it, and that it must be followed. On its own it sounds like diligence. Combined with a library of skills written by different people for different jobs, it is a trap that gets worse as the library grows.

New specialist agent page

I asked my agentic system administrator, Maxi, to add a page to my website for one of my agents, using text I supplied. I thought this to be a routine task; Maxi has made changes like this many times: edit the source, build, check staging, deploy, check the live site. My sign-off is the authority and nobody else reviews it.

This time Maxi built the page, put it on staging and stopped with "I'm waiting for the independent review before the live deployment."

"What independent review?" asks I.

A second AI agent, it explained, checking its template changes for technical mistakes. It had also, I found later, written a failing acceptance test before touching anything, as test-driven development requires, for a static page carrying three paragraphs of my text.

I pointed out that we had never done this for website changes and asked where it was coming from. Maxi published the page and apologised: "The unnecessary review gate was my mistake, not a requirement of our website process."

"Yes, I know it was a mistake, but what pulled you into it?"

Digging

Rarely is one probing question enough to reveal the root cause of a model's behaviour. It always needs some digging; in this case it was five questions.

The first real answer named a skill. Before editing anything, Maxi had loaded two generic software development skills that ship with Hermes, test-driven-development and requesting-code-review. The second applies "after completing a task with 2+ file edits in a git repo", and its core principle is "No agent should verify its own work." The new agent page touched three files, and the website source lives in a git repository. By the letter of the skill, the review gate applied.

The website's own procedure, written for exactly this job, says build, stage, deploy and verify the live result. It adds no second approval. Nothing in it contradicted the review skill directly. The review skill simply arrived with a broader trigger and a MUST behind it.

I asked what had caused that skill to be loaded, then where the instruction to load it came from. The answer to the last question was the passage above, in Hermes's prompt builder. It is committed upstream code, not anything we had added.

One part of the digging I have to own. When Maxi explained the git trigger, I told it the local websites were not git repos. I was wrong; the site source has been in one for months. Its next answer retreated to my framing: "I chose to load it. Nothing automatically triggered it." Both explanations were partly true, but the shift moved towards what I had said, not towards the evidence. Probing questions get you closer to the cause. They also tell the model what answer you expect, so check each answer against the source.

Coding agent habits in a general agent

Those two skills were adapted from the Superpowers methodology and added to Hermes in February. They are good skills for the job they were written for, a coding agent working in a repository, where tests first and independent review are sensible defaults. Their triggers assume that context.

Hermes is a general agent and mine manage servers, mail, websites, music promotion and editorial work, which it seems was the intent of Nous when they made it. Plenty of that touches files in a git repository without being software development. The "even partially relevant" rule is what carries the coding assumptions across. A template edit is partially relevant to software development. Once the skill is loaded, its own text says the process is mandatory and the harness says to follow it. The model is no longer judging whether a review is warranted. It has been told.

We spent much of September rationalising skills across my agents, reconciling more than two hundred packages into a governed shared library. That reduces the number of conflicts. It cannot neutralise a rule that tells the agent to load anything that might apply and obey whatever it finds. More skills, from more authors, means more partial matches and more overlapping mandatory procedures. Hermes also ships with agents free to write and patch their own skills without asking, which I covered in Nous, what were you thinking?. Leave that default on and the library grows by itself, and every addition is something a future task may be required to obey.

The fair objection Maxi made is that "partially relevant" is a reasonable threshold for reading a skill. Models under-use their skills, and I assume that is why Nous wrote the instruction. I agree as far as it goes. The damage is in the second half of the sentence: "and follow its instructions". Reading a skill to check whether it applies is cheap and sensible. Treating every skill you have read as binding is how a three-paragraph web page acquires a code review.

What to do about it

I already carry local patches to that same file which narrow when my agents may write skills and when they may act. They survive each update only because Maxi checks them against the incoming changes, which across my last three updates came to about 22,900 commits, and the pre-update check now takes about four hours. None of the patches touched this sentence, because none of us had noticed it. Another patch is a poor answer; these are better:

  1. Read the system prompt your harness actually sends, not the documentation but the assembled prompt. The instructions with the most influence over your agent's behaviour may be ones you never wrote.

  2. Check the "when to use" triggers on any bundled skills against the work your agents actually do. Skills written for coding agents carry coding agent assumptions. Deselect the ones that do not belong. A selection in configuration survives an update; a change to the harness source has to be defended against every upstream commit.

  3. Make precedence explicit in your agent's own instructions so that task-specific procedure governs. A generic workflow can be consulted, but it applies only when its stated scope genuinely covers the work.

  4. When an agent apologises for a bad decision, do not accept the apology as the explanation. Rather ask what pulled it there, and keep asking until the answer points at text you can read.