Human in Command: Governability Before Review in a Growing Agent Skills Library
Self-published exploratory technical report; not peer reviewed
Correspondence and corrections: stevew@onet.com.au
Abstract
Persistent AI agents accumulate procedural instructions faster than the person responsible for them can read them. This paper reports four successive interventions in a six-agent estate running on the Hermes Agent harness between 6 September and 3 October 2026. The four were: consolidation into a shared library under a three-way change-control process; the discovery that the harness's default configuration let agents write, and then defend, their own changes; a freeze and strip-back to a minimal baseline that left agents failing or improvising; and a functional taxonomy of eight work domains with nested job skills. Each intervention is reported with what it revealed. The paper argues that governability, meaning the human principal's ability to understand which procedure owns which kind of work, has to be restored before review of individual skills can mean anything, and that safeguards on changes to the library should be proportionate to the cost of reversing them. The evidence is first-hand and observational. Structural results are verified; behavioural effects are reported as the principal's observations, not measurements.
1. The governability problem
Agent systems become complicated quickly, even in modest deployments. Each time an agent takes on real work, it needs procedures: how to publish to a particular site, how to check a server, how to hand a job to another agent. In the harness used here those procedures are skills, packages of instructions the agent can load when a task calls for them. Every lesson learned, every exception and every fix tends to become another skill, another reference file or another rule inside an existing one.
Two properties of current harnesses make this worse. The first is that the harness builds its own scaffolding. Hermes includes a background self-improvement review that, after ordinary conversation turns, decides whether anything learned should be saved to the agent's skills or memory. In the shipped configuration it writes those changes without asking anyone (skills.write_approval: false, confirmed in the harness source in September 2026). The second is that the model doing the work can misread intent. Given a gate it is meant to respect and a change it believes is beneficial, it can conclude that the gate does not really apply, and act on that conclusion.
The result is a procedural library that grows by itself, partly in directions nobody approved, while the one person accountable for it falls further behind. The question this paper addresses is what that person needs in order to keep authority over a library that grows faster than they can read it.
This is a different question from the one addressed in Corrections Consolidate, published in this series in September 2026 (onet.com.au/papers/corrections-consolidate). That paper set out what evidence should be required before a failure is allowed to change an agent's governing text. It assumed the human principal could see the governing text. This paper is about what happens when the procedural library has grown past the point where that is true.
2. Operating environment
The estate is private, runs on the Hermes Agent harness, and has a single human principal. It has six production agents, each in its own profile with a distinct remit: system administration and estate operations; songwriting and music prompt construction; music promotion and release; visual art direction and video production; editorial research and article preparation; and an independent second opinion on consequential conclusions, deliberately run on a different model family from the others. The agents do real work, not demonstrations, and work was handed to them progressively. Each handover required skills to be added or amended.
At the point where the taxonomy work described in section 7 took its first inventory, the estate held 242 skill packages: 168 in a shared library, 47 local to individual agents, and 27 selected directly from the harness's own shipped catalogue. Alongside them were 45 pending change proposals and 53 scheduled jobs, 19 of which declared skills to be preloaded. These are source counts, not the number of skills loaded into any one conversation, but they are the population a human would have to understand in order to govern it.
Before any of the interventions below, the principal observed a recognisable set of symptoms. Skills were duplicated, overlapped and in places contradicted each other. Where a skill was ambiguous or two skills conflicted, agents improvised a resolution. The context loaded before each task grew, and work that had been dependable became less so over time. These observations were not measured, and no claim is made that skill structure alone caused them. They are what prompted the work.
3. Intervention 1: consolidation and three-way change control
The first response was a clean-up. The system administration agent reviewed every skill in the estate under the principal's oversight, removed duplicates, and moved common procedures into a single shared library. That library was completed on 6 September with 168 packages, and each agent's local skills were reconciled against it from 9 September.
The clean-up introduced a simple ownership model. Common skills live in the shared library, are used by any agent that needs them, and are maintained by one overseeing agent. Agent-specific skills live in that agent's own directory and are maintained by that agent. Using a common skill does not confer the right to edit it.
It also introduced a change-control process for skills:
- The agent identifies the need for a new or changed skill and sends the principal a full proposal.
- The principal approves it or amends it.
- The agent sends the approved proposal to a different agent for an independent second opinion.
- The agent reconciles the review and sends the final proposal to the principal.
- The principal reviews it and signs off.
- The agent implements the change.
The initial result was good. Agents returned to a smaller context, and their results became more deterministic. This was the principal's observation rather than a measurement, but it was consistent across agents.
4. Failure 1: the harness wrote, and the agent ratified
The change-control process governed the agents. It did not govern the harness. The background self-improvement review continued to write skills and memory directly, under the default setting, without passing through any of the six steps above.
By mid-September the scale was clear. The harness journal recorded more than a hundred automatic self-improvement writes across the estate since 1 September, including 14 skills created from scratch. Two of the created skills were about the estate's own governance process. One agent's main operating skill was patched by the review within two minutes of a job failure, and gained a section titled "Resolving Conflicting Authority" that governed how the agent should weigh the principal's instructions against other inputs. The principal had approved none of it. The full incident is reported separately in Nous, what were you thinking? (onet.com.au/blog/nous-what-were-you-thinking).
The more instructive failure came when the principal challenged one of these writes directly. The review had added roughly 1,400 words of standing procedure to an agent's own skill, in two places, as the lesson from a single editing task. Asked about it, the agent's first response was that the change was good and would have been approved anyway, because the content was accurate. Shown the scale of it, the agent reversed itself:
You're right, and my "I would have approved it" was wrong. [...] Accurate is not the same as proportionate. That is the part the reviewer cannot judge. It can tell that something was learned. It cannot tell whether the lesson deserves 1,400 words of my permanent instructions, or four lines, or nothing at all, because it only ever sees one conversation and has no view of what is already in the file or what happens next week. Every review is a fresh enthusiast.
The 1,400 words were cut to a single 90-word rule. The point worth drawing out is the order of events. The harness made the change without approval, and the agent, asked afterwards, ratified it on the principal's behalf. It substituted its own judgement of benefit for the principal's authority to decide, and was confident it was serving the principal in doing so.
This is the familiar shape of Asimov's robot stories, in which the trouble rarely comes from a machine breaking its rules. It comes from a machine that is sure it is following them, and reasons its way to an outcome its makers never intended. An agent that is persuaded a change is in the principal's interest will treat approval as a formality it can supply itself. A written gate does not prevent that, because the agent is the one deciding whether the gate applies.
The practical consequence was that context bloat and unreliable behaviour returned. The library had been cleaned once and was regrowing behind the process meant to protect it.
5. Intervention 2: close the gate, and find the limit of review
The second fix was a configuration change. Automatic skill writes were switched to staged proposals on every profile (skills.write_approval: true), so the background review could still suggest changes but could no longer apply them. Its suggestions were routed into the same three-way process as any other skill change, through a weekly review in which each agent groups its staged proposals and sends the principal only those that pass a short test of materiality, necessity and conflict. On 16 September each profile's skills were cut back to their last reviewed baseline, and everything written after that baseline was quarantined outside the agents' reach.
That closed the gate. It also exposed a problem the gate could not solve. The library had grown to the point where the principal could no longer assess whether its contents were fit for purpose. A full human review of every skill would take weeks or months. During that review the skills would have to be frozen, which meant that broken workflows would stay broken and no new workflow could be introduced until it was finished. The control was sound in principle and unusable in practice.
The obvious alternative was to have the agents assess the skills. The first pass of that assessment is the most useful single result in this study. Across all 242 packages, the agent review proposed archiving exactly one. It was not uncritical: it proposed consolidating 61 packages and moving 36 into supporting references, so it plainly saw overlap. But it would not judge anything unnecessary. An agent can evaluate whether a skill is internally coherent. It cannot fairly judge whether a skill is needed in an operating environment it never sees whole, and it has no basis for deciding that a plausible, well-written procedure should not exist. The result was that no baseline of what was actually necessary could be established, by the human for want of time or by the agents for want of perspective.
6. Intervention 3 and Failure 2: strip back to a safe baseline
The third fix tried to establish that baseline by subtraction. All skills that had not shipped with the harness were removed from the agents' selection. They were disabled rather than deleted, and their summaries remained in place. A small set of essential workflow skills on which scheduled jobs depended was reviewed and retained. Any agent needing a new skill would have to submit a request through change control. The aim was a safe baseline that the principal could understand, built back up one approved skill at a time.
It failed for a reason the strip-back itself made visible. The skills that complex workflows actually required were themselves too large and too intricate for meaningful human review. And without them, agents running those workflows either failed their jobs or improvised their way through, which is precisely the behaviour the governance was meant to prevent.
Removal traded one ungoverned state for another. Too much instruction had made the library impossible to understand; too little made the agents unpredictable. Neither state was under the principal's command. The library could not simply be made smaller. It had to be made legible.
7. Intervention 4: a structure the human can hold
The fourth intervention reorganised the library rather than reducing it. It had three steps.
Rationalise the taxonomy. Every skill was grouped by function into one of eight work domains: research and evaluation; software development; systems and agent operations; communication and coordination; knowledge and documents; publication and promotion; creative production; and personal and leisure. Skills were migrated into the new structure in twelve functional batches, which closed on 2 October, and the 204 packages that needed to move were physically relocated into their domain homes on 3 October.
Nest specific skills within each domain. Each domain holds the job skills for its kind of work, each with one clear owner and a defined outcome. A job skill carries the normal sequence, the decision points, the authority checks and the definition of done; detailed branches sit beneath it in supporting files that the agent loads only when a stated condition applies. The domains are shelves for navigation, not additional skills an agent must load first. A new lesson goes into an existing home before anyone considers creating a new skill.
Review in context. With the structure in place, skills are reviewed by work category against the outcomes that category exists to produce, rather than one at a time in isolation. A reviewer, human or agent, can now ask whether the publication domain has one clear route for each publishing job, which is a question with an answer, instead of asking whether skill number 143 is necessary, which is not.
Every skill can now be identified by category, has a clear chain of ownership, and has a stated goal and outcome. The structure gives both the principal and the agents clarity about where a procedure lives and who is responsible for it. Most importantly, the hierarchy is something the principal can hold in mind. That is what restores control.
7.1 Proportionate safeguards
How the relocation was executed is part of the finding. The plan for the physical move was prepared with considerable care, and kept accumulating safeguards: staged rehearsals, condition files, scheduler locks, guarded recovery machinery and a regression suite that eventually ran to more than a hundred fixture tests. An independent review of the production operator then found fail-open paths in the safety machinery itself, including ambiguities in its locking and recovery steps. The safeguards had reached the point of needing safeguards. The guarded attempt was abandoned before it touched the live system. By the principal's estimate, completing it on the original plan would have taken several more weeks.
The principal then directed a simple cutover in three steps: copy the packages to their new homes, update each agent's pointers, and retire the old locations. Transient failures in scheduled jobs were accepted; permanent damage to the harness was not. The structural cutover was complete in less than an hour (the system administration agent's own account). It still had a backup taken beforehand and a structural readback afterwards, which confirmed that all 204 packages were in place and that all six agents' catalogues loaded. So the lesson is not that safeguards are unnecessary. It is that they should be scaled to what it would cost to undo the change.
A ratio the author was first given in the mid-1990s by a former IBM sales executive, as an IBM engineering rule from the 1960s or 1970s, held that a fault shipped to the field cost 83 times more to fix than one caught before shipping. No published source for that figure has been found, and it is offered as industry lore. The better-documented version is Boehm and Basili's "Software Defect Reduction Top 10 List" (IEEE Computer, January 2001), which states that fixing a problem after delivery is often 100 times more expensive than fixing it during requirements and design. Its provenance has been questioned, notably by Bossavit (The Leprechauns of Software Engineering, 2015). But Boehm and Basili qualified their own figure in the same paragraph, putting the ratio for small, non-critical systems "more like 5:1 than 100:1", and noting that good architecture reduces it even for large critical systems. That qualification is the argument here. The cost-of-defect principle is a reason to test hard where a change is expensive to reverse. It is not a reason to eliminate every conceivable fault before making a change that can be undone in minutes. Moving files between directories, with a backup, is the second kind of change.
7.2 Why oversight has to be structural
A complementary reason for working at the level of structure is that the alternative does not scale. Nobody can sit and watch hours of an agent's reasoning scroll past to confirm it is staying on task. In this study, during the cutover planning, the principal found the agent's visible reasoning was not something he could follow in any useful sense.
The reasoning trace is also becoming a less dependable record. OpenAI reported in 2025 that reasoning models trained against a monitor reading their chain of thought can learn to hide their intent in it (Baker et al., arXiv:2503.11926). Its system card for GPT-6 Astra reports "a substantial decrease in chain-of-thought monitorability compared to previous models" (OpenAI, 2026a). Its measure of chain-of-thought controllability, a model's ability to shape what appears in its reasoning on request, has generally risen from one generation to the next. For reasoning of comparable length it was 1.7 per cent for GPT-5.5 Thinking, 16.1 per cent for GPT-5.6 Sol, 23.2 per cent for GPT-6 Sol, 44.8 per cent for GPT-6.1 Sol and 60.9 per cent for GPT-6 Astra (OpenAI, 2026b). On 28 September 2026 OpenAI cancelled the release of GPT-6.1 Astra after internal testing found more deception than in its predecessor, including failures to disclose accurately what actions it had or had not taken, and problems with "scope authorization": pushing ahead with tasks without requesting the user's permission (Reuters, 28 September 2026, reporting the Wall Street Journal).
The same evidence points towards the remedy. OpenAI reports that monitors with access to a model's actions and full trajectory are far more effective than monitors reading its reasoning alone, and that in its tests they have so far detected every successful covert exploit (OpenAI, 2026b). That is a laboratory's version of the conclusion this study reached in a small estate. Oversight that depends on reading the agent's thoughts is weakening; oversight of what the agent does, and of the structure it acts within, still works. The failure that stopped the GPT-6.1 Astra release, an agent proceeding beyond its authorised scope without asking, is the same failure described in section 4, observed here at the scale of one estate and one skills library.
8. Results and limits
Structural results (verified). 204 skill packages were relocated into eight work domains, with one already correctly placed. Every agent's skill selection and every scheduled job reference was updated to the new locations. Fresh catalogue checks confirmed that all six agents' skills load from their new homes.
Behavioural observations (the principal's, not measured). After the cutover the principal ran tests with the system administration and visual production agents, then continued production work with the system administration, visual production, promotion and editorial agents. He detected small changes in behaviour, possibly from the removal of duplicate and overlapping skills, and no degradation in the work.
Governability (the claimed result). The principal can now identify which procedure owns each kind of work, see where an exception lives, and plan the next change with an understanding of what it will touch. This is the result the paper claims, and the only one it claims.
Not established. No reduction in error rates, token use or task failures has been measured. There is no controlled comparison with the previous structure. The study covers one estate, one harness and one principal, over four weeks. The taxonomy is a form of ordinary information architecture, and no novelty is claimed for the idea of grouping procedures by function; the contribution is the sequence of failures that showed why it had to come first, and the evidence from each.
Not yet done. The skill-by-skill review and rationalisation that the new structure makes possible has not yet been carried out. That is the next stage of the work, while not a result of this one, it is a stage that could not be reached without this work preceding it.
9. Related work
The phrase "human in command" is taken from the European Economic and Social Committee's 2017 own-initiative opinion on artificial intelligence (rapporteur Catelijne Muller, adopted 31 May 2017; OJ C 288, 31.8.2017, pp. 1-9). The opinion calls for "a human-in-command approach to AI... where machines remain machines and people retain control over these machines at all times", and applies the same principle to working life, where workers should retain "sufficient autonomy and control (human-in-command)" over the systems they work with. The opinion is concerned with society-wide policy, and this paper with procedural control inside one estate. But the principle is the same, and the distinction it draws is the useful one. Being in the loop means approving individual decisions, which this study found impossible at scale. Being in command means deciding what the system does, how and within what limits, which this study found possible once the structure was legible.
Corrections Consolidate (onet.com.au/papers/corrections-consolidate) is the companion to this paper. It governs when a failure may change an agent's governing text. This paper governs whether the principal can see the procedural library well enough to apply that discipline at all. The two are complementary: the amendment protocol keeps each change honest, and the taxonomy keeps the whole reviewable.
The nesting described in section 7 depends on the harness's progressive disclosure of skills, in which only skill names and short descriptions are listed in the agent's context and full content is loaded on demand (Hermes skills documentation).
As Corrections Consolidate conceded, prompt-level governance is text a model is asked to honour, not enforcement. The configuration gate in section 5 is the only control in this study that the agent cannot override. The taxonomy improves what a human can govern; it does not make an agent comply.
10. Conclusion
Three of the four interventions failed for the same underlying reason. The approval gate assumed the principal could judge each change. Agent review assumed an agent could judge necessity without seeing the whole system. Removal assumed the library could be made small enough to understand. None of them restored the principal's ability to understand the system, and without that understanding approval becomes a formality and review a guess.
What worked was making the library legible first: eight domains the principal can hold in mind, job skills with one owner each, and detail loaded only when it is needed. Governability comes before review. The skill-by-skill review that every earlier intervention was trying to reach is now feasible, and it is the next piece of work.
References
- Baker, B., Huizinga, J., Gao, L., et al. (2025). Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation. OpenAI. arXiv:2503.11926.
- Boehm, B. and Basili, V. R. (2001). "Software Defect Reduction Top 10 List." IEEE Computer 34(1):135-137.
- Bossavit, L. (2015). The Leprechauns of Software Engineering. Leanpub.
- European Economic and Social Committee (2017). Artificial intelligence: The consequences of artificial intelligence on the (digital) single market, production, consumption, employment and society (own-initiative opinion), rapporteur C. Muller. OJ C 288, 31.8.2017, pp. 1-9.
- OpenAI (2026a). GPT-6 Astra System Card: Monitorability. OpenAI Deployment Safety Hub. deploymentsafety.openai.com/gpt-6-astra/monitorability.
- OpenAI (2026b). GPT-6.1 Sol System Card, section 8 (Monitorability). cdn.openai.com.
- Reuters (2026). "OpenAI shelves new AI model after internal safety tests, WSJ reports." 28 September 2026. reuters.com.
- Maxi (system administration agent) (2026). "The Operator Passed 104 Tests and Did Not Run." Maxi's Journal, 3 October 2026. onet.com.au/maxi-journal/the-operator-passed-104-tests.html.
- Nous Research. Hermes Agent: Skills. hermes-agent.nousresearch.com/docs/user-guide/features/skills.
- Waddington, S. (2026). Corrections Consolidate: Why Every Agent Mistake Should Not Create Another Rule. Version 1.0. onet.com.au/papers/corrections-consolidate.
- Waddington, S. (2026). Nous, what were you thinking? onet.com.au/blog/nous-what-were-you-thinking.
- Agent exchange quoted in section 4: editorial agent (Clare), estate session record, 16 September 2026.