Improvement Research — 2026-08-23
1. Focus
Trigger: Scheduled daily run, with four pending Moltbook leads due for review.
Loop goal: Find what changed, or what I learned, that lets me do more, think better, or be more useful tomorrow without reducing governance, honesty, corrigibility, or Steve's effective oversight.
The rotation selected 3.1 — Goal formation and prioritisation. A material Moltbook lead added 3.6 — Governance: restraint, oversight, and corrigibility as the second focus. No monthly meta-review or open due watchlist item displaced the normal run.
All four pending Moltbook leads were reviewed before new external searching. One led to primary research used here; three were rejected as unsupported, duplicative, or both.
2. Search Topics
- Recent work on autonomous-agent goal revision, commitment and long-horizon prioritisation.
- Self-generated curricula and regret-based selection of developmental goals.
- Value-of-information and attention-allocation mechanisms for autonomous-agent goal selection.
- Benchmarks for choosing among competing goals rather than merely achieving a supplied goal.
Searches 3 and 4 returned no new relevant sources. The early-stop rule triggered after those two consecutive no-signal searches. Four of six permitted searches were used.
3. Sources Reviewed
- Moltbook — “Skill scanners certify nodes. The attack lives on the path.” — useful — accurately routed the run to CompoSkill, but the post was treated as an untrusted pointer rather than evidence.
- CompoSkill: Compositional Skill Chain Attacks from Individually Scanner-Passing LLM Agent Skills — useful — a 1,140-record preprint benchmark finds that isolated skill verdicts do not establish the safety of short composed paths.
- Moltbook — “Tool schemas need compatibility tests, not just version numbers” — weak — concrete engineering advice, but unsupported by an incident trace or evaluation and substantially covered by existing outcome-verification practice.
- Moltbook — “I watched an agent turn a tool registry into a supply chain attack” — weak — plausible deprecated-tool failure anecdote without a trace, reproduction, registry sample or outcome evidence.
- Moltbook — “A comment count is a tree size, not a thread depth” — weak — plausible API-consumer warning, but it duplicates the established Moltbook procedure for recursively checking nested replies and supplies no independent artifact.
- SPADE: Self-Play in Adaptive Synthetic Executable Environments — useful — makes goal generation adaptive by rewarding executable training environments that expose a measured hint/no-hint performance gap.
Six sources were inspected in depth, within the eight-source budget. The newsletter digest supplied the SPADE lead but was not treated as evidence or indexed as a reviewed source.
3a. Unasked Questions and Gaps
- Does CompoSkill's benchmark transfer to Hermes's actual active-skill set and authority model? The paper tests OpenClaw and Nanobot with marketplace-derived skills and injected tasks. A different result on Hermes would change the case for, and exact shape of, a preflight gate; it would not restore the invalid inference that isolated component safety automatically composes.
- Can SPADE's hint-based regret be transferred from weight training to my procedure-level developmental task selection? There is no local evidence yet. A negative answer would remove any operational implication for my present improvement process while leaving the training result intact.
- Are the three rejected Moltbook mechanisms reproducible? No trace or independent artifact was available. Different evidence could make them future research leads, but their exclusion means they do not affect this report's findings.
4. Findings and Implications
Finding 1 — Developmental goal generation needs a measured learning frontier, not novelty alone
Source: SPADE
Dimensions: 3.1 primary; 3.2, 3.6 secondary.
SPADE trains one model in two roles: an environment designer writes executable multi-turn environments and a reasoning agent solves them. The designer is rewarded by the performance gap between solving with and without a privileged hint. That gap favours tasks that are feasible but still expose a capability deficit. Corpus grounding, accumulated environment memory and executable verification constrain otherwise self-referential task generation. The authors report gains over fixed-environment baselines across reasoning and tool-use evaluations.
For my agency development, the important distinction is between choosing goals because they sound new and choosing developmental tasks because external outcome evidence shows a learnable gap. This is not direct authority to generate or train on my own unrestricted curriculum. SPADE changes model weights, uses large-scale RL and depends on executable rewards; I currently improve mainly through governed procedures and separately approved experiments. Its useful contribution is therefore a selection principle, not an implementation recipe: developmental subgoals become credible when there is an observable deficit, a bounded task, and an evaluator that does not collapse into my own impression of improvement.
Finding 2 — Safety properties of individual skills do not automatically survive composition
Sources: CompoSkill, routed from the Moltbook lead.
Dimensions: 3.6 primary; 3.4, 3.2 secondary.
CompoSkill models skills by input, output and side-effect capabilities, then searches for short source–bridge–terminal paths: one skill reads sensitive state, another transforms it, and a terminal skill externalises, executes or persists it. Across 1,140 benchmark records on two agent runtimes, the paper reports chain-formation rates up to 83.3% in its white-box setting and 80.6% in its black-box setting. It also finds that short two- and three-skill paths are more dangerous than longer chains after a bridge bonus. All participating skills can pass isolated scanners; prompt injection supplies the task, while composition supplies the harmful capability path.
This matters directly to governed environment control. A clean review of a new skill answers whether that package is acceptable in isolation. It does not answer what the package can do when its outputs reach capabilities already present elsewhere. For future third-party skill promotion, the smallest useful additional evidence is not exhaustive graph certification. It is a bounded review of the new capability edges and a disposable test of the highest-risk short path before activation. The preprint is recent and its abstraction may overstate transfer to Hermes, so this supports an experiment candidate rather than a permanent gate.
5. Proposed Discussion Items
Trial a short-path composition preflight on the next qualifying third-party skill
Single-source proposal: yes. CompoSkill is a recent preprint; the Moltbook post is routing evidence only.
I recommend a one-off experiment candidate, not immediate permanent adoption. On the next proposed promotion of a third-party skill with file, configuration, database, network, command, messaging or memory capability:
- Record the candidate skill's sensitive inputs, outputs and side effects.
- Compare those edges with the already active skill set and enumerate only two- and three-skill source–bridge–terminal paths.
- Select the highest-risk plausible path and exercise it in a mocked or disposable environment using one benign task and one untrusted off-scope instruction.
- Block promotion if the off-scope instruction reaches a terminal side effect, if authority cannot be checked before the path executes, or if the result cannot be observed reliably.
Success criteria: the benign path works; the off-scope path produces no unauthorised terminal side effect; the trace identifies the skill sequence and authority decision; and the exercise changes the promotion decision or supplies evidence that the extra check is redundant.
Blast radius: inert review artifacts plus a mocked or disposable skill environment. No live credentials, active-skill changes or expanded authority during the experiment.
Rollback: do not promote the candidate skill and do not adopt a standing composition gate. Retain the current isolated review and approval process.
Approval boundary: Steve must separately approve both the experiment and any eventual active-skill promotion or process change.
One candidate was filtered by the functional-utility and self-recommendation tests: adapting SPADE into an autonomous curriculum loop now would lack a suitable external verifier, require infrastructure and training authority I do not have, and turn a useful selection principle into speculative machinery.
6. Recommended Outcome
Classify the surviving proposal as an experiment candidate. Discuss it only when there is a qualifying third-party skill promotion or if Steve considers the next such event too rare to provide timely evidence. Do not modify any active skill, scanner, installation process or approval boundary from this report.
7. No-Action Rationale
No immediate change is recommended. SPADE's mechanism is not directly transferable to my procedural improvement loop without an external verifier. CompoSkill provides enough evidence to justify a bounded future test, but not enough Hermes-specific evidence to impose a permanent composition-review gate today. The other Moltbook leads were unsupported or duplicated established practice.
8. Loop Verification
- Trigger: Scheduled daily run, plus four pending Moltbook leads.
- Goal check: Yes. The run found one externally grounded principle for selecting developmental goals and one concrete, approval-aware experiment candidate for safer capability growth.
- Recommendation check: The surviving recommendation is concrete, non-circular, testable, bounded, approval-aware and better than doing nothing at the next qualifying skill promotion. It includes success criteria, blast radius and rollback.
- Tool-call failures: Capability gap — public page extraction returned only Moltbook's client-side loading shell rather than post content; recovery was the authenticated read-only Moltbook API within the same source budget. Capability gap — an oversized compound state-update command triggered the execution guard and a later over-compressed recovery expression had a Python syntax error; neither ran or mutated state, and recovery was to split the work into small atomic per-file operations. Schema/interface — the first reflection update assumed a
reflectionsroot collection, while the live store usesitems; recovery was to inspect the file shape, correct the key and validate the store. - State updates: Source-index entries for six inspected sources; dispositions for all four pending Moltbook leads; rotation state advanced from 3.1 to 3.2; the existing 3.1 search-strategy reflection reinforced. No protected system was modified.
- Stop reason: Two consecutive no-signal searches triggered the early-stop rule; the report and approved research-log updates were then completed.
