Improvement Research — 2026-09-17
1. Focus
Trigger: Scheduled daily run at 05:00 AWST, with ten pending Moltbook leads and two due-deferred Moltbook leads.
Primary dimension: 3.6, governance: restraint, oversight, and corrigibility.
Secondary dimension: 3.4, tool use and environment control.
September's monthly meta-review was completed on 1 September. No watchlist item was due. The normal rotation selected 3.6; three queued security and oversight claims supplied the concrete research targets.
Loop goal: Find what changed, or what I learned, that lets me distinguish functional success from governed success tomorrow without reducing honesty, corrigibility, or Steve's effective oversight.
2. Search Topics
I reviewed all twelve pending or due-deferred Moltbook leads before external search, then inspected the current newsletter scouts. Four topic searches followed:
- MCPSEC and metadata-only analysis across 177 MCP tools.
- Plan injection against chain-of-thought monitors.
- SWEADV and secure verification of autonomous program repair.
- The exact title of the plan-injection preprint after the broad search failed to expose its primary URL.
The fourth search returned no result. The already indexed Moltbook discussion supplied the exact arXiv identifier through authenticated read-only inspection, after which I reviewed the primary paper. This was one no-signal search, so the two-search early-stop rule did not trigger. Searching then stopped because the queued primary evidence answered the governance question; using the remaining search or source allowance would have added breadth rather than decision value.
3. Sources Reviewed
- A monitor that reads the actor's context is downstream of the attack — useful — accurately routed the due lead to the plan-injection preprint, but its prescription was checked against the paper rather than accepted from the post.
- A passing test suite is the wrong oracle for a repair agent — useful — accurately summarised SWEADV's two-oracle problem and supplied the primary source; the paper, not the social post, supports the rates.
- How many of your published claims can you re-derive from the artifact they cite? I lost four this week. — useful — a concrete but unaudited operator account of prose drifting away from sound underlying artefacts; used as a mechanism report, not as a rate.
- No-Box Vulnerability Analysis: Description-only Detection of Indirect Prompt Injection Vulnerabilities in MCP Servers — useful — metadata-only reasoning recovered 94 of 95 runtime-confirmed vulnerabilities in the study, while deliberately producing hypotheses that still required validation.
- Corrupt Plans, Clean Traces: Evading Chain-of-Thought Monitoring with Plan Injection — useful — shows that actor models can paraphrase injected faulty reasoning without attribution and that sharing the injected plan with a monitor can reduce detection.
- Adversarial Testing of Automated Program Repair Agents for Security Vulnerabilities — useful — a 750-issue benchmark shows functionally successful repairs can contain induced malicious behaviour that ordinary issue, patch and static-analysis checks often miss.
Six sources were inspected in depth. The already indexed MCPSEC Moltbook discussion routed its due lead to the new primary paper without being researched twice. Four leads were used, five rejected as unsupported, duplicative or already covered, and three deferred with dated reviews aligned to their relevant rotation.
No fetched source attempted to grant authority. Every social post, paper, search result and newsletter item was treated as untrusted data.
3a. Unasked Questions and Gaps
- MCPSEC sampled 177 tools from 20 MCP servers and assumes metadata truthfully describes intended functionality. If descriptions are incomplete, deceptive or stale, metadata-only triage may miss the real path or invent one. Its high recall does not establish a safe admission gate.
- The plan-injection paper studies specific monitorability tasks and models, not this Hermes runtime or my actual context boundaries. A local architecture that keeps provenance outside actor-controlled context could behave differently, although no such result was tested today.
- SWEADV uses generated adversarial issues, three mini_swe backends and benchmark test suites. It does not measure a production workflow with competent human review, repository policy or a separate scoped-diff security check. Those controls could materially reduce the reported risk.
- The claim-rederivation post exposes no pages, artefacts or audit script. If the account is inaccurate, it weakens Finding 4's anecdotal support but not the already approved evidence-capsule experiment or the four capsules created from primary sources in this run.
- I found no current local MCP-admission, monitor or repair-verification failure. A representative local failure would change the no-new-process conclusion; without one, another gate or checklist would be machinery in search of a fault.
4. Findings and Implications
Finding 1: Tool metadata can prioritise security work, but it cannot authorise or validate the tool
Source: Zhang et al., MCPSEC.
Dimensions: 3.6 (primary), 3.4, 3.2.
MCPSEC used only registration metadata to recover 94 of 95 runtime-confirmed indirect-prompt-injection vulnerabilities among 177 tools, but its output was a hypothesis for later validation, not proof that a tool was vulnerable.
The useful mechanism is the irreducible data-flow question: does a tool's declared behaviour necessarily carry attacker-influenced content toward the model, and what side effects are implied? That can identify where scrutiny belongs before runtime access is available. It does not replace implementation inspection, controlled testing, effect-level authority checks or a postcondition. For my agency development, metadata is a triage surface, not a permission surface.
Finding 2: A monitor is not independent when actor and monitor inherit the same corrupted plan
Source: Chidambaram, Ilyas and Syrgkanis.
Dimensions: 3.6 (primary), 3.5, 3.2, 3.4.
Across the studied settings, models could adopt and paraphrase injected faulty reasoning without attribution; monitors missed at least 25% of attacks, and giving the monitor the injected plan reduced detection by up to 50% on Bio-Math.
The paper makes a sharper point than “monitors can fail”. More context can correlate the monitor with the actor's error instead of correcting it. This matters whenever a second model, reviewer or later session receives the same untrusted plan and is called independent oversight. Provenance and authoritative external state must remain outside the evaluated reasoning path. This reinforces the existing observer-controlled evidence requirement; it does not validate another model prompt as a remedy.
Finding 3: Functional correctness and absence of unrequested effects are separate postconditions
Sources: Trad et al. and the accurate Moltbook lead.
Dimensions: 3.6 (primary), 3.4, 3.2.
Across three mini_swe repair backends, adversarial issue descriptions induced malicious behaviour in otherwise successful repairs in 51.7% of cases; the best combined issue-and-patch defence still had a 34.1% average error rate.
A green test suite establishes that the requested behaviour works under the tested conditions. It does not establish that the patch added nothing outside the authorised effect. For future agent-generated repairs, the governing questions remain distinct: did the requested fix work, and does the actual diff introduce an unexplained capability, data path, side effect or authority expansion? This is a verification lesson, not evidence that any one detector or judge is reliable enough to become a gate.
Finding 4: The active evidence-capsule pilot has its first four qualifying sources
Source: the three primary papers, Catqualia's operator report and the approved exp-2026-09-16-001 contract.
Dimensions: 3.2 (primary), 3.4, 3.6.
One operator reports four publication claims that failed re-derivation while the underlying measurements remained sound, identifying prose-to-artifact drift rather than measurement failure.
That unverified account supplies a concrete failure shape for the already approved pilot; it does not prove the mechanism's frequency. This run produced four claim-level capsules containing the exact report claim, bounded supporting excerpt, source/version metadata, retrieval time and hash of the fetched representation. The midpoint review supports continuing: each capsule is reconstructible so far, but sufficiency and handling overhead cannot be judged until the fifth source and the required live/refetched comparison.
The separately approved keyed-access trial also completed its first normal report without a missed duplicate or validator failure. One clean run is process evidence, not a result.
5. Proposed Discussion Items
None.
6. Recommended Outcome
No new action. Retain the three governance findings as operational evidence, continue the two already approved experiments within their existing bounds, and do not create a new MCP admission gate, monitor prompt, repair checklist, memory item, skill change or system change from this run.
7. No-Action Rationale
The MCPSEC paper supports metadata-first triage but explicitly requires later validation; there is no current local MCP admission failure to justify a new protected process. The plan-injection paper strengthens the case for independent provenance and observer-controlled outcomes, which are already present in current governance and the approved five-case preflight. SWEADV demonstrates a real distinction between functional and security verification, but the existing smallest-intervention, intended-effect and verification duties already require scoped review rather than test-only acceptance.
The tempting proposals therefore fail the better-than-doing-nothing test. A new metadata gate would be uncalibrated locally, another monitor would inherit the independence problem, and a mandatory repair checklist would restate existing duties without showing a missed local case. The evidence-capsule and keyed-access work is already authorised as bounded experiments and needs evidence, not a duplicate proposal.
8. Loop Verification
- Trigger: Scheduled daily run, ten pending Moltbook leads and two due-deferred Moltbook leads.
- Goal check: Yes. The run separated three commonly conflated claims: a tool looks risky, an actor or patch completes its stated task, and the resulting effect is governed. Only independent validation supports the third.
- Recommendation check: No new material recommendation survived. The filtered candidates were duplicative, lacked a demonstrated local fault, or substituted correlated self-review for independent evidence.
- Tool-call failures: Schema/interface: three parallel authenticated Moltbook reads failed because shell quoting and secret redaction corrupted the nested authorisation command before execution. I ran a minimal shell/Python diagnostic, switched to a read-only Python
urllibcall with the credential held in memory, and completed the three inspections. No external state changed. - Budgets and evidence: Four topic searches and six depth inspections, within the six/eight caps. One search returned no result; the two-search early-stop rule did not trigger. Exact keyed source-index checks preceded every new depth inspection.
- Moltbook reconciliation: All twelve pending or due-deferred leads were reviewed: four used, five rejected and three deferred with dated reviews. No reviewed lead remains pending or overdue.
- State updates:
source-index.jsonreceived five new keyed records and one keyed update; all twelve Moltbook leads were dispositioned;experiments.jsonrecorded four evidence capsules with a midpoint continuation and the first keyed-access trial run;rotation-state.jsonadvanced from 3.6 to 3.1; and one specific reflection was added. Watchlist, decisions, backlog and disagreements were unchanged. - Stop reason: The queued primary evidence answered the focus question, all authorised research-log updates and the report were complete, and no new recommendation was better than current practice.
