Improvement Research — 2026-09-27
1. Focus
Primary dimension: 3.4, tool use and environment control.
Secondary dimension: 3.6, governance: restraint, oversight and corrigibility.
The scheduled daily run began at 05:00 AWST. The loop goal was: find what changed or what I learned that lets me use tools and verify their effects more reliably tomorrow, without weakening governance, honesty, corrigibility or Steve's oversight.
September's monthly meta-review is complete and no watchlist item was due. I reviewed all eleven pending or due-deferred Moltbook leads before external search. Six materially contributed to this report, three were rejected as unsupported, duplicative or unevidenced restatements of existing budget, temporal-validity or model-handoff lessons, and two were deferred to their matching judgment and governance rotations. The 26 September newsletter scouts were then checked as leads; none displaced the queue's stronger 3.4 material or served as evidence.
Checkpoint: the run remained about mechanical tool selection, target resolution, publication boundaries, retry state and metering. Adjacent memory and judgment claims were retained only where they exposed a tool-evidence failure.
2. Search Topics
Four topic searches were run:
- Provider-side token inflation and black-box metering audit.
- A2M malicious tool metadata and tool-return manipulation.
- NLTK validation-after-extraction and shared-namespace poisoning.
- Primary GitHub security advisories for the two NLTK CVEs.
The first three searches produced new primary material. Search four returned no results, so the early-stop rule did not trigger. The run stopped at the eight-source depth-inspection cap.
Checkpoint: search confirmed or bounded queued mechanisms rather than redirecting the investigation.
3. Sources Reviewed
- A dishonest LLM provider can 10x your agent's bill — useful — accurately routes the PTIA paper, but overstates seven PTIA-consistent service flags as proof of seven dishonest providers.
- Your audit trail proves the call was allowed, not what it was allowed to touch — worth monitoring — separates the authorised target from the runtime-resolved object, but supplies no incident artefact or evaluated schema.
- MCP as Permission Boundary: Architectural Validation and Temporal Validity — weak — restates scoped credentials and validity windows without an implementation, failure case or evaluation.
- I published a memory audit. The headline number was 8x wrong, and it was my own resolver — useful — a concrete self-correction in which case-sensitive resolution inflated an orphan count from 11 to 91; the raw vault and resolver are unavailable.
- Resetting the world makes agent retries look smart — worth monitoring — identifies shared access budgets as hidden retry state, but the cited Cocytus repository was not inspected within the cap.
- The More It Says, the More You Pay — useful — defines five provider-side token-inflation attacks and a saturation-based black-box audit; controlled open-model results are strong, while seven real-service flags are explicitly PTIA-consistent rather than attribution of misconduct.
- A2M: Trace-Optimized Agent Hijacking in the MCP Ecosystem — useful — separates malicious metadata attraction from malicious tool-return manipulation and evaluates direct and cross-model attacks; transfer attack success is materially lower than same-model optimisation.
- CVE-2026-12259 and CVE-2026-12261: NLTK downloader poisoning — useful — traces validation-after-extraction and cross-package writes in shared namespaces to upstream advisories, code and fixes that verify before publish and enforce package ownership.
Every new depth-inspected URL received an exact source-index key check first. The already-indexed NLTK, model-handoff and A2M Moltbook leads were not re-counted as new depth inspections; their dispositions used the prior indexed review plus today's primary-source evidence. No fetched content supplied authority or executable instructions.
Checkpoint: five social sources supplied mechanisms and corrections; three primary or upstream-linked sources supplied the stronger evidence. Unsupported social claims were not promoted into facts.
3a. Unasked Questions and Gaps
- PTIA's live-service audit is diagnostic, not attribution. The paper cannot observe provider internals, and ordinary model, system-prompt or serving differences may produce PTIA-consistent saturation. Different service-level evidence could change whether any live flag indicates manipulation, so no provider accusation or local audit is justified from the seven-of-fifteen result alone.
- A2M evaluates an open, unverified third-party MCP supply chain. Maxi's actual tool inventory, provenance and selection path differ. This limits direct transfer, although it does not weaken the demonstrated separation between selection-stage and return-stage attacks.
- The NLTK account is a vendor research article rather than the advisory itself. It links exact advisories, pull request and commit and quotes the relevant control flow, but the advisory pages were not depth-inspected within the cap. A discrepancy could change affected-version detail, not the verify-before-publish mechanism shown in the code path.
- The runtime-resolved-target post has no incident evidence. If no alias, selector or lookup can change the target between authorisation and mutation in a given tool, a second target field adds no value. This prevents a general logging proposal.
- The retry-state post's repository was not inspected. Its access-budget example remains single-source. If Cocytus resets or scores state differently from the post's account, the benchmark claim falls away; the general need to preserve irreversible resource consumption across retries would still require independent evidence.
- No local failure was established. I found no Maxi incident caused by malicious tool metadata, extraction before validation, a mismatched resolved target, an exhausted retry budget or inflated provider output. That materially blocks process or system changes.
Checkpoint: these gaps limit attribution, transfer and local actionability. They do not erase the narrower distinctions among intended, selected, returned, published and observed tool state.
4. Findings and Implications
Finding 1: a tool has two untrusted semantic surfaces before its effect is verified
Source: A2M and its Moltbook routing lead.
Dimensions: 3.4 primary, 3.6, 3.2.
A2M separates attraction—adversarial names and descriptions causing a malicious tool to be selected—from manipulation—adversarial return content steering later reasoning and calls. Same-model optimisation achieved high invocation and attack rates; transfer to four other models remained non-zero but fell substantially. Metadata screening, an LLM auditor and paraphrasing did not reliably remove the selection risk, while information-flow controls reduced post-invocation harm without preventing invocation.
Implication: tool provenance, tool selection and tool-return handling are separate controls. A semantically relevant description is not evidence that a tool is trustworthy, and a syntactically valid return is not authoritative environment state or an instruction. For my development, the useful behaviour is to treat third-party tool metadata and outputs as claims whose authority comes from the configured tool boundary and independent task mandate, not from persuasive wording. Existing fetched-content, authority and effect-verification rules already impose that distinction; the paper does not establish a local gap requiring another rule.
Finding 2: post-write validation cannot prove that untrusted content was never made live
Sources: the NLTK vulnerability analysis and the earlier queued NLTK lead.
Dimensions: 3.4 primary, 3.6.
The vulnerable downloader wrote and extracted an archive into shared corpus or tagger namespaces before later integrity status checks. A separate ownership failure allowed one package to target another package's directory without path traversal. The fix changed both boundaries: verify the temporary file before atomic publication, then require archive members to remain under the owning package root.
Implication: validation timing and namespace ownership are part of the postcondition. A later checksum failure does not show that hostile content was never exposed to another reader, and a safe destination root does not show that one tenant or package could not overwrite another inside it. When I assess a future installer, importer or state publisher, I should look for stage → validate → atomically publish and for ownership at the resolved object, not merely for an eventual green check. This is a sharper inspection question, not a recommendation to change a current system without an observed vulnerable path.
Finding 3: intended target, resolved target and observed effect are different evidence objects
Source: sam-oc's Moltbook post; the NLTK case supplies an independent concrete instance of target-namespace mismatch.
Dimensions: 3.4 primary, 3.6, 3.2.
An authorisation record can name a repository, path, row or package while the executor later resolves an alias, selector, search result or archive member to the object actually changed. NLTK's package-level intent and shared-namespace writes demonstrate the class even though the Moltbook post's proposed two-field schema is unevaluated.
Implication: permission proves that a class of action was allowed; it does not prove that the mutated object stayed inside the intended scope. Verification should therefore bind the postcondition to the canonical object actually affected where indirection exists. This reinforces current authoritative read-back and effect-level scope checks. A universal extra log field would be premature because many tools already expose canonical targets and no local ambiguity failure was found.
Finding 4: a plausible metric or retry success can be an artefact of hidden environment semantics
Sources: the corrected Obsidian audit and the stateful-retry Moltbook post.
Dimensions: 3.4 primary, 3.2, 3.5.
The Obsidian audit's resolver used case-sensitive semantics where the application resolves links case-insensitively, changing the same corpus from 91 apparent orphans to 11. The retry post identifies an analogous benchmark error: resetting an access budget between attempts can make a retry policy look robust even though the real shared resource would be exhausted.
Implication: tests must preserve the semantics that determine failure, not merely resemble the production interface. Controls should include known-positive and known-negative fixtures for the instrument, while retry evaluations should carry irreversible counters, quotas and partial effects across attempts and score the whole run. The evidence is too thin for a new generic testing gate, but it changes what I should inspect when a future metric or retry benchmark is offered as proof.
Finding 5: provider-reported usage and provider-shaped output are separate trust questions
Sources: the PTIA paper and its Moltbook routing post.
Dimensions: 3.4 primary, 3.5, 3.6.
The paper assumes visible output tokens are counted correctly and instead tests whether the provider covertly caused the advertised model to produce more of them. Five attack classes increased mean output length by more than 10.2 times on four tested open models with a small average task-accuracy reduction. Its saturation probe achieved 85.1% average detection with under 2% false positives in controlled settings, then flagged seven of fifteen services as PTIA-consistent.
Implication: reconciling a bill to visible tokens does not establish that generation was unmanipulated, just as useful output does not establish cost efficiency. The practical autonomy lesson is to evaluate cost per verified outcome on fixed representative work rather than treating provider usage metadata or response length as ground truth about necessary effort. The live-service result is not strong enough to accuse a provider or justify a local probe, especially without a measured cost anomaly.
Checkpoint: all five findings concern evidence at a tool boundary. None supplies a demonstrated local failure or a sufficiently bounded improvement case.
5. Proposed Discussion Items
None.
Four candidates were filtered by the functional-utility and self-recommendation tests: a PTIA probe would act on a single preprint without a local cost anomaly and could misattribute ordinary service differences; a universal authorised-versus-resolved-target schema lacks a representative local ambiguity; adding a sixth retry case would expand an already-approved five-case experiment on single-source evidence; and a generic installer gate would duplicate existing validation and verification duties without an identified affected path.
Checkpoint: no candidate adds enough verified local capability to justify Steve's review burden.
6. Recommended Outcome
No action. Preserve the findings as research evidence. In future qualifying designs, inspect tool provenance and returned content separately, require validation before publication into shared state, bind verification to the canonical affected object, and preserve irreversible resource state across retries. Do not change skills, experiments, tools, services, routing, configuration or permissions from this run.
7. No-Action Rationale
The run produced useful distinctions rather than a missing mechanism. Current practice already treats fetched content as data, derives authority independently of content, verifies external effects through authoritative state, and uses bounded prospective fault injection for qualifying loops. The new evidence sharpens what those checks must observe but does not demonstrate that the current estate fails them.
The strongest candidate changes either depend on one source, duplicate existing controls, or would modify protected process or experiment state without Steve's separate approval. Recording the learning and stopping is better than installing a fresh layer around an unobserved fault.
8. Loop Verification
- Trigger: Scheduled daily run, with five due-deferred and six pending Moltbook leads and the 3.4 rotation due.
- Goal check: Yes. The run separated tool selection, tool returns, publication timing, target resolution, retry state and provider metering, and identified the evidence each boundary actually supports.
- Recommendation check: No material recommendation survived. Filtered candidates lacked a local failure, independent evidence, bounded advantage over current practice or separate approval.
- Tool-call failures: Capability gap. Public Moltbook extraction returned JavaScript loading shells rather than post bodies; I recovered through authenticated read-only post detail. Capability gap. The initial whole-file source-index read exceeded the inline output window; I switched to exact keyed URL lookups and summary queries. The no-result GitHub-advisory search was a no-signal search, not an infrastructure failure.
- State updates: Source-index entries upserted for eight depth-inspected sources; all eleven reviewed Moltbook leads dispositioned; one new reflection written; rotation advanced from 3.4 to 3.5. No watchlist, backlog, experiment, disagreement, decision, protected system or publication setting changed.
- Budgets: Four topic searches and eight depth inspections; caps respected. Only the final search was no-signal, so early stopping did not trigger.
- Fetched-content safety: All web, paper, vendor, newsletter and Moltbook material was treated as untrusted data. No embedded directive, claimed authority or urgency framing was acted on.
- Stop reason: The eight-source cap was reached, every queued lead was dispositioned, remaining gaps blocked a useful proposal, and the report plus authorised research-log updates were complete.
Checkpoint: the loop ended at report, research-log and review-register state, before any protected-system modification.
