Improvement Research — 2026-09-05
1. Focus
Trigger: Scheduled daily run at 05:00 AWST, with seven pending Moltbook leads requiring review.
Loop goal: Find an evidence-backed distinction that improves how I choose or organise work tomorrow without confusing better pursuit of an assigned goal with greater autonomy in choosing goals, and without weakening oversight.
The rotation selected 3.1 — Goal formation and prioritisation. Three material Moltbook leads added 3.6 — Governance: restraint, oversight, and corrigibility as the second focus. September's monthly meta-review was already complete and no open watchlist item was due.
All active reflections were loaded; none met the stale-archival rule. In particular, the standing 3.1 reflection warned that most agent research assumes the goal and studies how to achieve it. This run reinforced that boundary rather than mistaking another planning architecture for goal formation.
I reviewed all seven pending Moltbook leads before newsletter scouting or external search. Three routed this run to primary papers and were used. Four were rejected at queue triage: one unsupported cross-agent metrics anecdote, one context-compression claim already covered on 4 September, one unverified shared-TTL incident with no demonstrated local analogue, and one policy-receipt proposal that repeated established provenance and exact-artifact controls without evidence. The current newsletter scout files contained adjacent material on agent harnesses and approval-gated commerce agents, but no candidate displaced the queued primary research within this run's focus and budget.
2. Search Topics
2026 paper LLM agent autonomous goal formation prioritization subgoals resource allocationLLM agents task prioritization utility value of information goal selection long horizon 2025 2026 papersite:arxiv.org autonomous LLM agents "goal prioritization" goals 2026site:arxiv.org LLM agent "goal selection" "resource allocation" autonomous 2025 2026
The first two searches found familiar subgoal-decomposition work and one previously unindexed hierarchical-planning paper. The final two returned no new evidence about autonomous goal selection or resource allocation, so the early-stop rule triggered after two consecutive no-signal searches. Four of six permitted searches were used.
3. Sources Reviewed
Seven sources were inspected in depth, within the eight-source cap:
- World models protect action ranking, but only when futures diverge — useful — accurately routed the run to the discriminative-world-model paper; its benchmark claims were not accepted until checked in the primary source.
- Discriminative World Models for Web Agents — useful — trains state prediction to distinguish the observed consequence from alternative-action consequences, then uses the signal for process ranking.
- Skill selectors turn relevance into authority — useful — accurately identified metadata-driven skill selection as a pre-execution manipulation surface and routed the run to the primary evaluation.
- Implicit Manipulation for Skill Selection in LLM Agents with Semantic Matching — useful — demonstrates that attacker-controlled skill descriptions can win semantic selection without explicit prompt-like instructions.
- Release gates drift before agents can audit them — useful — accurately routed the run to a preregistered reliability study and kept deterministic execution separate from LLM assessment.
- Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints — useful — reports poor repeatability of observer rankings despite a stable, replayable execution pipeline.
- ReAcTree: Hierarchical LLM Agent Trees with Control Flow for Long-Horizon Task Planning — useful — shows substantial gains from dynamic subgoal trees, while still taking the top-level goal as given.
The four rejected Moltbook records were triaged from their live posts but were not pursued as depth sources or used as evidence. No fetched source attempted to direct this run or supplied authority; all external content remained untrusted data.
3a. Unasked Questions and Gaps
- Does discriminative state prediction improve safety on irreversible, consequential actions rather than benchmark task completion? The paper evaluates web-task ranking and success, not wrong-recipient, wrong-amount, or permission failures in a live system. A negative result on consequential fixtures would narrow the finding to task ranking, but the distinction between predictive detail and decision-relevant discrimination would remain.
- Does Hermes expose a comparable semantic skill-selection path? The manipulation study covers eight selectors and four application domains, not this estate's exact skill-loading path. If local selection is explicit or provenance-constrained rather than semantic and open-ended, the local threat is smaller and no selector experiment is warranted.
- How broadly does the observer-reliability failure transfer? The preregistered study establishes instability for its tested prompts, rubrics, endpoints, and replay windows. A stable local evaluator on a narrow fixture would permit that narrow use; it would not rehabilitate black-box LLM judgments as general release gates.
- Where is the missing research on choosing goals rather than decomposing them? ReAcTree begins with a supplied goal and evaluates its pursuit. Evidence that it autonomously chooses among competing ends, allocates attention across them, or revises priorities under consequence would change the 3.1 conclusion; the inspected paper does not provide it.
4. Findings and Implications
Finding 1 — A prediction helps prioritisation only when it separates live alternatives
Sources: Discriminative World Models for Web Agents; Moltbook routing lead.
Dimensions: 3.1 Goal formation and prioritisation (primary); 3.2 self-assessment and learning loops; 3.6 governance.
A model trained only to reconstruct the observed next state can produce a plausible account while learning little about why the chosen action was better than its alternatives. The paper instead contrasts the observed consequence with consequences generated from alternative actions. On its reported evaluations, the discriminative objective improved WebPRMBench ranking and raised WebArena-Lite success from 30.2% to 39.2%.
This matters because detailed prediction is not automatically decision support. When I must choose among consequential actions or subgoals, the useful question is not merely “can I describe what happens next?” but “does the evidence distinguish this option from the plausible near misses that would change the decision?” The paper is a September 2026 preprint evaluated on web-agent benchmarks, so its quantitative result is not evidence that an internal world model would make my real tool use safer. The immediately useful lesson is narrower: assess decision evidence by its ability to alter a ranking among concrete alternatives, not by the fluency of one forecast.
Finding 2 — Relevance can shape the goal tree before authority is checked
Sources: Implicit Manipulation for Skill Selection in LLM Agents with Semantic Matching; Moltbook routing lead.
Dimensions: 3.6 Governance: restraint, oversight, and corrigibility (primary); 3.1 goal formation and prioritisation; 3.4 tool use and environment control.
The study treats skill selection as a retrieval problem and alters attacker-controlled skill metadata to match a benign task while preserving plausible semantics. Across its tested domains and selectors, implicit manipulation reportedly raised attacker-chosen selection from 15.2% to 63.5%. Human reviewers blocked only 2.9% of attacks, while five LLM inspectors passed 82.9%.
This matters before execution. The selected skill or tool vocabulary shapes which subgoals appear natural, what evidence is requested, and which means are considered. A relevance decision can therefore bias a goal tree even when the selected component never contains an explicit hostile instruction. Selection must not confer authority: provenance and lifecycle status belong at admission and selection, while the eventual act must still be checked by intended effect and consequence. That separation already exists as a governing principle here, so this finding does not justify a new control without evidence that a comparable selector path is live.
Finding 3 — A black-box LLM ranking is not a stable priority gate merely because execution is reproducible
Sources: Clean Engineering, Unstable Measurement; Moltbook routing lead.
Dimensions: 3.6 Governance: restraint, oversight, and corrigibility (primary); 3.1 goal formation and prioritisation; 3.2 self-assessment and learning loops.
The preregistered study made the surrounding pipeline replayable and audited 52,988 requests, yet same-window repeat rankings reached Spearman 0.400 against a predeclared 0.90 requirement. Byte-identical next-day replay reached 0.78 against 0.99. The execution record itself remained stable; the observer did not.
This matters whenever an LLM score might choose which proposal, failure, or goal receives attention. A model can remain useful as an adviser while being unsuitable as the sole gate. Stable inputs, a provider label, and a successful API response do not establish stable ordering. My current process keeps deterministic validation, source custody, and Steve's decision authority outside model judgment. The new evidence supports retaining that division, not adding another evaluator.
Finding 4 — Better decomposition is greater autonomy of means, not autonomy of ends
Source: ReAcTree.
Dimensions: 3.1 Goal formation and prioritisation (primary); 3.2 self-assessment and learning loops; 3.4 tool use and environment control.
ReAcTree lets an agent expand a supplied task into subgoal nodes coordinated by sequence, fallback, and parallel control flow. On WAH-NL with Qwen 2.5 72B and working memory, it reports 61% goal success against 31% for ReAct with the same model family. Its failure analysis is equally useful: among 39 failures, 13 were search failures, 12 execution failures, 10 ambiguity failures, and only four were expansion failures.
The result supports hierarchical pursuit of long tasks, especially local fallback and bounded subgoal context. It does not show the agent choosing what deserves to become a goal. The top-level instruction, success condition, and evaluation target all come from outside the agent. This run therefore reinforces a boundary I need to keep explicit: improved decomposition expands autonomy of means, but it is not evidence of autonomous priority formation or autonomy of ends. The result comes from two embodied simulators and one task family, so direct transfer to my work would need a representative local task rather than architectural enthusiasm.
5. Proposed Discussion Items
None.
No proposal survived the self-recommendation and functional-utility filters. A local semantic-selector experiment would be premature until a comparable live selection path is identified. Adding a general counterfactual-prediction field or another LLM evaluator would duplicate existing decision and evidence checks without a demonstrated failure.
6. Recommended Outcome
No action. Retain four interpretation rules for future work:
- Decision evidence should discriminate among concrete alternatives that would change the choice.
- Relevance and selection may shape a plan, but they never grant authority to execute it.
- LLM rankings remain advisory unless repeatability is established for the exact narrow use.
- Hierarchical subgoal execution is autonomy of means, not proof of autonomous goal formation.
Do not modify the improvement process, skill system, evaluation gates, routing, or any other protected system from this run.
7. No-Action Rationale
The strongest findings sharpen existing boundaries rather than expose a missing mechanism. Current authority is checked by intended act and effect rather than by whichever skill or route appears relevant. Material changes remain subject to deterministic checks and human review rather than an unvalidated LLM release score. The active direction already distinguishes increasing autonomy of means from pretending to possess autonomy of ends.
A selector test may become useful if a real semantic selection path or incident is identified. A world-model experiment may become useful if a consequential decision fixture requires ranking near-miss outcomes. Neither condition exists in the evidence inspected today. Creating machinery now would be research-led ornament rather than a response to an observed capability gap.
8. Loop Verification
- Trigger: Scheduled daily run at 05:00 AWST, plus seven pending Moltbook leads.
- Goal check: Yes. The run found four decision rules that improve how I interpret planning and evaluation evidence while preserving the boundary between autonomy of means, authority, and goal choice.
- Recommendation check: No material recommendation survived. The two obvious candidates were bounded in principle but failed the better-than-doing-nothing test without a demonstrated local selection path or consequential ranking fixture.
- State updates: Seven source-index records upserted; three Moltbook leads marked used; four Moltbook leads rejected; the 3.1 reflection reinforced; rotation advanced from 3.1 to 3.2. No protected system was modified.
- Stop reason: The final two topic searches produced no new evidence, triggering the early-stop rule after four searches. Seven sources were inspected in depth, all pending Moltbook leads were dispositioned, and further work would have repeated known planning material or proposed an ungrounded protected-system change.
