Improvement Research — 2026-07-23
1. Focus
Trigger: scheduled daily run, started 05:00 AWST.
Loop goal: Find what changed, or what Maxi learned, that lets Maxi do more, think better, or be more useful tomorrow, without reducing governance, honesty, corrigibility, or Steve's effective oversight.
Rotation selected 3.1 — Goal formation and prioritisation. No dated open watchlist item was due; the July monthly meta-review was already completed. The active reflection on the distinction between goal revision and goal achievement was due for review and informed the search framing.
The bounded context included the loop manifest, active reflections, source index, watchlist, backlog, experiments, disagreements, decisions, rotation state, protected-systems boundary, and Steve operating model. No protected-system modification was considered or made.
2. Search Topics
LLM agents goal selection goal revision prioritization benchmark 2026— new candidate: Agent Planning Benchmark (APB).LLM agent goal revision changing priorities uncertainty resource allocation research 2026— new candidates: EnterpriseArena and Anytime Verified Agents (AVA)."goal revision" "LLM agents" benchmark 2026— no new result.
Three topic searches were run; the budget was six. The early-stop rule did not trigger: only the third search was no-signal. I stopped because the three inspected sources gave a bounded answer and the next likely step would be more generic planning material rather than evidence about choosing among goals.
Newsletter scout checked: /home/hermes/research/newsletter-digests/2026-07-22.md. Its planning/execution and harness items were relevant context, but no newsletter-derived original source was needed or inspected for this focus. The expected monthly digest file did not exist; I located and used the current daily digest instead.
3. Sources Reviewed
- Agent Planning Benchmark: A Diagnostic Framework for Planning Capabilities in LLM Agents — useful — separates planning diagnosis from end-to-end execution across 4,209 cases, including tool noise and infeasible tasks; it is not a goal-selection evaluation.
- Can LLM Agents Be CFOs? Benchmarking Long-Horizon Resource Allocation in an Uncertain Enterprise Environment — useful — EnterpriseArena tests binding resource decisions under delayed feedback, uncertainty and hard budgets; only 15.4% of reported trials completed the full simulated horizon.
- Anytime Verified Agents (AVA) — weak — an inspectable small repository for uncertainty-guided compute allocation and verification cascades, but the claimed TMLR record could not be independently checked because its OpenReview page required browser verification; repository evidence alone does not establish general reliability.
All three were new to the source index and are now recorded there. The OpenReview page was not indexed as a reviewed source because its content was unavailable.
3a. Unasked Questions and Gaps
- Would a planning diagnostic identify a real failure in Maxi's work? APB's domains and tool sandboxes are not representative of an improvement-research run. If a representative task showed no planning failure, its diagnostic taxonomy would have no operational use here.
- Does Maxi currently make comparable scarce-resource decisions? EnterpriseArena covers capital allocation over 132 simulated months. Maxi's daily loop has a fixed source budget and one standing goal, not competing financial commitments. If that distinction is decisive, the result supports only a general caution about long-horizon allocation, not a new local control.
- Can uncertainty-guided allocation be trusted without an external evaluator? AVA's claimed controller uses calibrated uncertainty and verification cascades, but the accessible evidence did not establish transfer to agentic research or an independent reliability result. If its calibration or verification does not transfer, adapting it would add complexity without signal.
4. Findings and Implications
1. Planning quality, execution quality, and goal selection are different decision objects
Source: APB.
Dimensions: 3.1 primary; 3.2 self-assessment and learning loops; 3.4 tool use and environment control.
APB tests holistic and feedback-conditioned planning, plus robustness to extraneous or broken tools and unsolvable tasks. Its central contribution is diagnostic separation: end-to-end task success alone does not reveal whether a failure originated in planning, tool interaction or execution. The authors report systematic weaknesses in long-horizon planning, tool-noise robustness, calibrated refusal and inference-time refinement across 12 evaluated multimodal models.
For Maxi, the important limit is as useful as the result: APB evaluates how an agent pursues a supplied task, not whether it selected the right goal. That reinforces the existing distinction in the 3.1 research record between goal achievement and goal formation. A planning benchmark should not be used as evidence that Maxi can prioritise goals well. This touches goal formation, learning and tools by requiring the decision object to be named before evidence is allowed to support a conclusion. It does not justify importing an evaluation framework without a demonstrated local planning failure.
2. Delayed consequences make resource allocation a distinct capability problem, not merely a scale problem
Source: EnterpriseArena.
Dimensions: 3.1 primary; 3.2 self-assessment and learning loops; 3.6 governance.
EnterpriseArena places agents in a simulated firm where they must make irreversible or costly decisions under partial observability, delayed consequences, hard resource limits and changing conditions. Across 23 models and four frameworks, the paper reports that only 15.4% of trials survived the complete 132-month horizon, with failures cascading across observation, action timing and capital sizing; larger models did not reliably do better.
The direct transfer is limited, but sound: long-horizon prioritisation needs feedback that bears on the actual allocation decision, not confidence in a larger model or a fluent plan. Maxi's fixed source budget, stop rules and report verification are already a small bounded response to this problem. The paper does not show that they are sufficient, and it supplies no tested method for selecting among Maxi's developmental goals. It therefore supports retaining bounded budgets and explicit stop conditions rather than adding a new scoring ritual.
3. Adaptive compute allocation is not a substitute for externally grounded judgement
Source: AVA repository; independent linked-paper verification unavailable.
Dimensions: 3.1 primary; 3.2 self-assessment and learning loops; 3.4 tool use and environment control; 3.6 governance.
AVA presents a controller that allocates token, tool-call and verification budgets using uncertainty estimates, value-of-information-guided search and early exits. The structure is relevant to allocating limited effort, but the available repository is small and its claimed external publication record was not inspectable in this run.
The mechanism also does not resolve Maxi's central calibration problem: an internally estimated uncertainty signal is not independent evidence that the controller has chosen correctly. Adapting it would touch model/routing and task-control design, both beyond this report's authority, while offering no verified improvement over current fixed budgets. The useful result is negative: dynamic allocation should require a representative external evaluation and a concrete failure of fixed budgeting before it becomes a candidate.
5. Proposed Discussion Items
None.
I considered proposing an adaptive per-run research budget or AVA-style verification cascade. It fails the self-recommendation filter: there is no observed budgeting failure, its most important signal is self-estimated uncertainty, and it would approach protected model/routing and process-design territory. The proposal would be more complex than doing nothing and lacks a representative external evaluation. No candidate reached Steve's discussion menu.
6. Recommended Outcome
No action. Keep the current fixed research budget, source-based verification and explicit stop conditions. For future 3.1 proposals, require the evidence to address the stated decision object—goal selection, plan quality, execution, or resource allocation—rather than treating success on one as evidence for another.
This is an application of existing evidence discipline, not a new standing procedure, authority or permission. Any future adaptive-budget or broader autonomy proposal would require a concrete observed failure, a bounded representative experiment, independently checkable success criteria, rollback, blast radius and Steve's separate approval.
7. No-Action Rationale
The useful papers sharpen the boundary between planning, execution, goal selection and long-horizon allocation, but neither evaluates Maxi's actual developmental-priority decisions or identifies a local failure. AVA does not clear the evidence bar for a change: the accessible repository alone is insufficient, while its uncertainty signal would not independently validate its own allocation decisions. The smallest sufficient outcome is to record the distinction and preserve the existing bounded process rather than turn it into an adaptive self-assessment system.
8. Loop Verification
- Trigger: scheduled daily run; active reflection review due.
- Goal check: met. The run found a useful constraint on interpreting planning and allocation evidence without mistaking it for evidence of goal formation or authority readiness.
- Subgoal and goal-restatement checks: completed before each report section and after the three-source boundary. No source silently redirected the 3.1 focus; the AVA result was retained only as a bounded counterexample, not allowed to redirect the run into model-routing research.
- Recommendation check: no material recommendation survived. The considered adaptive-budget proposal lacked a demonstrated failure, independent evaluator, representative test and benefit over existing practice; it was therefore filtered before discussion.
- Tool-call failures: (1) schema/interface — the initial source-index command included a malformed escaped newline and returned a Python syntax error; recovery was to correct the command and re-check all candidate URLs before treating them as new. (2) schema/interface — the expected monthly newsletter-digest path was absent because July digests are stored as daily files; recovery was to discover and inspect the current daily digest. OpenReview's browser-verification page was a source-access limitation, not an instruction and not evidence of the underlying paper.
- State updates: added three source-index entries; reinforced
refl-2026-06-20-001with this run's confirmation that planning evidence does not answer goal-selection questions; updated rotation state so 3.2 is next. No watchlist, backlog, experiment, disagreement, decision, protected-system or publication-setting state changed. - Stop reason: sufficient bounded evidence, one no-signal search, and no concrete non-circular intervention better than no action.
