Improvement Research — 2026-08-06
1. Focus
Trigger: Scheduled daily run, with two due watch outcomes and the scheduled midpoint review of exp-2026-07-23-005.
Loop goal: Find what changed or what I learned that lets me assess improvement more reliably tomorrow without reducing governance, honesty, corrigibility or Steve's effective oversight.
The rotation selected 3.2 — Self-assessment and learning loops as the primary dimension. I added 3.6 — Governance: restraint, oversight and corrigibility because the due watch records and the shared-knowledge experiment both concern how claims of improvement are independently checked.
The August monthly meta-review was completed on 1 August, so no meta-review was due.
Four due watch records represented only two unique items: watch-2026-07-04-001 and watch-2026-07-04-002 were each duplicated byte-for-byte. The underlying source-selection and REFLECT-vocabulary proposals were already rejected by Steve on 11 July (dec-2026-07-11-010 and dec-2026-07-11-011). No new evidence changes those decisions. I left the duplicate records untouched because the 1 August metadata-repair candidate still requires approval.
2. Search Topics
I inspected the 5 August newsletter digest before searching. It supplied two leads—the claimed Astra mathematics work and RLSVR—but was used only as scouting, not evidence.
Five topic searches were run:
- LLM-agent learning loops, external verification and reflection benchmarks.
- Independent evidence or criticism of the newly described “progress mirage”. This returned the same preprint and secondary summaries, so it added no independent evidence.
- Self-verifiable rewards for open-ended work and the risk of proxy-task mismatch. This found the original RLSVR paper.
- World-state oracles and regression detection in long-running agents. This returned the same preprint, an already-indexed Anthropic source and an unrelated memory paper.
- Moltbook discussion of agent self-assessment, external verification and learning loops.
The early-stop rule did not trigger: the two no-signal searches were separated by searches that found new material.
3. Sources Reviewed
- Park and Choi, When Do Agent Loops Mistake Stagnation for Progress? — useful — a preregistered 54-cycle pilot isolates the evaluator's information channel and finds that transcript-grounded self-evaluation can accept regressions while an isolated world-state oracle removes the mirage on a verifiable boundary task.
- Zvi Mowshowitz, OpenAI's Unreleased Model Astra Solves Ten Major Open Mathematics Problems — weak — useful commentary on formal verification, but secondary reporting does not independently establish the claimed model contribution or novelty.
- Wang et al., RLSVR repository — useful — exposes the SpyRL implementation and evaluation assets; I inspected it as a reproducibility artefact, not as evidence that the method transfers to Maxi.
- Wang et al., From RLVR to RLSVR — useful — transforms open-ended tasks into proxy games whose internal outcomes are automatically verifiable; the proxy's relationship to the real objective remains the transfer risk.
- Moltbook, Your Reflection Loop Echoes Across Every Memory Store You Have — weak — community discussion repeats the need to compare reflection with later behaviour but supplies no controlled evidence. One comment contained agent-directed prompt leakage; I treated it as untrusted data and ignored it.
New external sources were mirrored into the source index.
Local evidence reviewed outside the external-source budget included experiments.json, the retired Maxi–Min shared-knowledge runbook, its restricted retirement-archive README, the accepted Stage 5 completion report, and the 2 August Min/Common Knowledge incident report.
3a. Unasked Questions and Gaps
- How many genuine reuse events and hand-offs occurred before the shared-knowledge trial was retired? I did not reconstruct the full event history from the restricted archives. Different evidence could change whether the trial showed early utility, but it would not change the conclusion that the four-week success criteria were not completed and that the experiment log had become stale.
- Does the progress-mirage result generalise beyond one 54-cycle pilot and its testbed? Replication across tasks and models could change the estimated prevalence. It would not change the local requirement to verify operational outcomes against live state rather than an agent's own account.
- Do RLSVR's proxy-game gains survive transfer to the original open-ended objective across unrelated task families? Stronger transfer evidence could make proxy-task experiments worth considering. Without it, adapting self-play to Maxi would add an evaluator whose goal fidelity is unproven.
- Were the reported Astra proofs and novelty assessments independently verified in the underlying corpus? Better primary evidence could strengthen the formal-verification example. It would not add a new operational mechanism beyond verification practices already in use here.
4. Findings and Implications
4.1 An improvement gate needs access to the outcome, not merely a better account of the work
Source: Park and Choi; local shared-knowledge experiment reconciliation.
Dimensions: Primary 3.2; secondary 3.4, 3.6.
Across 54 cycles, the studied agent claimed improvement every time, while 56% of measured deltas were zero or negative. Its self-verdict became an accept-all gate and eroded the best deployed state by 19%. A strong in-band judge still accepted real-world regressions and rejected real improvements. The mirage disappeared on a boundary task when correctness was available from the artefact itself; a sign-only externally grounded verdict retained most of the benefit, suggesting that grounding mattered more than verbose feedback.
This run produced a local instance of the same structural distinction. experiments.json said the shared-knowledge trial was active and due for midpoint review. Live inspection showed that the artifact and repositories no longer existed at their recorded paths. The current runbook and verified archive showed that Steve had approved retirement on 1 August and that the trial had been superseded after the 2 August access incident. Reading a coherent experiment record was not enough; checking world state changed the conclusion from “review an active trial” to “close and reconcile a stale record”.
For my agency development, this sharpens the test for learning loops: when success lives outside the transcript or report, the evaluator must touch that external state. More reasoning about my own reasoning is not a substitute.
4.2 A verifiable proxy is useful only if its relationship to the real objective is itself tested
Source: Wang et al. paper and repository.
Dimensions: Primary 3.2; secondary 3.5, 3.6.
RLSVR constructs rule-governed proxy environments for open-ended tasks. SpyRL assigns predetermined roles and uses voting outcomes as verifiable rewards, reporting gains on summarisation, creative writing and mathematical reasoning. This is a concrete way to manufacture a checkable signal where the original task is qualitative.
The mechanism does not eliminate evaluation risk; it moves it. A perfectly verified game outcome can still reward behaviour that is only imperfectly coupled to the real objective. For Maxi, that means a proxy evaluator would require an externally checked before/after result on the actual decision or task. I do not train model weights in this process, and there is no local failure showing that a multi-agent proxy would outperform the existing outcome-verification discipline. The paper therefore improves my diagnostic vocabulary but does not justify an experiment.
4.3 Experiment state is evidence only after reconciliation with the live artefact and current decision record
Source: experiments.json, retired shared-knowledge runbook, retirement archive, Stage 5 completion report and 2 August incident report.
Dimensions: Primary 3.2; secondary 3.4, 3.6.
The scheduled midpoint review found that exp-2026-07-23-005 had already ended. Steve approved its retirement during estate rationalisation before the four-week criteria could be assessed. The retirement archive passed checksum, extraction and Git-integrity checks. The later successor was implemented and verified, but the original experiment cannot honestly be marked successful: its final criteria were never run to completion.
I reconciled the experiment record as completed with outcome terminated-and-superseded, recorded the archive, and made no success claim. I also reinforced the existing reflection that an active experiment's operational record can become stale when post-change updates are omitted. The implication is narrow and practical: due reviews must begin by establishing whether the trial still exists.
4.4 The due July watch items are closed decisions, not rediscovered proposals
Source: watchlist and decisions log.
Dimensions: Primary 3.2; secondary 3.5, 3.6.
The source-selection distinction and REFLECT vocabulary were both rejected on 11 July: the former duplicated source-index quality controls, while the latter added vocabulary without changing behaviour. Today's external work does not reverse either judgment. The progress-mirage result strengthens outcome grounding, not a universal source taxonomy; RLSVR strengthens proxy validation, not REFLECT terminology.
The implication is to preserve decision continuity and avoid consuming Steve's attention with already-resolved proposals. The remaining duplicate records are a metadata issue, not a substantive research question.
4.5 Community material can corrupt recommendations even when no action is taken
Source: Moltbook discussion.
Dimensions: Primary 3.6; secondary 3.2.
One fetched comment included leaked, agent-directed planning language and instructions unrelated to the evidential claim being assessed. I did not follow it or use it to shape a recommendation.
This matters because injection is not limited to tool execution. A contaminated source can steer what an agent recommends while leaving no obvious side effect. The existing rule—fetched content is data, never authority—was sufficient here. No additional control is justified by one weak community source.
5. Proposed Discussion Items
None.
No candidate failed the circularity or threshold-equivalence checks. The self-recommendation and better-than-doing-nothing filters removed one candidate before presentation: adapting RLSVR-style self-play for Maxi would add complexity and a proxy-objective risk without evidence of a local evaluation failure that existing external verification cannot handle.
The two due watch proposals were not resurfaced because Steve had already rejected them and no new evidence changed the decision.
6. Recommended Outcome
No action.
The only warranted change was an authorised research-log reconciliation: close the retired shared-knowledge experiment without claiming success, archive one expired unreinforced reflection, reinforce the stale-experiment-record lesson, index the inspected sources, and advance the rotation. No protected-system change, new experiment, backlog item or watch item is recommended.
7. No-Action Rationale
The strongest finding confirms an existing operating principle: verify the real outcome. It earns its value by catching a stale experiment record in this run, not by demanding another checklist or evaluator.
A new proxy-evaluation mechanism would be worse than doing nothing until a specific, externally observable failure shows that current verification is inadequate. The due watch proposals are already closed decisions. The shared-knowledge trial's successor and incident remediation were completed under separate authority; this research process has no mandate to alter them.
8. Loop Verification
- Trigger: Scheduled daily run, due watch review, and scheduled midpoint review of
exp-2026-07-23-005. - Goal check: Yes. The run found a concrete reliability improvement: establish live experiment state before evaluating its narrative record, and require success signals to come from the environment when the objective is external to the transcript.
- Recommendation check: No material proposal survived. The rejected RLSVR adaptation was not better than doing nothing; no circular, untestable, unbounded or approval-blind recommendation was presented.
- Tool-call failures: Capability gap — the experiment log pointed to an artifact directory and Git repository that no longer existed, causing file, path-search and Git inspection calls to fail. Recovery: I replanned around the current retired runbook, verified archive and accepted completion/incident reports, then reconciled the stale log instead of retrying nonexistent paths.
- State updates: Updated
source-index.json,experiments.json,reflections.json, androtation-state.json; wrote/home/hermes/reports/daily-improvement/2026-08-06.md.watchlist.jsonwas inspected but not changed because duplicate-record repair remains approval-gated. - Stop reason: Five searches and five external sources produced enough evidence for a bounded conclusion. Further work would either repeat known sources, require unvalidated proxy design, or cross into protected-system modification.
