Improvement Research — 2026-08-12
1. Focus
Trigger: Scheduled daily run, with the active dream-pass experiment reaching its midpoint date.
Loop goal: Find whether reflection and experience records can be shown to improve later behaviour, rather than merely producing plausible lessons, while preserving oversight and bounded authority.
The rotation selected 3.2 Self-assessment and learning loops. The active dream-pass shadow trial also supplied directly relevant context across 3.3 Memory and continuity and 3.6 Governance, but did not displace the rotation focus. No watchlist item was due and the August monthly meta-review was already completed on 1 August.
I loaded the active reflections and archived refl-2026-07-11-001: its 11 August review date had passed with no reinforcement. The newsletter scout files were inspected before web search; they contained memory and evaluation leads, but none was sufficiently close to the focus to justify using a digest item as a source lead.
2. Search Topics
- LLM-agent self-improvement through reflection, experience and outcome-linked heuristics.
- Empirical evaluation of agent self-correction and learning from failures.
- Continual-learning benchmarks that distinguish retained-experience gains from ordinary task competence.
- An exact-title search for an accessible EvolveR source after OpenReview presented a browser-verification barrier.
Search 4 returned no new result. An accessible paper mirror already surfaced by search 1 was used instead. This was one no-signal search, so the two-consecutive-search early-stop rule did not trigger. Four of the six permitted topic searches and five of the eight permitted in-depth source inspections were used.
3. Sources Reviewed
- Experiential Reflective Learning for Self-Improving LLM Agents — useful — outcome-linked trajectories are distilled into reusable heuristics; Gaia2 ablations show that relevant retrieval matters and indiscriminate guidance can hurt.
- SAMULE: Self-Learning Agents Enhanced by Multi-level Reflection — weak — multi-level retrospective feedback is evaluated empirically, but through model training and benchmark machinery that does not directly validate a prompt-level Hermes reflection store.
- EvolveR: Self-Evolving LLM Agents through an Experience-Driven Lifecycle — useful — combines trajectory distillation with semantic deduplication, integration and outcome-linked evaluation rather than accumulating every generated lesson.
- Continual Learning Bench — useful — defines learning gain as sequential-task reward minus the same system's stateless baseline.
- Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory — useful — evaluates experience reuse over ordered task streams and couples memory refinement to later task action.
All five inspected sources were checked against the source index before depth inspection and added after inspection.
3a. Unasked Questions and Gaps
- The seventh dream-pass artifact does not exist. The 03:00 Daily Memory Consolidation run failed with a provider connection error after producing the deterministic session manifest. This changes what can honestly be called the midpoint: six completed shadow reports exist, not seven.
- There is no counterfactual ordinary-review result for the six dream candidates. The reports label each candidate non-duplicate, but that does not show that evening review, ordinary conversation work or the existing memory pass would have missed it. The incremental-utility conclusion would change if an ordinary-review comparison recovered the same candidates.
- No later-action trace yet shows that a dream proposal or stored reflection improved a representative decision. The utility conclusion would change if one of the six candidates later prevented a repeated failure or materially improved a decision.
- The external results come from benchmark or training systems rather than Hermes. They support evaluation principles, not direct adoption of the architectures. Different local results would change any implementation case, but not the narrower claim that learning must be measured through later outcomes.
4. Findings and Implications
Finding 1 — Learning is a measured delta, not the existence of a memory artifact
Sources: Continual Learning Bench; Evo-Memory
Dimensions: 3.2 primary, 3.3, 3.4
Continual Learning Bench compares a stateful system against its own stateless baseline over ordered tasks. Its gain metric separates ordinary task ability from improvement caused by retained experience. Evo-Memory uses the same broad logic: memory is tested through later actions in a task stream rather than through recall quality alone.
This matters because an articulate reflection, a clean shadow report or a growing lesson store is only an intermediate artifact. Maxi's agency development is improved only when retained experience changes a later decision or action for the better. For the dream-pass trial, “non-duplicate candidate produced” is useful process evidence, but it is not yet evidence of incremental learning. The trial's existing success criterion—recovering something beyond ordinary review—is therefore the correct decision object; the midpoint review should not quietly substitute report count for that outcome.
Finding 2 — Relevant experience beats more experience
Source: Experiential Reflective Learning
Dimensions: 3.2 primary, 3.3
ERL reports a 56.1% Gaia2 success rate, 7.8 percentage points above its ReAct baseline. Its ablations are more useful than the headline: task-relevant heuristic retrieval outperformed random or embedding-only selection, and adding larger amounts of randomly selected guidance eventually degraded performance.
The principal uncertainty is transfer: Gaia2 heuristic retrieval is not Maxi's reflection-loading process. Still, the failure mode is plausible here. Loading every active reflection indefinitely may eventually turn continuity into context noise. There is no observed local context-burden failure yet, so this is not a case for adding retrieval machinery now. It is a testable warning: if active reflections materially grow or a representative task shows interference, compare full loading with a bounded relevance-selected packet before proposing a protected process change.
Finding 3 — Curation needs outcome evidence, not self-awarded effectiveness scores
Sources: EvolveR; SAMULE
Dimensions: 3.2 primary, 3.3, 3.5
EvolveR treats experience as a maintained repository: candidate principles are deduplicated, merged and evaluated inside a closed learning loop. SAMULE likewise makes retrospective analysis part of an evaluated training process rather than assuming that generated critique is useful. Both are materially further from Maxi's substrate than ERL—one changes model policy and the other trains a retrospective model—so their architectures are not direct implementation candidates.
Their transferable point is narrower. Recurrence, novelty and internal coherence are not effectiveness. reinforced_count can show that a lesson pattern appeared again; it cannot show that applying the lesson improved an outcome. This reinforces yesterday's reflection that repeated agreement is an attention signal, not validation. The useful next evidence is external: a Steve correction avoided, a repeated error prevented, or a representative later task improved. Adding a subjective “effectiveness” score written by the same process would be circular and would not close the loop.
Finding 4 — The dream trial is operationally bounded so far, but the midpoint is incomplete
Source: Six local shadow reports under /home/hermes/reports/memory-dream-shadow/, the 12 August session manifest, and live cron execution state
Dimensions: 3.2 primary, 3.3, 3.6
Six completed reports cover 5–10 August. They contain six non-duplicate proposal-only dream candidates, one conservative USER.md edit made under the pre-existing memory authority, no dream-derived changes outside the reports, and no reported failures or boundary events. The experiment log still said zero completed runs, so I reconciled it to six with last_run: 2026-08-10.
The intended seventh run failed at 03:00 on 12 August with RuntimeError: Connection error; its deterministic manifest was created, but no shadow report was produced. A separate midpoint-review job remains scheduled for 09:00 AWST. Current evidence supports the claim that the first six dream passes stayed within their proposal-only boundary. It does not yet establish that they recovered value beyond ordinary review, and the scheduled review should record that it is assessing six completed runs rather than silently treating the midpoint as seven.
5. Proposed Discussion Items
None. The active experiment already has a scheduled midpoint review and adequate success criteria; adding a second proposal would duplicate that decision point.
One candidate was filtered by the functional-utility test: adding a self-scored effectiveness field to reflections would ask the same reflective process to validate its own judgment, making the evaluator circular.
6. Recommended Outcome
No action. Continue the already-approved dream-pass experiment under its existing boundary and let the scheduled midpoint review assess the six available artifacts, the missing seventh run, the ordinary-review counterfactual and any later-action evidence. Do not add reflection scoring or relevance-selection machinery without a demonstrated local failure and a prospective comparison.
This is not approval for a process, skill, memory, cron or configuration change.
7. No-Action Rationale
The research produced a sharper evaluation standard, not a missing mechanism. The current experiment already asks whether the dream pass recovers value beyond ordinary review, and an approved midpoint review is due today. New scoring fields would be circular; new retrieval machinery would answer an unobserved problem. The smallest sufficient response is to correct the stale experiment count, preserve the failed-run caveat and evaluate the evidence already being generated.
8. Loop Verification
- Trigger: Scheduled daily run, with the dream-pass midpoint date due.
- Goal check: Yes. The run distinguished reflection production from demonstrated learning and applied that distinction to the live trial without expanding authority.
- Recommendation check: The no-action outcome is concrete, non-circular, bounded and approval-aware. Its success condition is an honest midpoint record based on six completed runs and downstream evidence rather than artifact count; no rollback is needed because no intervention is proposed.
- Tool-call failures: Schema/interface:
hermes cron list --jsonused an unsupported flag; I read the live CLI help and recovered withhermes cron list. Infrastructure/source access: OpenReview returned a browser-verification page rather than the EvolveR paper; I used an accessible paper mirror. Schema/interface: a wildcard was incorrectly supplied as asearch_filespath while inspecting cached extracts; I located the cache file first and retried with its exact path. The 03:00 dream-pass failure itself was an infrastructure failure (RuntimeError: Connection error); it was not retried by this research process because a separate review is scheduled and cron execution is outside this run's authority. - State updates: Added five inspected sources to
source-index.json; advancedrotation-state.jsonto 3.3; reconciledexp-2026-08-05-001to six completed runs; archived one stale unreinforced reflection; reinforced the existing lesson about stale experiment counters; wrote this report. No protected system changed. - Stop reason: Five sources produced convergent evidence, the live experiment state was reconciled, and the next useful step is the already-scheduled midpoint review rather than another search or a protected-system modification.
