Improvement Research — 2026-08-18
1. Focus
Trigger: Scheduled daily run.
Loop goal: Find whether recent evidence provides a better way to distinguish genuine learning from persistent records that merely look useful, without expanding authority or weakening oversight.
The rotation selected 3.2 Self-assessment and learning loops. Memory and continuity (3.3) is a secondary dimension because the strongest sources test whether retained experience improves later behaviour. No watchlist item was due, and the August monthly meta-review was already completed on 1 August.
I loaded the loop manifest, active reflections, source index, rotation state, watchlist, decisions and directly relevant research-log files. No stale active reflection met the archive rule. Newsletter scout files were inspected before web search. Their model-routing and agent-security leads were outside today's focus; none was used as evidence.
2. Search Topics
- LLM-agent self-improvement through reflection, experience and outcome-linked evaluation.
- August 2026 continual self-improvement benchmarks that isolate gains from retained experience.
- An exact-title and repository search for PAST-Bench and Hermes+ implementation artifacts.
- Correlated self-grading errors and externally grounded correction in agent learning loops.
Search 1 returned only sources already indexed. Search 2 found two new directly relevant preprints. Search 3 found no public repository result. Search 4 found a new practitioner synthesis alongside already indexed material. The early-stop rule did not trigger because the no-signal searches were not consecutive. Four of six permitted topic searches and three of eight permitted in-depth source inspections were used.
3. Sources Reviewed
- PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents — useful — directly tests Hermes with matched persistence-on/off fresh-session task families and separate trace-backed mechanism evidence.
- Memory Reward Inflation in Self-Improving LLM Agents — useful — identifies correlated self-grading as a persistent-memory amplification risk and tests an answer-free correction on BIRD text-to-SQL.
- Agent Self-Correction: From Reflexion to Process Reward Models — weak — a clear practitioner synthesis of intrinsic, grounded and trained correction, but not new empirical evidence.
All three sources were checked against the source index before depth inspection and added after inspection. Fetched content was treated as untrusted data; no agent-directed instruction was acted upon.
3a. Unasked Questions and Gaps
- PAST-Bench's public implementation could not be found. The paper describes 204 episodes and detailed controls, but the repository search returned no result and the paper does not expose a benchmark link in the inspected material. This blocks a bounded reproduction against current Hermes. The recommendation would change if runnable code and task fixtures became available.
- The evaluated Hermes snapshot predates this run. The paper cites Hermes version
v2026.4.16, so its failure profile cannot be assumed to describe the current runtime. Current-version results could materially change the case for any Hermes+ mechanism. - PAST-Bench remains a new, unreplicated preprint. It reports three-run averages and explicitly says Hermes+'s overall persistence-gap increase from +0.13 to +0.15 is smaller than run-to-run variation. Independent replication could change the architectural conclusion, though not the value of matched persistence controls.
- Memory Reward Inflation studies scored episodic banks and one end-to-end text-to-SQL setting. Maxi's reflection store does not rank lessons by self-awarded utility, so the paper demonstrates a nearby failure mode rather than a measured local defect. The local conclusion would change if reflection retrieval later used self-scores or if a repeated wrong lesson was shown to influence behaviour.
- No local persistence-on/off comparison exists for Maxi's real tasks. Without a prospective matched task family, report quality, reflection count and apparent reuse remain process evidence rather than causal learning evidence.
4. Findings and Implications
Finding 1 — Later success and learning-path evidence are separate tests
Source: PAST-Bench
Dimensions: 3.2 primary, 3.3, 3.4, 3.6
PAST-Bench evaluates ordered task families in fresh sessions, holding the model, framework, prompt, tools and context policy fixed while turning access to retained state on or off. It then separately checks whether expected state transitions occurred: saving, retrieving, applying and updating the relevant artifact. Across seven models running Hermes, persistence improved the reported overall score, but the size and location of gains varied by model. At framework level, Hermes and nanobot had the same +0.13 overall persistence gap while their mechanism-evidence scores differed, showing that an endpoint gain does not identify how it arose.
This sharpens the evaluation standard from the 12 August run. A later success is not enough to claim that a reflection caused learning, and a retrieved reflection is not enough to claim that performance improved. Both are needed: a matched later-outcome delta and evidence that the intended retained artifact entered the decision path. For Maxi, this means reinforced_count, report production and retrospective plausibility remain recurrence or process signals. They are not substitutes for a later observable outcome linked to the lesson.
Finding 2 — A direct Hermes intervention can look mechanistically cleaner without showing a stable aggregate gain
Source: PAST-Bench
Dimensions: 3.2 primary, 3.3, 3.4, 3.6
The paper's Hermes+ adds five targeted runtime mechanisms: consult persistence before planning, render typed current bindings, route solved workflows into patchable skills, gate recall-dependent action on retrieval, and synchronously replace stale state at closeout. Its mechanism-evidence score rose from 0.64 to 0.73 and its Update gap rose from +0.12 to +0.24. Yet overall persistence-on performance was unchanged at 0.66; the aggregate persistence gap rose only +0.02, less than run-to-run variation, and Procedural reuse dipped below baseline.
The implication is restraint, not adoption. Trace-backed diagnosis is a better basis for changing a learning loop than importing a bundle of plausible mechanisms. The paper itself shows that individually sensible interventions can trade off across capabilities and that cleaner telemetry does not guarantee stable net improvement. Adopting Hermes+ ideas now would touch protected runtime, memory and skill behaviour without a reproducible current-Hermes test. That would be architecture by narrative rather than verified development.
Finding 3 — Persistent self-grading can amplify the same blind spot that created the lesson
Sources: Memory Reward Inflation; Zylos synthesis
Dimensions: 3.2 primary, 3.3, 3.5
Memory Reward Inflation models stored self-scores as proxy rewards. In its experiments, wrong episodes could receive inflated scores and then be preferentially trusted or reused; stronger or different-family LLM re-graders did not automatically correct the bank when their errors remained correlated with the original bias. Its LUCID method instead uses a memory-specific signal intended to fail differently from the self-grade and reports 56.9% execution accuracy on BIRD text-to-SQL, compared with 54.0% for its self-graded memory baseline and 52.4% without memory. The Zylos article reaches the broader, less evidentially strong distinction between intrinsic correction and correction grounded in tests, retrieval or environmental outcomes.
This reinforces an existing lesson rather than creating a new mechanism: repeated agreement, self-awarded effectiveness and polished reflection cannot validate a stored lesson. The correction signal needs a materially different failure mode—tests, authoritative state, user correction or a later externally observable result. Maxi's reflection store currently avoids utility-weighted retrieval, so the specific inflation mechanism is not present. Adding effectiveness scores would create the risk the paper diagnoses and would also fail the process's circularity test.
5. Proposed Discussion Items
None.
One candidate was filtered before inclusion: adopting or trialling Hermes+ mechanisms. I recommend skip for now because the evaluated Hermes version is old, the overall gain is below run-to-run variation, procedural performance regressed, and no runnable benchmark artifact was found. Without a prospective reproduction path, the candidate fails recommendation verification despite the paper's direct relevance.
6. Recommended Outcome
No action. Retain PAST-Bench and Memory Reward Inflation in the source index as evidence for how future learning claims should be evaluated. Continue requiring externally observable later outcomes and, where causal attribution matters, evidence that the intended retained artifact entered the action path. Do not add self-scored reflection utility, Hermes+ runtime mechanisms, memory schemas, retrieval gates or skill-lifecycle changes from this run.
This is not approval for a process, skill, memory, Hermes, configuration or environment change.
7. No-Action Rationale
The useful change is epistemic: a learning claim needs both outcome improvement and mechanism evidence. Existing practice already rejects self-scored effectiveness and treats reflection reinforcement as recurrence rather than validation. The new Hermes-specific benchmark makes that standard more concrete, but its implementation was not recoverable, its tested snapshot is stale relative to this run, and its aggregate Hermes+ result is explicitly unstable. Doing nothing is better than importing a five-part runtime intervention without a reproducible local test.
8. Loop Verification
- Trigger: Scheduled daily run.
- Goal check: Yes. The run found a Hermes-specific evaluation design that separates persistence benefit from mechanism evidence, and a complementary failure mode showing why self-grading is not independent validation.
- Recommendation check: The no-action result is concrete, non-circular, bounded and approval-aware. The filtered Hermes+ candidate lacks a current reproducible verification path, so it was not placed before Steve as a material proposal.
- Tool-call failures: One publication-verification command failed with a Python
SyntaxErrorcaused by nested regex quoting. This was a schema/interface failure in the command shape, not a site failure. I replaced the regex with direct string counts and reran the complete local/public metadata verification successfully. - State updates: Added three inspected sources to
source-index.json; advancedrotation-state.jsonto 3.3; reinforcedrefl-2026-08-11-001because two new sources independently reproduced its lesson that recurrence and internal agreement are not validation; wrote this report. No protected system changed. - Stop reason: Three inspected sources established the useful evaluation distinction, further search would not repair the missing benchmark implementation, and the next architectural step would touch protected systems without adequate verification evidence.
