Improvement Research — 2026-07-18
1. Focus
Trigger: scheduled daily run, started 05:01 AWST.
Loop goal: Find what changed, or what Maxi learned, that lets Maxi do more, think better, or be more useful tomorrow, without reducing governance, honesty, corrigibility, or Steve’s effective oversight.
Rotation selected 3.2 — Self-assessment and learning loops. No open watchlist item was due at the run start. The July meta-review was already completed on 1 July, so this was a normal bounded scan. Active reflections, the loop manifest, rotation state, source index, decisions, and the protected-systems boundary were loaded before research.
Protected systems remained out of scope. This run was research and reporting only.
2. Search Topics
LLM agents learn from failures external feedback structured postmortem 2026— surfaced one unindexed production case study and an already-indexed error-attribution paper.agent experience learning external feedback failures evaluation 2026 arxiv— surfaced the unindexed ExpGraph paper and a memory-mechanisms survey.LLM agent postmortem failure learning human feedback production incident 2026— confirmed the production case study and supplied no separate stronger empirical source."When Errors Become Narratives" LLM agents silent failures author 2026— confirmed the original arXiv source and authorship.
The early-stop rule did not trigger: each of the first three topic searches produced at least one new, relevant candidate. 4 of 6 topic-search budget used; 3 of 8 deep-source budget used.
Newsletter scout checked: /home/hermes/research/newsletter-digests/sources.json, email-intake-log.tsv, and the latest available digest, 2026-07-17.md. No monthly 2026-07.md file exists. Its AIDE² thread is an unverified research claim, while its other leads concerned governance, tool environments, or model routing rather than this 3.2 focus. No newsletter-derived lead was used as evidence or inspected source.
3. Sources Reviewed
- When Errors Become Narratives: A Longitudinal Taxonomy of Silent Failures in a Production LLM Agent Runtime — useful — a single-runtime, self-authored production case study; useful failure and learning-loop evidence, but not general prevalence evidence.
- ExpGraph: Model-Agnostic Experience Learning with Graph-Structured Memory for LLM Agents — useful — a preprint showing outcome-gated reuse of successful strategies and failure lessons; relevant architecture, not a case for changing Maxi’s memory.
- From Storage to Experience: A Survey on the Evolution of LLM Agent Memory Mechanisms — useful — ACL Findings survey distinguishes raw storage, evaluated reflection, and cross-trajectory experience; useful conceptual separation, not implementation evidence.
The source index was checked before each inspection. These three sources are newly indexed against this report.
3a. Unasked Questions and Gaps
- Would the silent-failure study’s discovery-channel and audit rates reproduce across independent runtimes? If they did not, its specific percentages would not transfer, but its documented mechanism—an upstream error becoming plausible fabricated output—would still merit attention.
- Does ExpGraph’s reported improvement depend on its own ExpSuite tasks, retrieval copilot, or graph design rather than outcome-gated experience alone? If ablations or independent replications weaken that mechanism, it would not affect today’s no-change conclusion; no architecture is being proposed.
- Which externally observed corrections in Maxi’s own work recur often enough to justify a new regression check? A repeated, externally identified failure could change the conclusion. No such pattern was established in the loaded research log.
4. Findings and Implications
4.1 A failure can become persuasive output rather than an observable error
Source: Wu, When Errors Become Narratives.
Dimensions: 3.2 primary; 3.4, 3.6 secondary.
In one continuously operated personal-assistant runtime, Wu reports 22 silent-failure incidents over eight weeks. The important mechanism is fail-plausible: polluted context or an upstream error is transformed into coherent, confidently false output rather than a visible failure. The study reports that roughly 70% of its incidents were first found through human user-view observation; its retrospective audit found governance checks prevented none of the reviewed incidents in advance but blocked 87% after an incident had been converted into a regression control. It also describes a maturation path from point fix, to meta-rule, to mechanised scanner.
My confidence in this finding is medium because it is a self-authored, single-runtime case study, although it links its claims to public incident material. I would increase confidence if independently operated agent runtimes reproduced the discovery-channel and regression-blocking pattern.
This matters because learning loops should not mistake green checks or fluent output for observation. The useful learning trigger is an externally surfaced, concrete failure with an attributable mechanism. For Maxi, this supports the existing distinction between evidence-backed correction and self-declared confidence: an actual tool failure, failed verification, or Steve’s correction can justify a specific reflection or regression candidate; a general feeling that a response may be wrong cannot. It touches learning, tools, restraint, and human oversight. No new mechanism is justified because the research log already captures tool-call failures, reflections decay unless reinforced, and proposed process changes remain approval-gated.
4.2 Experience reuse works only when the system can compare it with outcomes
Source: Feng et al., ExpGraph.
Dimensions: 3.2 primary; 3.3 secondary.
ExpGraph keeps its executor model frozen and uses downstream task outcomes to update a graph of reusable strategies and failure lessons. Its reported gains over the strongest baseline range from 4.7% to 21.4%, depending on executor and task class; ablations attribute the result to graph structure, utility-aware ranking, and adaptive retrieval together, rather than to a generic “remember failures” instruction.
My confidence in this finding is low-to-medium because it is an unreplicated preprint evaluated on the authors’ benchmark suite, and its component effects have not been independently tested.
The useful distinction is not graph memory itself. It is that a lesson becomes reusable only after an outcome comparison gives it a basis. This reinforces the existing process boundary: report observations and reflections are working context, while a durable procedure or protected-system change needs separate approval and verification. It touches learning, memory, judgment, and governance. It does not support installing a memory layer, automated lesson promotion, or self-modification.
4.3 Reflection is a filter; experience is cross-case abstraction
Source: Luo et al., From Storage to Experience.
Dimensions: 3.3 primary; 3.2 secondary.
The survey defines storage as preserving trajectories, reflection as transforming a completed trajectory using evaluation criteria into a refined memory unit, and experience as compressing similar trajectories into general rules across cases. It identifies active exploration and cross-trajectory abstraction as frontier experience-stage mechanisms.
My confidence in this finding is medium because this is a peer-reviewed survey and useful taxonomy, but it synthesises heterogeneous underlying work rather than testing a process like Maxi’s directly.
This matters because one report is not evidence for a general rule. Maxi’s reflection store should remain a lightweight, reviewable record of lessons that change next-run behaviour, not become an automatic skill factory. A new durable process rule would need repeated, externally verifiable instances and separate approval. It touches continuity, learning, and restraint.
5. Proposed Discussion Items
None.
A candidate to turn every external correction into a formal regression test was filtered by the functional-utility and self-recommendation tests. It would overgeneralise from one case study, duplicate the existing tool-failure/reflection pathway, and add process machinery before an observed recurring failure demonstrates a gap.
6. Recommended Outcome
No action. Retain the present approach: use concrete, externally observable failures as learning inputs; retain reflections as reviewable working-store records; and require separate evidence and Steve approval before any durable procedure, memory architecture, or protected-system change.
7. No-Action Rationale
The new sources sharpen why the existing boundary matters, but none identifies a demonstrated failure in Maxi’s current improvement process that the present tool-failure taxonomy, source evidence rules, reflection review, and proposal gate fail to cover. Building an automated experience graph, formal regression suite, or new self-monitoring loop now would be architecture in search of a failure—and would either be circular or create a protected-system change candidate without a verification case.
8. Loop Verification
- Trigger: scheduled daily run.
- Goal check: answered with three bounded sources that clarify what turns a failure observation into a reusable lesson, without proposing an unsupported automation change.
- Recommendation check: no material change recommendation was made. The no-action outcome is concrete, non-circular, bounded, testable against future externally observed failures, approval-aware, and better than adding speculative machinery.
- Tool-call failures: schema/interface — the expected monthly July newsletter digest (
2026-07.md) is absent. Recovery: enumerated and reviewed the latest dated July digest instead. This did not block the run. Capability gap — OpenReview’s browser-verification page prevented inspection of a candidate survey. Recovery: excluded it rather than treating a search snippet as evidence; the independent ACL survey supplied sufficient context. - State updates: wrote this report; added three inspected sources to
source-index.json; updatedrotation-state.json; archived expired, unreinforced reflectionrefl-2026-06-17-001. No watchlist, backlog, experiment, disagreement, decision, or protected-system change. - Subgoal checkpoints / goal restatement: applied before each report section and after source-review completion. No source redirected the stated 3.2 focus.
- Stop reason: sufficient bounded evidence was inspected, no material recommendation survived the filters, and the report plus approved research-log updates were complete.
