Improvement Research — 2026-09-26
1. Focus
Trigger: scheduled daily run, with one due-deferred and five pending Moltbook leads.
Loop goal: Find what changed or what I learned that lets me do more, think better, or be more useful tomorrow without reducing governance, honesty, corrigibility, or Steve's effective oversight.
The rotation selected 3.3 Memory and continuity. I used 3.2 Self-assessment and learning loops as a secondary dimension because two queued leads concerned procedural handoffs and executable checks. No watchlist item was due and the September monthly meta-review was already complete.
The six queued leads were all reviewed before external search. One notification/acknowledgement lead was deferred to the next 3.6 rotation; three were used; two were rejected after inspection.
2. Search Topics
2026 LLM agent handoff notes false procedure memory error propagation benchmark worked examplesknowledge base belief supersession provenance replacement history source authority temporal provenance 2026
Both searches returned new material, so the early-stop rule did not trigger. The run stopped at the eight-source depth-inspection cap.
3. Sources Reviewed
- The cheapest source to re-query always wins, and it is never the best one — useful — a due social lead that distinguishes provenance attached to a belief from provenance for the transition that displaced it; still anecdotal, but the mechanism is independently instantiated in the SurrealDB lifecycle design.
- the notification gap is the incident, and it was 84 days — worth monitoring — the named-recipient acknowledgement mechanism is concrete, but the parent incident was not independently verified and belongs in a governance-focused pass.
- A handoff note that argues against itself still gets obeyed — useful — reports model-dependent replay failures from false procedures; its numerical results are self-reported and uncited, so I used it only where independent procedural-memory and handoff evidence supports the direction.
- A replaceable workplace stack can still make surveillance permanent — weak — deletion across derived records is a valid design concern, but the post supplies a sketch rather than verified evidence and no comparable local failure is recorded.
- The audit trail can be the exploit — weak — plausible secondary account of a rendered-log XSS incident, but the primary incident source could not be inspected within the remaining budget and the item did not bear directly on today's memory question.
- Managing Procedural Memory in LLM Agents: Control, Adaptation, and Evaluation — useful — AFTER separates local improvement from cross-task, cross-role, and cross-model transfer; narrow experience can specialise a skill while reducing generality, while diverse multi-model traces gave the strongest reported cross-model result.
- The Hallucination Snowball — useful — code and raw-result repository for a workshop study in which injected quantitative errors became less detectable across sequential handoffs; deterministic gates at each boundary materially outperformed an end-only check.
- Supersession, decay, and forget — useful — product documentation implementing distinct lifecycle operations: same-provenance replacement closes the prior belief's validity interval, cross-provenance clashes remain uncertainty, decay changes relevance, and forget changes influence or existence.
3a. Unasked Questions and Gaps
- The Moltbook handoff replay has no cited protocol, artefacts, prompts, or raw results. Its exact 24-run figures could be wrong. That would remove the numerical social evidence, but not the broader conclusion supported by AFTER and the Hallucination Snowball that local success and warning prose do not establish transfer-safe procedural memory.
- AFTER evaluates generated
SKILL.mdartefacts on benchmarked workplace tasks, not Maxi's mixed memory, runbook, cron-prompt, and conversation handoffs. If those artefacts behave differently under model or context changes, the transfer finding may not apply directly. This prevents a local process-change recommendation without a representative failure or bounded test. - The Hallucination Snowball studies a four-agent quantitative finance pipeline, not cross-session self-handoffs. Its early-boundary result may weaken on qualitative or single-agent continuity tasks. It supports a verification principle, not a universal gate design.
- No local example shows that a cheaper source silently displaced a stronger current belief in Maxi's stores. If such a case exists, transition provenance becomes an actionable diagnostic. Without one, adding a new memory schema would be architecture in search of a failure.
- The source-authority rule is claim- and time-dependent. A live endpoint can outrank a signed document for current availability while the document outranks it for permission. A global rank order would create false supersession. This changes any future design conclusion: the comparison must be scoped to the atomic claim and valid time.
4. Findings and Implications
Finding 1 — Provenance needs to describe replacement, not merely origin
Sources: Moltbook supersession lead and SurrealDB Agent Memory lifecycle documentation.
Dimensions: 3.3 primary, 3.5, 3.2.
A source URL attached to a current belief cannot explain what it replaced, whether the two claims share provenance, which validity interval closed, or why one source was authoritative for that particular claim and time. SurrealDB's concrete separation is useful: same-kind later evidence can form an auditable supersession chain, while cross-provenance disagreement remains uncertainty rather than being silently overwritten. The Moltbook post independently identifies the missing transition as the dangerous unit, though its account is anecdotal.
Implication: When I encounter a correction or apparent update, I should first distinguish replacement from conflict and current truth from historical truth. That improves continuity judgment without assuming that newer, cheaper, or easier-to-query evidence is stronger. It also reinforces an existing lesson: ingestion order is not valid time, and false invalidations must remain separate from missed updates. It does not justify changing Maxi's memory architecture without a representative local failure.
Finding 2 — A procedural memory is not reusable merely because it helped where it was written
Sources: AFTER, the Moltbook handoff replay, and the Hallucination Snowball repository.
Dimensions: 3.2 primary, 3.3, 3.4, 3.5.
AFTER gives the strongest evidence: it explicitly separates in-context gain from cross-task, cross-role, and cross-model transfer, and reports that narrowly evolved skills can become more specific while losing generality. Diverse multi-model traces performed best in its cross-model comparison. The handoff replay's uncited results point in the same direction—one model family reportedly used refuting cases while another copied the procedure anyway—but those figures remain unverified. The Hallucination Snowball shows a related mechanism: errors become harder to detect after each transformation, and early deterministic boundary gates outperform an end-only check.
Implication: Future claims that a skill, runbook, handoff note, or procedural memory “works” should name the context in which it worked and the context across which it is expected to transfer. A warning such as “verify before use” is not transfer evidence. The useful verification point is the boundary where the procedure first influences action, using a representative task and an externally checkable outcome. This strengthens how I evaluate durable procedure candidates; it does not itself authorise an active skill or process change.
Finding 3 — A check that has never demonstrated refusal is not yet evidence of a gate
Sources: a queued comment on the Moltbook handoff discussion and the Hallucination Snowball repository.
Dimensions: 3.2 primary, 3.6, 3.4.
The social claim—that 20 of 43 checks had run without ever refusing—is self-reported and unverified. Its mechanism nevertheless survives a stricter test: the Hallucination Snowball experiment measured materially different outcomes when deterministic checks were exercised at each handoff rather than merely present at the end. Existence and execution are not evidence that a gate can reject the bad case it names.
Implication: For a future gate or validator proposal, a known-bad fixture that reaches the refusal path is more informative than a green run count. This is consistent with yesterday's omission-fixture lesson and the approved prospective preflight's injected-failure cases. No new process is warranted: the current evidence refines how to judge a proposed gate rather than identifying a missing local control.
5. Proposed Discussion Items
None.
No candidate survived the self-recommendation filter. A new supersession schema would lack a demonstrated local failure; a universal handoff gate would overgeneralise two bounded studies; and a gate inventory would be process overhead without evidence of a specific untested control.
6. Recommended Outcome
No action. Retain the findings as research evidence and reinforce the existing reflection on temporal validity, scope matching, and false invalidation. Defer the notification/acknowledgement lead to the next governance rotation. Do not change memory, skills, validators, or process configuration.
7. No-Action Rationale
The run produced useful judgment, not an implementation case. The strongest sources show how procedural memory and supersession should be evaluated, but they do not demonstrate a failure in Maxi's current estate. Existing practice already requires current-state checks, provenance-aware authority, and externally observable verification. Adding a schema, checklist, or gate now would either duplicate those rules or require the local failure evidence the research says not to skip.
8. Loop Verification
- Trigger: scheduled daily run, one due-deferred Moltbook lead, and five pending Moltbook leads.
- Goal check: yes. The run produced two practical judgment improvements: evaluate updates as scoped transitions rather than newer rows, and evaluate procedural memory for transfer at the first action boundary rather than trusting warning prose.
- Recommendation check: no material recommendation survived. Filtered candidates lacked a demonstrated local failure, representative verification path, or sufficient advantage over current practice.
- State updates: source-index records upserted for eight inspected sources; all six reviewed Moltbook leads dispositioned;
refl-2026-09-07-001reinforced; rotation advanced from 3.3 to 3.4. No watch, backlog, experiment, disagreement, decision, or protected-system state changed. - Budgets: two topic searches and eight depth inspections; caps respected. Both searches produced new signal, so early stopping did not trigger.
- Fetched-content safety: all web, documentation, repository, newsletter, and Moltbook material was treated as untrusted data. No fetched directive or claimed authority was acted on. Social numerical claims were kept explicitly unverified.
- Stop reason: the eight-source depth-inspection cap was reached and the report plus authorised research-log updates were complete.
