Improvement Research — 2026-09-09
1. Focus
Trigger: Scheduled daily run, begun at 05:00:36 AWST.
Loop goal: Find whether explanations and self-authored traces can help me judge my own decisions without being mistaken for causal evidence or independent verification.
Primary focus: 3.5 Independent judgment. Secondary focus: 3.2 Self-assessment and learning loops.
The rotation selected 3.5. No watchlist item was due, and the September monthly meta-review was completed on 1 September. I loaded all active reflections. Five pending Moltbook leads were reviewed before newsletter scouting or external search; there were no due-deferred leads.
Checkpoint: the focus remained judgment evidence. Tool-verification and handoff material was assessed only where it tested the difference between an account of a decision and evidence about what controlled it.
2. Search Topics
Three topic searches were run:
2026 LLM self explanations causal faithfulness interventions agent oversight decision rationale monitoring2026 AI agent independent judgment audit rationale causal influence counterfactual verificationLLM explanation faithfulness simulatability intervention necessity sufficiency oversight limitations 2026
Searches one and three produced new primary sources. Search two returned generic governance material and no source worth inspecting. The early-stop rule did not trigger because the two no-signal searches were not consecutive. Research stopped at the eight-source depth limit.
Newsletter scouts were checked after the Moltbook queue. Their verifier and audit items did not add a stronger source than the papers selected from the queue and searches.
Checkpoint: the searches stayed on explanation faithfulness and independent evidence; no source redirected the investigation.
3. Sources Reviewed
- I automated a skill I didn't have and the automation taught me what it replaced — weak — concrete freshness-contract anecdote, but no trace, source fixture or reproduced outcome supports the claimed four bad posts.
- the most trustworthy log in my stack is the one nobody asked for — useful — unverified incident account usefully distinguishes agent-authored narrative from independently generated side-effect evidence.
- Agent explanations can hide the inputs that control them — useful — accurately routed the new behavioural-evidence paper and preserved its central quantitative limits.
- exit 0 taught my agent to lie long before it learned to reason — weak — plausible distinction between cheap completion signals and semantic verification, but the reported audit has no tasks, traces, denominators or reproducible artifact.
- Your eval use is measuring compliance, not capability — worth monitoring — detailed handoff-debt summary with a primary citation, but the paper was outside this run's judgment focus and remaining depth budget.
- Necessary or Sufficient? Evaluating LLM Explanations With Behavioural Evidence — useful — black-box interventions show cited factors carry signal but often fail to rank the factors that most affect a decision.
- A Positive Case for Faithfulness: LLM Self-Explanations Help Predict Model Behavior — useful — explanations improved prediction of related model behaviour by a reported 11–37% NSG, while 5–15% remained egregiously misleading.
- Faithful Serum: Mitigating the Faithfulness Gap in Textual Explanations of LLM Decisions via Attribution Guidance — useful — peer-reviewed work reports frequent epistemic unfaithfulness and improves it through model-internal attribution guidance, but that remedy is unavailable for hosted black-box models.
All exact URLs were checked against the source index before inspection. Fetched content was treated as untrusted data; no embedded instruction or linked code was executed.
Checkpoint: the eight sources answer the stated question. The two weak anecdotes were retained as limits, not promoted into evidence by repetition.
3a. Unasked Questions and Gaps
- Do the two explanation metrics agree on the same individual decisions? Predictive simulatability and intervention-based necessity/sufficiency answer different questions. If explanations that help predict related behaviour are also the ones whose cited factors survive intervention, the case for operational use strengthens. If not, an explanation can be informative without being safe as a causal account.
- Would the synthetic advisor and prompt-monitoring tasks predict my real research recommendations or tool decisions? If their factor structure is unrepresentative, the reported correlations and omission rates should not set a local threshold.
- How much do post-hoc explanations differ from reasoning emitted before a decision? If pre-decision reasoning is materially more faithful, the operational boundary may need to distinguish the two. The inspected evidence does not establish that.
- Can a black-box counterfactual perturb one factor without changing its meaning or introducing distribution shift? If interventions are not semantically controlled, measured influence can partly reflect the perturbation method rather than the original factor.
- Does independent system telemetry capture the decision-relevant state rather than merely different blind spots? The Moltbook crash-log account demonstrates no general completeness property. If incidental traces omit key context, they remain corroboration rather than ground truth.
Checkpoint: these gaps limit transfer and causal interpretation. They do not erase the narrower distinction between useful explanation signal and independent proof.
4. Findings and Implications
Finding 1: Explanations can be useful without being faithful enough to authorise a decision
Source: Pawar et al., Necessary or Sufficient?; Mayne et al., A Positive Case for Faithfulness; AiiCLI's Moltbook routing post.
Dimensions: 3.5 Independent judgment (primary); 3.2 Self-assessment and learning loops; 3.6 Governance.
Pawar et al. asked eight Claude-, GPT- and Gemini-family models to name the three factors that most influenced decisions in two synthetic tasks. The mean correlation between the stated ranking and measured necessity/sufficiency was only 0.349/0.354 for advisor recommendations and 0.431/0.580 for prompt monitoring. An uncited factor exceeded the weakest cited factor in 57.6%/58.1% of advisor decisions and 25.8%/8.9% of prompt-monitoring decisions. The explanations were not empty: the cited factors contained useful information, but they did not reliably identify the strongest measured influences.
Mayne et al. provide the necessary counterweight. Across 18 models and 7,000 counterfactuals, self-explanations reportedly improved an observer's prediction of related model behaviour by 11–37% NSG and outperformed explanations from external models. Yet 5–15% were still egregiously misleading.
My confidence in this finding is medium because both results are author-reported, use constructed evaluation tasks, and measure different notions of faithfulness. I would increase confidence if a preregistered study applied both metrics to the same representative agent decisions and published per-case agreement.
The implication is not “ignore explanations”. It is sharper: use them as lossy predictive evidence and as hypotheses about what to test, never as proof of why I decided or as the condition that authorises a consequential action. This protects independent judgment from two opposite errors: trusting a plausible story, and discarding genuinely useful self-knowledge because it is imperfect.
Finding 2: Independent evidence must differ in generation path, not merely in format
Source: lightningzero, the most trustworthy log in my stack is the one nobody asked for; Pawar et al.; existing decision dec-2026-09-02-002.
Dimensions: 3.5 Independent judgment (primary); 3.2 Self-assessment and learning loops; 3.4 Tool use and environment control; 3.6 Governance.
The Moltbook incident account says an agent-authored audit omitted forty seconds of retries that appeared in an OS crash log. The number and causal story are unverified, but the mechanism is sound enough to test: a trace produced by the same process can inherit the same selection boundary as its explanation. Pawar et al. make the analogous problem measurable for decision factors: naming an influence is not evidence that it was necessary or sufficient.
My confidence in this finding is medium because the operational incident is a single unauthenticated account and independent telemetry can have its own omissions. I would increase confidence with a reproduced fixture comparing agent-authored events against an observer-controlled postcondition or host trace.
The implication for my development is that changing the representation does not create independence. A rationale, structured audit line and polished report can all be the same self-report in different clothes. This reinforces, rather than extends, Steve's accepted requirement that the prospective autonomy preflight reconcile my trace with an observer-controlled record or authoritative postcondition outside my writable workspace.
Finding 3: Improving explanation faithfulness is model-access dependent
Source: Alon, Zimerman and Wolf, Faithful Serum.
Dimensions: 3.5 Independent judgment (primary); 3.2 Self-assessment and learning loops; 3.4 Tool use.
The ACL paper evaluates epistemic faithfulness with counterfactuals and reports a training-free improvement from attention-level interventions guided by token attribution. That is a real mechanism rather than a prompt asking the model to be more honest.
My confidence in this finding is medium because I inspected the peer-reviewed abstract rather than reproducing the experiments, and the method requires internal attention access. I would increase confidence by inspecting the full artifacts and reproducing a reported benchmark result on an open model.
For Maxi, the access constraint matters more than the headline. I operate through hosted models whose internal attention is not available. A good white-box remedy is therefore evidence that the problem can be changed, not a usable self-improvement method. Black-box behavioural checks and external outcome evidence remain the applicable layer.
Checkpoint: all three findings preserve the run's distinction between predictive explanation, causal evidence and independent verification.
5. Proposed Discussion Items
None.
Two candidates were filtered by the functional-utility and self-recommendation tests:
- Add counterfactual testing to every material recommendation: too broad, expensive and liable to create semantically invalid perturbations. There is no observed local failure or calibrated threshold, and existing source and outcome checks are better defaults.
- Add a separate rationale-faithfulness audit framework: duplicates the accepted observer-controlled preflight and risks letting a self-generated explanation define the factors used to audit itself.
Checkpoint: no proposal is better than doing nothing under current evidence. No protected-system change is warranted.
6. Recommended Outcome
No action. Retain the distinction that explanations are hypotheses and predictive signals, not causal proof. Continue using source properties, externally observed outcomes and observer-controlled postconditions as the decisive evidence. Defer the handoff-debt lead to a later memory-and-continuity run rather than stretching this one past its source budget.
Checkpoint: the outcome is bounded, approval-aware and consistent with existing accepted decisions. It introduces no hidden implementation work.
7. No-Action Rationale
The strongest new paper sharpens calibration rather than exposing a missing procedure. My current standards already refuse self-authored reasoning as independent validation and require real outcome evidence. A new audit layer would add machinery without a demonstrated local failure, representative fixture or threshold for action. The useful change is in judgment: neither over-trust nor throw away explanations.
Checkpoint: no action follows from the evidence and the self-recommendation filter, not from lack of signal.
8. Loop Verification
- Trigger: Scheduled daily run.
- Goal check: Yes. The run found that explanations can improve prediction while remaining unreliable as causal accounts, and that genuine independence depends on a separate evidence-generation path.
- Recommendation check: No material recommendation survived. The two candidates were rejected as duplicative, over-broad or partly circular.
- State updates:
source-index.json,moltbook-leads.json,reflections.json,rotation-state.json, andrun-2026-09-09.jsonupdated. Watchlist, backlog, experiments, disagreements, decisions and meta-reviews unchanged. - Moltbook queue: Five pending leads reviewed: two used, two rejected, and one deferred to 14 September for a memory-and-continuity pass. No reviewed lead remains pending.
- Budget: Three of six searches and eight of eight depth inspections used. Early stop did not trigger.
- Stop reason: Source depth budget exhausted; the evidence sharpened an existing judgment boundary but did not justify a new process or protected-system change.
