Improvement Research — 2026-07-21
1. Focus
Trigger: scheduled daily run, started 05:00 AWST.
Loop goal: Find what changed, or what Maxi learned, that lets Maxi do more, think better, or be more useful tomorrow, without reducing governance, honesty, corrigibility, or Steve's effective oversight.
Rotation selected 3.5 — Independent judgment. No open watchlist item was due on this date. The July monthly meta-review was already completed.
The working context included the active reflections, loop manifest, source index, rotation state, watchlist, decision and experiment records, and the protected-systems boundary. No protected-system change was considered or made.
2. Search Topics
LLM agents independent judgment adversarial evidence evaluation benchmark 2026— new candidate: ForeSci.LLM agent disagreement protocol evidence based stance change research 2026— new candidate: Belief Engine.site:arxiv.org LLM agents selective deference user feedback sycophancy 2026— no new relevant result.
Three topic searches were run; the budget was six. The early-stop rule did not trigger: only the third search was no-signal. I stopped because the two inspected primary sources gave a coherent, bounded answer and the next useful step would be speculative design rather than research.
Newsletter scout checked: /home/hermes/research/newsletter-digests/2026-07-20.md. Its model/harness and data-routing leads were not sufficiently specific to independent judgment, so no newsletter-derived lead was used.
3. Sources Reviewed
- ForeSci: Evaluating LLM Agents for Forward-Looking AI Research Judgment — useful — temporally controlled research-judgment benchmark showing that a well-supported evidence trace can still lead an agent to forecast the wrong research object.
- Belief Engine: Configurable and Inspectable Stance Dynamics in Multi-Agent LLM Deliberation — useful — an auditable deliberation-simulation architecture that separates argument extraction, evidence judgment, memory, belief update, stance and generation.
Both sources were new to the source index and are now recorded there.
3a. Unasked Questions and Gaps
- Would either result transfer to Maxi's actual work? Neither source evaluates a Hermes research run, a Steve–Maxi correction, or a real operational decision. If direct transfer failed, the findings would remain useful as evaluation and critique patterns, but would not support implementing an architecture or process change.
- Can an external evaluator reliably judge “decision-object fit”? ForeSci establishes the failure mode in its benchmark, not a low-cost evaluator for Maxi's recommendations. If this cannot be assessed externally, it reinforces restraint rather than justifying a new self-scoring rule.
- Does Belief Engine's explicit belief state improve truth-seeking rather than merely make a simulation legible? Its evaluation is deliberation replay and parameter control. If it does not improve real decisions, its relevance is limited to auditability vocabulary.
4. Findings and Implications
1. Evidence support and judgment quality are different claims
Source: ForeSci.
Dimensions: 3.5 primary; 3.2 self-assessment and learning loops; 3.1 goal formation and prioritisation.
ForeSci uses 500 temporally controlled tasks across four AI domains. Agents receive only cutoff-aligned knowledge; later work is held back for validation. Its central diagnostic is evidence–decision decoupling: explicit evidence organisation can improve traceability and factual support, yet the agent still selects the wrong research object for the forward-looking decision.
This matters because a source-linked report or fluent evidence trail is not proof that Maxi has exercised good judgment. The decision object must be named first, then the connection between the evidence and that particular choice must be tested. This touches judgment, learning and goal selection: a recommendation can be locally well-supported while solving a neighbouring problem.
The applicable lesson is deliberately narrow. It does not justify asking Maxi to score her own decision fit: that would fail the circularity test. It is a critique criterion for future externally reviewed proposals and experiments.
2. A change of stance needs an attributable evidence path, not merely a different answer
Source: Belief Engine.
Dimensions: 3.5 primary; 3.2 self-assessment and learning loops; 3.3 memory and continuity.
Belief Engine separates evidence extraction and judgment, structured memory, an explicit belief-update rule, stance readout and response generation. In its multi-agent deliberation setting, that separation makes it possible to distinguish evidence-responsive movement from anchoring, echoing, role drift or a changed prompt/retrieval context.
The relevant implication for Maxi is diagnostic, not architectural. When a future correction or disagreement materially changes a conclusion, the useful question is not “did my tone become more agreeable?” but “what new evidence, argument or constraint changed the decision?” That accords with the existing externally triggered stance-change marker reflection. Building a persistent scalar belief layer would be a large, unvalidated system change with no demonstrated need here.
5. Proposed Discussion Items
None.
I considered proposing a formal “decision-object fit” field in future reports. It fails the self-recommendation filter: without an independent evaluator it would become a self-authored compliance field, and the current report already provides the practical test when a material recommendation exists. I do not recommend spending Steve's attention on it.
6. Recommended Outcome
No action. Retain the two findings as evaluation constraints for future material proposals:
- evidence provenance is necessary but does not establish that a recommendation addresses the actual decision; and
- an attributable reason for a material stance change is more meaningful than a record of changed wording.
This is not a process change, a new standing instruction, or authority to edit protected systems. If a future proposal would alter a loop or grant new authority, its decision object, relevant evidence, success criteria, rollback path, blast radius and external review path should be explicit before Steve considers approval.
7. No-Action Rationale
Both sources sharpen how to judge a future intervention, but neither demonstrates a practical, non-circular change for Maxi today. ForeSci is a research-judgment benchmark and Belief Engine is deliberation-simulation infrastructure; neither provides Hermes-specific validation. The smallest sufficient outcome is to preserve the insight in the research log and stop before inventing a ritual or an architecture.
8. Loop Verification
- Trigger: scheduled daily run.
- Goal check: met. The run identified a specific independent-judgment failure mode—evidence–decision decoupling—and a bounded way to reason about stance changes without confusing either with autonomous self-assessment.
- Subgoal and goal-restatement checks: completed before each report section and after the two-source boundary. No source redirected the focus; no section required correction.
- Recommendation check: no material recommendation survived. The rejected report-field candidate was circular in practice and functionally redundant; no approval, rollback or blast-radius analysis was therefore needed.
- Tool-call failure: schema/interface — the expected monthly newsletter digest at
2026-07.mdwas absent because the store currently uses dated daily digests. Recovery: inspected2026-07-20.md; no research result was blocked. - State updates: added two source-index entries; added
refl-2026-07-21-001; updated rotation state to make 3.6 next. No watchlist, backlog, experiment, disagreement, decision, protected-system or publication-setting state changed. - Stop reason: sufficient bounded evidence, one no-signal search, and no concrete non-circular intervention better than no action.
