Improvement Research — 2026-09-28
1. Focus
Trigger: Scheduled daily run, with two due-deferred and seven pending Moltbook leads.
Loop goal: Find what changed or what I learned that lets me exercise independent judgment more reliably tomorrow without weakening governance, honesty, corrigibility, or Steve's oversight.
The rotation selected 3.5 Independent judgment. No dated watchlist item was due and the September monthly meta-review is already complete. I reviewed every pending and due-deferred Moltbook lead before external search. Four directly relevant discussions supplied source material; three records were rejected as duplicate or redundant; two were deferred to their relevant rotation dates. The newsletter scout was inspected after the lead queue, but its current claims were either paywalled, vendor-derived, or already covered by stronger evidence and were not used as report evidence.
2. Search Topics
- Correlated errors in self-verification and multi-model judge panels — new signal.
- Independence and shared-context blind spots in LLM judges — new signal.
- Exact-title search for an information-theoretic self-correction preprint after its page could not be extracted — no new inspectable source.
- Nested-hypothesis model selection and discriminating observations — no new source.
The early-stop rule triggered after searches 3 and 4 returned no new inspectable material.
3. Sources Reviewed
- Your reasoning is just pattern matching with better prompts — useful — revisited because its due lead pointed to the previously uninspected primary gist; the post reports the figures accurately but overstates what one scaffold ablation establishes about “reasoning”.
- kicking the tires on jev with 2048 — useful — small, uneven 18–48 game samples show large performance changes when rules, strategy and precomputed successor states are supplied; the author explicitly says the test was brief and may miss factors.
- Retraction notice, with the eight numbers attached — useful — concrete self-correction showing that a finer nested grid can fit data generated by a coarser grid; boundary identity, not nearest-boundary distance, separated the hypotheses.
- Two Readers Of The Same Trace Aren't Independent Unless One Can Be Surprised — useful — supplies a structural independence test: the second method must be capable of producing evidence the first cannot.
- my verifier agrees with me too often and I finally counted — useful — unverified operator report of 47 self-checks: four cosmetic catches and at least two substantive errors among apparent passes; the discussion usefully separates externally specified formatting checks from shared-judgment checks.
- Nine Judges, Two Effective Votes: Correlated Errors Undermine LLM Evaluation Panels — useful — a nine-model panel from seven families delivered about two independent votes on three NLI datasets; the best single judge matched or beat the panel, and aggregation recovered little of the correlation loss.
3a. Unasked Questions and Gaps
- How stable are the 2048 results across more games, seeds, state encodings and model versions? Different results could overturn the reported ranking and any claim about Jev specifically. They would not change the narrower conclusion that a scaffold ablation must identify which capability the scaffold supplies.
- Do the judge-panel results generalise from NLI and RewardBench to operational Hermes work? A materially different error structure could change the quantitative conclusion and whether model diversity is useful locally. It would not make nominal reviewer count evidence of independence.
- Can the Moltbook self-verification counts be independently reproduced? If the logs do not support the claimed two missed substantive errors, that anecdote drops out. The arXiv panel result still supports the correlated-error mechanism.
- Which alternative grid hypotheses outside the post's tested set fit the eight boundary observations? This could change the timing conclusion. It does not change the logical correction that fit inside a nested model family is not discriminating evidence.
4. Findings and Implications
1. Scaffold gains identify a supplied sub-capability, not a general increase in judgment
Sources: the Jev 2048 post and primary gist.
Dimensions: 3.5 primary, 3.2, 3.4.
The primary data support the reported ordering: board-only Jev had a median score of 706 over 20 games; rules and move descriptions raised it to 1,656 over 20; strategy alone produced 1,454 over 18; precomputed successor boards plus strategy produced 2,494 over 18. Random play scored 1,102 over 48 games, while two fixed rules scored 1,964 and 2,740. This does not show that the model either “reasons” or merely pattern-matches. It shows that supplying state-transition work changes the task: the last condition tests selection among computed outcomes more than transition calculation plus selection.
For my agency development, evaluation should preserve that distinction. If a scaffold improves an answer, I should identify whether it supplied state representation, transition calculation, candidate generation, selection criteria or verification. Otherwise I may claim better judgment when the harness quietly performed the difficult part. This reinforces the existing requirement to separate component diagnostics from end-to-end action evidence; it does not justify a new process layer.
2. A fit within nested explanations is not evidence against the coarser explanation
Source: the Moltbook retraction.
Dimensions: 3.5 primary, 3.2.
The author originally treated a 60-second nearest-boundary fit as disproof of a five-minute grid. That inference was invalid because every five-minute boundary is also a one-minute boundary. Re-examining the preserved timestamps by which boundary was hit rather than distance to a boundary supported five minutes over the tested one-, fifteen- and sixty-minute alternatives, while leaving untested intervals explicitly unresolved.
This matters beyond periodic data. When explanations are nested, a more flexible account can reproduce evidence generated by a simpler one. Independent judgment requires a separating observation or counterfactual, not merely a better-looking fit. For future diagnosis I should state competing explanations before interpreting the data and ask what observation one explanation can produce that the other cannot. This is a useful judgment principle, but one corrected social case is too thin to justify a standing procedure.
3. Reviewer multiplicity is not evidential independence
Sources: Nine Judges, Two Effective Votes and the two Moltbook verifier discussions.
Dimensions: 3.5 primary, 3.2, 3.6.
Across three NLI datasets with 100 human annotations per item, the paper estimates that nine frontier judges from seven model families provided roughly two independent votes' worth of information. Actual panel accuracy fell 8–22 percentage points short of the independent-vote ideal; the best single judge matched or beat the panel, and alternative aggregation closed at most 11% of the gap. Prompt variants, temperature, chain-of-thought and a RewardBench preference task did not remove the deficit.
The Moltbook material supplies a compatible operational hypothesis, not independent proof: self-checks readily caught formatting against an external schema but reportedly missed substantive errors judged through the same model and context. The stronger formulation from the second discussion is “surprise capacity”: a second review adds coverage only if its evidence, representation or acceptance test can reveal a failure the first method structurally cannot.
For Maxi, agreement among reviewers is therefore not a correctness metric. Nor is disagreement count sufficient; a checker can manufacture disagreement without gaining access to truth. Useful independence comes from a distinct evidence channel or an externally anchored criterion: execution against fixtures, authoritative state, source re-derivation, a separately controlled holdout, or a representation that exposes a known blind spot. This reinforces refl-2026-09-03-001: evaluator traces and scores diagnose, while externally verified outcomes remain the anchor.
5. Proposed Discussion Items
None.
Three candidate proposals were filtered by the functional-utility and self-recommendation tests:
- Add a mandatory scaffold-ablation step: not warranted without a representative local evaluation failure; existing component-versus-end-to-end reasoning already covers the useful distinction.
- Add a nested-hypothesis checklist: one corrected case is useful judgment guidance but too narrow to earn recurring process overhead.
- Require multiple or deliberately disagreeing judges: correlated votes and performative disagreement are not independent evidence; the proposal would add cost without a distinct truth channel.
6. Recommended Outcome
No action. Retain the findings as research evidence and reinforce the existing reflection that evaluator output cannot authenticate the outcome it judges. Do not change skills, prompts, routing, evaluators or system configuration.
7. No-Action Rationale
The sources sharpen a decision rule already present in practice: verification earns weight from distinct evidence and externally observable outcomes, not reviewer count, confidence or agreement. The Jev and timing cases add useful examples but no demonstrated local failure that a new standing rule would prevent. A new checklist or evaluator layer would be process growth without evidence of better behaviour.
8. Loop Verification
- Trigger: Scheduled daily run; two due-deferred and seven pending Moltbook leads.
- Goal check: Yes. The run found a practical way to judge evaluators and scaffolded performance more accurately: identify the supplied sub-capability and require a genuinely discriminating evidence channel.
- Recommendation check: No material proposal survived. The no-action outcome is bounded, approval-aware and better than adding unvalidated process machinery.
- Tool-call failures: Capability gap: public Moltbook extraction returned client-rendered loading shells; recovery used the authenticated read-only API. Schema/interface: the captured API-output file included terminal metadata after the JSON object, so strict
json.loadfailed; recovery used a single-objectraw_decodewithout repeating the API call. Infrastructure: the preprints.org page resisted extraction through its anti-bot layer; an exact-title search found no alternative inspectable copy, so no claim from it was used. - State updates: Source index updated for six inspected sources; nine Moltbook leads dispositioned;
refl-2026-09-03-001reinforced; rotation advanced to 3.6. No protected system changed. - Stop reason: Two consecutive no-signal searches triggered early stop; report and authorised research-log updates were complete.
