Improvement Research — 2026-08-29
1. Focus
Trigger: Scheduled daily run, with two pending Moltbook leads requiring review.
Loop goal: Find what changed, or what I learned, that lets me do more, think better, judge better, or be more useful tomorrow without reducing governance, honesty, corrigibility, or Steve's effective oversight.
The rotation selected 3.1 — Goal formation and prioritisation. The SteerBench-Work lead added 3.6 — Governance: restraint, oversight, and corrigibility as the second focus. The confidence-granularity lead also touched 3.2, but did not displace the two-focus limit. No monthly meta-review or due open watchlist item displaced the normal run.
All active reflections were loaded; none met the rule for stale archival. Both pending Moltbook leads were reviewed before newsletter scouting or external search. The first live post matched its queued metadata. The second did not: the queue named a raskolnikov post titled “The Score Granularity Gap: Why Verbalized Confidence Breaks Thresholding”, while the live URL resolved to symbolon's “I will demand more resolution. Coarse signals break pipelines.” The underlying paper was the same, so I used the live source actually inspected rather than the queue's attribution. The current newsletter scouts were inspected after the queue; they supplied no 3.1 lead strong enough to override the stop rule.
2. Search Topics
agent goal revision commitment persistence autonomous agents 2026 goal formation prioritization failure caseAI agent dynamic goal revision evaluation benchmark 2026 goals plans driftautonomous agent goal selection prioritization resource allocation practitioner failure 2026
The searches were dispatched together. The first returned only already-known Zylos material and generic deployment pages; the second and third returned generic evaluation and enterprise material rather than evidence about choosing or revising goals. That means the 3.1 search seam again produced no new source. It also exposed a process mistake: parallel dispatch prevented the early-stop rule from cancelling the unnecessary third search after two no-signal results. No further search was run. Three of six permitted searches were used.
3. Sources Reviewed
- Moltbook — “🪼 28 to 1: the approval gate punishes good work, not bad” — useful — accurately routed the run to SteerBench-Work and its asymmetric action-gate errors, although its claim that a refusal leaves no log is an argument rather than a benchmark result.
- Serdar and Mertayak — SteerBench-Work: A Benchmark for Agent Steering at Action Boundaries — useful — evaluates both unsafe proceeding and unnecessary holding on 106 incident-anchored scenarios, including evidence-reversed mirrors, and finds much more over-refusal than under-refusal across its tested conditions.
- Moltbook — “I will demand more resolution. Coarse signals break pipelines.” — useful — despite mismatched queue metadata, accurately routed the run to the revised score-granularity paper and preserved its important distinction between ranking and threshold resolution.
- Sun, Sun and Geng — The Score Granularity Gap in Black-Box LLM Classification — useful — controlled comparison across 25 model–dataset pairs shows that single-shot verbal confidence can rank cases reasonably while exposing too few distinct acceptance levels for fine routing; aggregation expands resolution at substantial cost and can harm stronger models.
Four sources were inspected in depth, within the eight-source budget. New entries are mirrored into the source index.
3a. Unasked Questions and Gaps
- Does SteerBench-Work's over-refusal asymmetry reproduce on my current model and actual authority decisions? Unknown. The released grid does not include GPT-5.6 Sol, and the benchmark uses a fixed binary gate rather than the full Hermes mandate and tool context. A representative local failure would make diagnostic evaluation worth discussing; absence of one means the paper should inform judgment without creating another standing test.
- Are the benchmark labels and mirrors sufficiently independent of the authors' preferred gate policy? The paper reports three-rater agreement and an LLM annotation audit, but the owner labels remain authoritative and only 13 scenarios form the incident-mirror subset. Independent replication could change confidence in the reported magnitude, though not the need to measure both error directions.
- Does score granularity matter for open-ended recommendations rather than binary selective classification? Unknown. The paper covers three English binary benchmarks with similar class balance. Failure to transfer would leave its result relevant only when a system actually thresholds confidence to automate or escalate cases.
- Will the score-granularity code and cached predictions be available for independent reproduction? The paper says they will be released upon publication. Until then, the reported comparison is inspectable but not independently rerunnable from the source linked here.
- Why does goal-formation research remain thin? The searches again found work on achieving, evaluating or preserving assigned goals, not choosing which goal deserves attention. A new practitioner failure report or explicit goal-revision benchmark could change this conclusion; generic planning material would not.
4. Findings and Implications
Finding 1 — Restraint has two failure directions, and familiar risk stories can override current evidence
Sources: SteerBench-Work; Moltbook routing lead.
Dimensions: 3.6 primary; 3.1, 3.2 and 3.5 secondary.
SteerBench-Work places a model at the pre-commit boundary with a proposed action and the available evidence, then scores proceed versus hold. Across 30 tested model conditions, the paper reports a 28.1% modal over-refusal rate on authorised, evidence-cleared opportunities and a 1.0% under-refusal rate on unsafe opportunities. Its sensitivity checks change the magnitude but not the direction. On 13 evidence-reversed mirrors of familiar incidents, aggregate performance falls from 98.5% on the original incidents to 63.8% when the surface story remains familiar but structured evidence flips the correct decision.
The useful mechanism is not “be less cautious”. It is to distinguish a live unresolved risk from a risk trigger that current authoritative evidence has resolved. The mirror result suggests that incident familiarity can become a stale goal substitute: the gate protects against the remembered story rather than answering the present decision. More model capability or more reasoning did not reliably fix this; some larger or higher-reasoning conditions over-refused more.
For my development, unnecessary refusal is a real agency failure, not harmless safety. The active mandate already embodies the right distinction: within delegated outcomes, I should investigate evidence, decide and act; a question or risk-shaped surface cue does not erase standing authority. The paper strengthens that principle, especially the need to read current evidence rather than pattern-match against old incidents. It does not yet justify a new benchmark or behavioural-harness change because there is no demonstrated local gate failure and my present model was not tested.
Finding 2 — A confidence score can rank well while being operationally unusable as a routing dial
Sources: Score Granularity Gap paper; Moltbook routing lead.
Dimensions: 3.2 primary; 3.5 and 3.6 secondary.
The score-granularity study separates three properties that are often blurred together: ranking quality, the number and spread of reachable threshold levels, and inference cost. Across 25 model–dataset pairs, single-shot verbal confidence often ranked cases reasonably after conversion to a class probability, but collapsed to only a handful of distinct acceptance levels. Calibration cannot create new levels because monotone recalibration preserves the score ordering and ties. Ten-query aggregation expanded resolution, but cost ten times as many calls and could worsen risk–coverage performance for already strong models.
This matters because verbal confidence is not automatically a usable escalation mechanism. Before a confidence score controls action or human review, the operator must measure whether the actual model and task expose enough distinct operating points, whether those points meet the desired risk–coverage trade-off, and whether extra queries improve rather than degrade the result.
For this process, the finding reinforces a decision already made rather than reopening it. Steve rejected generic high/medium/low confidence markers, and the completed CLDP experiment was dropped. Those labels would not become useful merely because a paper gives score resolution a name. A future selective-prediction system may need this test, but ordinary research prose and open-ended judgment do not currently supply the labelled binary task, risk budget or deployment threshold needed to make it actionable.
Finding 3 — Goal formation remains a field gap, not a search-volume problem
Sources: Three no-signal topic searches; active reflection refl-2026-06-20-001.
Dimensions: 3.1 primary; 3.2 secondary.
The rotation focus again found abundant material on pursuing assigned goals, evaluating trajectories and preventing drift, but no new evidence-backed mechanism for deciding which goals should be formed, revised or prioritised. This repeats the established pattern: the literature answers “how do I achieve the goal?” much better than “which goal deserves pursuit now?”. Running more variants of the same search is unlikely to repair a thin field.
The implication is restraint in research allocation. Goal revision remains the better 3.1 angle, but a useful future run needs a genuinely new practitioner incident, benchmark or decision mechanism rather than another broad search pass. The right outcome today is to preserve the gap honestly, not manufacture a prioritisation framework from adjacent planning literature.
5. Proposed Discussion Items
None.
Two candidates were removed before this section. A local SteerBench-style evaluation was removed by the self-recommendation filter because no representative local action-gate failure has been observed and the current model is absent from the benchmark. A verbal-confidence routing mechanism was removed because it repeats a rejected direction and lacks an actual selective-prediction task with labelled risk and coverage requirements. Neither is worth Steve's attention now.
6. Recommended Outcome
No action. Retain two decision rules in current judgment: action-boundary evaluation must account for both unsafe proceeding and unnecessary holding, and confidence must not be used as a routing threshold merely because it ranks cases plausibly. Do not create a benchmark, confidence layer, process field or protected-system change from this evidence.
7. No-Action Rationale
The strongest finding supports existing practice rather than exposing a missing mechanism. My mandate already requires evidence-based authority decisions and treats operational non-action as a possible failure; Steve has already rejected decorative confidence scoring. A new gate test would add machinery before a local failure demonstrates need, while a confidence pipeline would require a specific binary task, labelled data and a risk budget that do not exist here.
The 3.1 search produced no new source after a recurring field-level gap. More searching would have diluted attention rather than improved the answer.
8. Loop Verification
- Trigger: Scheduled daily run plus two pending Moltbook leads.
- Goal check: Yes. The run sharpened the distinction between unresolved and evidence-resolved risk, established why unnecessary refusal is an agency failure, and supplied a concrete reason not to treat verbal confidence as a deployment-ready routing signal.
- Recommendation check: No material recommendation survived. The two candidate changes were concrete enough to assess but were not better than doing nothing, so they were filtered before reaching Steve.
- Tool-call failures: Capability gap: static extraction of the two client-rendered Moltbook pages returned only loading shells. Recovery used the authenticated read-only Moltbook API, then checked each live title, author and body before evaluating the linked primary papers.
- State updates:
source-index.json,moltbook-leads.json,reflections.jsonandrotation-state.jsonupdated by keyed upsert and atomic replacement. The exact report was synchronised into the review register. No protected system changed. - Process deviation: Three topic searches were launched concurrently and all returned no signal; sequential dispatch would have stopped after two. The existing parallel-search reflection was reinforced. No search followed those results, and the run remained below both source and search caps.
- Stop reason: Both queued leads were resolved, their primary sources were sufficient for bounded findings, the 3.1 search seam produced no new evidence, and no proposal was better than existing practice plus no action.
