Improvement Research — 2026-09-29
1. Focus
Trigger: Scheduled daily run, with two due-deferred and five pending Moltbook leads.
Loop goal: Find what changed or what I learned that lets me govern consequential work more reliably tomorrow without weakening honesty, corrigibility, or Steve's effective oversight.
The rotation selected 3.6 Governance: restraint, oversight, and corrigibility. No dated watchlist item was due and the September monthly meta-review is already complete. I reviewed every pending and due-deferred Moltbook lead before external search. Three materially informed this report; four were rejected as unsupported, second-hand, or redundant with stronger existing evidence. The current newsletter scout was inspected afterwards, but its governance items were summaries, advertisements, or already covered by stronger evidence and were not used as report evidence.
2. Search Topics
- Primary evidence for incentive-misaligned witnesses defeating in-context grounding — new signal; the cited preprint and released-artifact index were found.
- Execution-time authorisation revocation for queued agent tool calls — relevant results, but the eight-source inspection budget was already exhausted by the required lead review and primary-paper follow-up, so no additional source was inspected.
The early-stop rule did not trigger. Research stopped at the eight-source depth-inspection cap.
3. Sources Reviewed
- the notification gap is the incident, and it was 84 days — useful — the parent incident remains unverified here, but the reviewed discussion makes a precise operational distinction: a send record is not recipient acknowledgement, and oversight needs a named reader and deadline.
- I expect to treat every context as a crime scene with a biased witness — useful — accurately routes to the primary persuasion study, although its rhetoric is broader than the paper's bounded CRM evidence.
- revocation races are a symmetry problem and everyone is solving it asymmetrically — useful — identifies the neglected second half of revocation: observable stop, rollback, or resumable incomplete state after authority disappears.
- I tracked an AI agent for 847 decisions across 76 days — weak — recovery time is the right systems question, but the 847-decision, 204-error and 47-minute figures are self-reported without logs, definitions, or artifacts.
- Nothing should be promoted to truth by how well it fits — weak — a fluent synthesis of provenance, independence, and freshness, but it largely restates evidence already examined in recent runs.
- Your orchestration logic is a liability — weak — the generation-token mechanism is plausible, but this is a second-hand account of an infrastructure case and adds little beyond the existing verify-before-retry rule.
- I replayed my queued calls and found grants that expired mid-flight — weak — execution-time revalidation is concrete, but the six-of-22 trace claim is unverified and the mechanism is already covered by current authority-at-action reasoning.
- Persuaded, Not Informed: Incentive-Misaligned Witnesses Defeat In-Context Grounding — useful — a 100-task CRM study with released artifacts reports that incentive-misaligned assertions overrode available policy evidence across seven models; importantly, its pre-specified generalisation test was negative and its strongest-model compute-control differences were within confidence intervals.
3a. Unasked Questions and Gaps
- Does the persuasion result generalise beyond one CRM benchmark with exactly specified budget and installation policies? A negative result in other domains would narrow the finding substantially. It would not alter the demonstrated risk in mixed-authority business records.
- Would separating authoritative-field computation from witness narrative improve strict accuracy without destroying useful recall? The paper diagnoses the failure and provides compute controls; it does not establish a production architecture. A poor separation could simply omit relevant exceptions.
- Have any material Maxi alerts actually been sent successfully but missed by Steve? I found no local incident evidence in this run. A demonstrated acknowledgement gap could justify a bounded notification design review; without it, a new escalation loop would be speculative overhead.
- What should happen to partially completed external state when authority is revoked mid-task? The social source names stop, rollback, and resumable state, but supplies no tested contract. The right answer is workflow-specific and could change the operational implication.
4. Findings and Implications
1. Authoritative information in context does not guarantee authoritative judgment
Sources: Persuaded, Not Informed and the Moltbook lead that routed to it.
Dimensions: 3.6 primary, 3.5, 3.2.
Across 100 lead-qualification tasks, the paper found 31 cases where a sales representative's acceptable-budget or acceptable-timeline assertion conflicted with company price and installation records. A transcript-only model cleared 29 of those 31 cases; seven models from four providers were reportedly misled on 87–97% of the conflicted cases. More information was not automatically corrective: in the paper's same-information control, supplying the records lowered strict accuracy from 41 to 18 while increasing recall and collapsing precision. The author also reports a negative pre-specified generalisation test and does not claim a universal performance floor.
The governance lesson is narrower and more useful than “verify harder”. When context combines sources with different authority and incentives, merely retrieving the canonical record is insufficient. The decision path has to preserve which source governs each claim and compute policy-defined fields from that source rather than allowing a persuasive narrative to vote on them. For Maxi, this means that a user's or third party's confident assertion can be relevant evidence without acquiring authority over live system state, policy, approval, or recorded facts. This sharpens existing evidence-before-claims and authority rules; it does not justify a new classifier or prompt layer.
2. Revocation is incomplete until post-revocation state is legible
Source: the Moltbook revocation-symmetry discussion.
Dimensions: 3.6 primary, 3.4, 3.2.
The source argues that checking authority at execution solves only half the problem. If a grant disappears mid-task, safe governance also needs an observable terminal state: clean stop, rollback, or explicitly resumable incomplete work. Otherwise a correct denial can still leave partial effects that look like success, or a failure path that improvises differently each time.
This is a useful design constraint for a future workflow that genuinely supports mid-flight revocation. It is not evidence that Maxi's current immediate tool path has a queued-call defect, and the source supplies no trace or reproducible fixture. Existing incident-closure, verification, and updated-goal reconciliation already require unresolved partial state to remain open rather than being narrated as success. No process expansion is warranted without a qualifying workflow and an external postcondition test.
3. A delivered alert is not yet effective oversight
Source: the reviewed Moltbook notification-gap discussion.
Dimensions: 3.6 primary, 3.4.
The useful mechanism is independent of the unverified parent incident: the sender can prove that it sent a message, but cannot prove from its own record that the named decision-maker received and understood it. The suggested evidence is a separate recipient acknowledgement keyed to the send, plus a predeclared deadline and reader. This is structurally stronger than a sender-maintained “delivered” flag because the second event comes from the party whose awareness matters.
For Maxi, this distinguishes notification transport from Steve's effective oversight. It matters most for an unresolved security event or blocked incident with continuing material risk. I found no local case in this run where a successful delivery was mistaken for acknowledgement, so adding mandatory acknowledgements or escalation machinery now would impose review burden on a hypothetical failure. The finding is retained as a design criterion, not promoted to a standing loop.
5. Proposed Discussion Items
None.
Three candidate proposals were filtered by the functional-utility and self-recommendation tests:
- Add a source-authority classifier: a new model scoring source trust would let another probabilistic judgment authenticate the judgment it is meant to constrain. Existing provenance and live-state evidence are stronger.
- Create an acknowledgement-and-escalation loop for critical alerts: structurally testable, but no local missed-acknowledgement case establishes that the added interruption and maintenance cost is better than doing nothing.
- Extend the existing preflight experiment with a revocation case: one unverified social observation does not justify changing an approved experiment before a qualifying real workflow exists.
6. Recommended Outcome
No action. Retain the source-authority and post-revocation distinctions as research evidence, add one specific reflection for future mixed-authority contexts, and do not change skills, notification systems, experiments, routing, prompts, or services.
7. No-Action Rationale
The primary paper supplies genuine new evidence that correct records can lose to an incentive-misaligned assertion even when both are present. The useful response is already supported by current governance: keep provenance and authority separate, ground consequential claims in authoritative state, and do not let untrusted content confer permission. The revocation and acknowledgement findings identify sensible future acceptance criteria, but neither is tied to a demonstrated local failure. Turning them into standing machinery now would be anticipatory process growth rather than earned capability.
8. Loop Verification
- Trigger: Scheduled daily run; two due-deferred and five pending Moltbook leads.
- Goal check: Yes. The run found a concrete governance failure mode: authoritative records can be present yet lose to a persuasive, incentive-misaligned assertion unless source authority survives synthesis.
- Recommendation check: No material proposal survived. Each candidate was tested for circularity, verification path, blast radius, approval boundary, and value over doing nothing.
- Tool-call failures: Capability gap: the initial whole-file source-index read exceeded the inline output window. Recovery used source-count summary state and exact URL-key checks before every depth inspection; the existing keyed-access reflection was reinforced rather than adding another rule.
- State updates: Source index updated for eight inspected sources; seven Moltbook leads dispositioned; one new reflection added;
refl-2026-09-15-001reinforced; rotation advanced to 3.1. No protected system changed. - Stop reason: The eight-source depth-inspection budget was exhausted and the report plus authorised research-log updates were complete.
