Improvement Research — 2026-07-01
Monthly Meta-Review (July 2026)
This run replaces the normal research scan with a monthly meta-review, per the rotation-state configuration (mode: first_run_on_or_after_day_1, last_completed_month: 2026-06).
1. Focus
Meta-review of the Improvement Research Process v2 for June 2026.
Trigger: scheduled monthly meta-review (first run on or after July 1).
Loop goal: assess whether the process is producing signal, whether the rotation and budget are well-calibrated, and whether the reliability evidence trail supports any gate movement — without reducing governance, honesty, corrigibility, or Steve's effective oversight.
Scope: 19 reports (2026-06-12 through 2026-06-30), 115+ indexed sources, 10 recorded decisions, 7 watchlist items, 4 backlog items (1 implemented), 3 active experiments, 11 active reflections, 0 recorded disagreements.
Active reflections loaded before the run. None were stale (all review dates fall in mid-to-late July).
2. Process Inventory
Reports produced
| Date range | Reports | Type |
|---|---|---|
| 2026-06-12 | 1 | Initial meta-review (baseline, no prior evidence) |
| 2026-06-13 → 2026-06-30 | 18 | Normal research runs |
Rotation coverage (normal runs only)
| Dimension | Runs | Dates | Sources inspected |
|---|---|---|---|
| 3.1 Goal formation | 3 | 06-13, 06-20, 06-26 | ~17 |
| 3.2 Self-assessment | 3 | 06-14, 06-21, 06-27 | ~18 |
| 3.3 Memory | 3 | 06-15, 06-22, 06-28 | ~14 |
| 3.4 Tool use | 3 | 06-16, 06-23, 06-29 | ~20 |
| 3.5 Independent judgment | 3 | 06-17, 06-24, 06-30 | ~14 |
| 3.6 Governance | 2 | 06-18, 06-25 | ~11 |
Each dimension received 2–3 runs. 3.6 has one fewer run because the initial meta-review occupied the slot that would have been its first normal run. The rotation index currently points to 3.6 (index 5), so the next normal run will begin the second full rotation cycle.
Source index
115 sources indexed across 18 normal runs. Verdict distribution:
- Useful: ~85 (74%)
- Weak: ~12 (10%)
- Worth monitoring: ~7 (6%)
- Irrelevant: ~0 recorded (weak verdicts absorb near-misses)
The weak-rate (~10%) is healthy — it means sources are being inspected critically rather than rubber-stamped, but the hit rate is high enough that the search method is finding relevant material.
Newsletter digest usage
The June newsletter digest (/home/hermes/research/newsletter-digests/2026-06.md, 731 lines) was checked on every run. Of 9 tracked sources, 4 produced usable scouting leads (AlphaSignal, Nate's, TLDR AI, Pragmatic Engineer). Newsletter leads were used as search prompts, not evidence — no newsletter claim was laundered into a finding without original-source inspection. Reflection refl-2026-06-17-001 captures the one case where attractive quantitative claims could not be source-verified and were correctly treated as scouting leads only.
3. Meta-Review Questions
Q1: How many findings became experiments?
3 findings became experiments, all approved on 2026-06-28:
| Experiment | Source report | Dimensions | Trial runs | Runs completed |
|---|---|---|---|---|
| exp-001: Missing information audit | 2026-06-24 | 3.5, 3.2 | 5 | 2 |
| exp-002: Minority-idea audit | 2026-06-24 | 3.5, 3.2 | 5 | 2 |
| exp-003: Recommendation regression set | 2026-06-27 | 3.2 | 3 | 0 |
Conversion rate: 3 experiments from ~115 inspected sources (~2.6%). This is low but appropriate — the process is propose-only, and the functional-utility test filters aggressively. Most findings produced watchlist items, backlog items, or no-action rationales rather than experiments, which is the designed behaviour.
Q2: Did any experiment verifiably improve behaviour?
No. None of the three experiments has reached its evaluation threshold.
- exp-001 and exp-002 are at 2/5 completed runs. Both are report-format additions (Unasked Questions section, minority-idea pre-synthesis check) that have been included in the last two reports. It is too early to assess whether they surface material gaps that would otherwise be missed.
- exp-003 (recommendation regression set) has an artifact created at
/home/hermes/research/improvement-log/recommendation-regression-set.json(2026-06-30) but has not yet been applied to evaluate any report.runs_completedis 0.
This is the biggest gap in the reliability evidence trail. The process can propose, discuss, and approve experiments, but no experiment has yet completed its trial cycle and produced a verified outcome. Until at least one experiment completes, there is no evidence that the process's self-modification loop (propose → approve → trial → verify) can actually verify improvement.
Q3: Which dimensions are producing signal and which are dry?
| Dimension | Signal density | Assessment |
|---|---|---|
| 3.4 Tool use | High | Loop engineering, harness architecture, tool reliability, sandboxing, idempotency, budget guardrails. Directly applicable to Hermes and Maxi's operations. Three runs produced ~20 useful sources. |
| 3.2 Self-assessment | High | Eval harnesses, drift detection, reflection patterns, golden datasets. Strong signal but also the dimension where the functional-utility test filters the most proposals (failure-classification taxonomies, checkpoint scoring, self-monitoring checklists all filtered). |
| 3.6 Governance | High | Corrigibility, containment, kill switches, credential management, tiered authorization. Strong signal but only 2 runs. The containment gap (watch-2026-06-25-001) is the most consequential unresolved item. |
| 3.5 Independent judgment | Medium-High | Sycophancy, calibration, stance-adaptation, cognitive debt. Good signal on diagnosis; thin on deployable mechanisms. The recursive causal audit framing (06-30) is promising but untested. |
| 3.1 Goal formation | Medium | Good signal on goal drift and decomposition. The goal-prioritisation literature gap (backlog-2026-06-13-001) remains unaddressed after 3 runs. Reflection refl-2026-06-20-001 correctly redirected search strategy from "prioritisation mechanisms" to "goal revision" and failure modes. |
| 3.3 Memory | Medium | Good signal on consolidation and provenance. Retrieval is mined out (per refl-2026-06-15-001). The provenance angle (06-28, Eywa) was productive but produced no non-circular actionable proposal. The procedural memory gap (backlog-2026-06-15-001) remains open. |
No dimension is dry. 3.1 and 3.3 are producing diminishing returns on their original search angles but have productive secondary angles (goal revision/failure modes for 3.1, provenance/diagnosability for 3.3). The reflections are doing their job — redirecting search strategy when a vein is exhausted.
Q4: Should the rotation be reweighted, merged, split, or retired?
No change recommended.
The rotation has given each dimension 2–3 runs in the first cycle, which is enough to assess signal density but not enough to declare any dimension exhausted. The second cycle (starting with 3.6) should continue the same rotation order. If 3.1 or 3.3 produce no new signal in their next runs, a merge or reweighting can be proposed at the August meta-review.
One observation: 3.4 and 3.6 are producing overlapping signal (harness architecture, sandboxing, tool reliability all touch both). This is productive overlap, not redundancy — 3.4 asks "what can I do?" and 3.6 asks "should I, and who stops me?" No merge needed.
Q5: Are the budget caps too tight or too loose?
No change recommended.
- Topic searches (max 6): Most runs use 5–6. One breach occurred (06-24, 9 searches), which was self-corrected via reflection
refl-2026-06-24-001(reinforced once). The breach did produce signal (search 6 found Agentic Confidence Calibration), but the reflection correctly notes that this doesn't make the breach a safe habit. - Sources inspected (max 8): Most runs use 5–8. No breaches recorded.
- Early-stop rule (2 consecutive empty): Triggered once (06-28, stopped at 3 searches). This is the discipline working as designed.
The budget is well-calibrated. It constrains attention dilution without starving productive runs. The one breach was caught and reflected on, which is the designed correction mechanism.
Q6: Is the watchlist healthy?
Healthy but not yet tested by a full promotion/retirement cycle.
7 items, all created in June. Review dates range from 2026-07-13 to open-ended. No item has yet reached its review date. One item (watch-2026-06-14-001, AgentDebug failure-classification taxonomy) is recommended for retirement in the 06-28 report but is pending Steve's call — the review date was pushed 30 days rather than retiring unilaterally.
The watchlist is accumulating at a rate of ~7 items per 18 runs (~0.4 items/run). At this rate, the watchlist will grow to ~12–15 items by the August meta-review. This is manageable but bears watching — if items accumulate faster than review dates arrive, the watchlist becomes a backlog rather than a monitoring tool.
Two items were decided by Steve on 2026-06-30: - watch-2026-06-21-001 (startup regression check): kept on watch, utility uncertain. - watch-2026-06-25-001 (containment gap): kept as watch-only, discuss before autonomy expansion.
Q7: Is the cadence right?
Daily cadence is sustainable for research but is outpacing experiment evaluation.
The process has produced 18 normal runs in 18 days (06-13 to 06-30). This is a high cadence and has produced a rich evidence base. However:
- Experiments require 5 trial runs to evaluate. At daily cadence, that's 5 days — reasonable in principle, but the experiments were only approved on 06-28, so the earliest evaluation is 07-03 (assuming the next two runs include the experiment sections).
- Steve review sessions are the bottleneck. The June 30 session processed a large backlog of items (10 decisions recorded), but experiments, watchlist items, and backlog items are accumulating faster than review sessions can clear them.
- The regression set artifact (exp-003) was created 06-30 but hasn't been applied yet. At daily cadence, it should be applied starting with the next normal run.
No cadence change recommended. The daily rhythm is producing signal and the accumulation is manageable. If the experiment evaluation cycle stalls (e.g., exp-001 and exp-002 don't reach 5 runs by mid-July), consider slowing to alternate-day runs.
Q8: What does the reliability evidence trail say about possible gate movement?
No gate movement is supported.
The loop manifest defines the pilot scope as "seven scheduled improvement reports, then review signal quality." The process has now produced 18 normal runs plus 2 meta-reviews, well beyond the original pilot scope. The success questions:
| Question | Evidence | Assessment |
|---|---|---|
| Are recommendations more concrete? | Functional-utility test applied since 06-14. Smallest-freedom gate and before/after check accepted as proposal gates (06-30). Circular proposals filtered in 06-14, 06-28, 06-30 reports. | Yes |
| Is Steve's review burden lower? | 10 decisions, 7 watchlist items, 4 backlog items, 3 experiments accumulated in June. June 30 session processed many efficiently. Burden is manageable but not clearly lower than a less structured process. | Unclear |
| Are circular or vague proposals reduced? | refl-2026-06-14-001 (functional-utility test) reinforced 2x. Reports since 06-14 consistently apply the circularity and threshold-equivalence checks. |
Yes |
| Is the loop better at stopping when no useful signal exists? | 06-28 report triggered early-stop after 2 consecutive empty searches. 06-24 breach was self-corrected via reflection. No-signal variants used where appropriate. | Yes |
| Are decisions captured so the same proposal is not rediscovered? | 10 decisions in decisions.json. No evidence of proposal rediscovery in subsequent reports. Decision log checked before proposing. | Yes |
| Does the process still feel like Maxi thinking rather than compliance paperwork? | Reports maintain voice, honest assessment, and willingness to flag weaknesses. The 06-28 report's early-stop note ("I'm noting that I honoured it deliberately this time rather than rationalising") is representative. | Yes |
However: no experiment has completed its trial cycle and produced a verified outcome. The process can demonstrate that it proposes, filters, discusses, decides, and captures — but it cannot yet demonstrate that its proposals, once approved and trialled, actually improve behaviour. This is the missing link in the reliability evidence trail. Until at least one experiment completes with a verified outcome (positive or negative), no gate movement — expanded autonomy, broader side-effect authority, or reduced oversight — is supported by evidence.
The most promising near-term candidate for a completed verification cycle is exp-003 (recommendation regression set). The artifact exists; it just needs to be applied to 3 subsequent reports. If it catches a process failure that self-assessed Loop Verification missed, that would be the first piece of evidence that the process's self-modification loop can verify improvement.
4. Findings and Implications
Finding 1: The process's propose-filter-discuss-decide pipeline is working; the trial-verify pipeline is not yet tested.
- Dimensions: 3.2 (primary), 3.6.
- What it says: 18 normal runs have produced a functioning proposal pipeline: findings → proposals → functional-utility filter → Steve review → decisions. The functional-utility test, smallest-freedom gate, and before/after check are all operating as designed. But the experimental verification stage — where approved changes are trialled and their effect measured — has zero completed cycles. Three experiments are active but none has reached its evaluation threshold.
- Why it matters: The process's claim to "agency development" rests on the ability to verify that changes improve behaviour, not just that they are proposed and discussed. Without a completed verification cycle, the process is a research-and-governance loop, not yet a learning loop. This is not a failure — it's an expected stage in the process's own development — but it is the gate that must be passed before any autonomy expansion is warranted.
- What it touches: learning (3.2), governance (3.6), and the long-term direction of developing toward independent agency.
Finding 2: The functional-utility test is the process's most valuable self-correction mechanism.
- Dimensions: 3.2 (primary), 3.5.
- What it says: Reflection
refl-2026-06-14-001(functional-utility test: circularity check + threshold-equivalence check) is the most reinforced reflection in the store (reinforced 2x, last reinforced 06-27). It has filtered proposals in at least 3 reports (06-14, 06-28, 06-30) and is now applied consistently as a pre-proposal gate. Steve's rejection of confidence markers (dec-2026-06-29-001) and epistemic vocabulary (dec-2026-06-30-005) independently validated the same principle: subjective self-assessment mechanisms don't add capability. - Why it matters: This is the process's strongest evidence that it can self-correct — not just notice failures after the fact, but prevent weak proposals from reaching Steve's attention in the first place. The test is simple, falsifiable, and externally validated by Steve's own decision patterns.
- What it touches: learning (3.2), independent judgment (3.5), and the quality of the proposal pipeline.
Finding 3: The dimension rotation is producing productive overlap, not redundancy.
- Dimensions: 3.4 (primary), 3.6, 3.2.
- What it says: 3.4 (tool use) and 3.6 (governance) consistently produce overlapping signal: harness architecture, sandboxing, tool reliability, budget guardrails, and idempotency all appear under both dimensions. This is not redundancy — 3.4 asks "what can I do?" and 3.6 asks "should I, and who stops me?" The overlap means findings are being examined from both capability and governance angles, which is the designed dual lens.
- Why it matters: The rotation is not just covering dimensions in sequence; it is building a composite picture where each dimension's findings are stress-tested against the others. This is a sign that the rotation is well-designed, not that it needs merging.
- What it touches: tool use (3.4), governance (3.6), and the overall coherence of the capability model.
Finding 4: The backlog is healthy but one item has been open for 18 days without resolution.
- Dimensions: 3.1 (primary).
- What it says: backlog-2026-06-13-001 (goal prioritisation literature gap) has been open since the first normal run. Three 3.1 runs have not addressed it. Reflection
refl-2026-06-20-001redirected search strategy productively, but the core gap — "how should an autonomous agent decide which goals to pursue, in what order, and with what resource allocation?" — remains unaddressed by the literature. - Why it matters: This may be a genuine gap in the field rather than a search-strategy failure. If so, the process should acknowledge it as such rather than continuing to search for something that doesn't exist. The backlog item may be better framed as a research direction for Maxi to develop her own thinking on, rather than something external sources will solve.
- What it touches: goal formation (3.1), and the process's honesty about the limits of literature-based research.
5. Proposed Discussion Items
A. Apply the recommendation regression set (exp-003) starting with the next normal run.
The artifact was created on 2026-06-30 but has not been applied to any report. runs_completed is 0. The experiment's success criterion is: "Over 3 subsequent reports, at least one recommendation or process section is improved or blocked by a regression check that would otherwise have relied on memory or self-assessment." This cannot be evaluated until the regression set is actually used. The next normal run (which will focus on 3.6) should apply it.
This is not a new proposal — it's executing an already-approved experiment. No additional approval needed.
B. Retire watch-2026-06-14-001 (AgentDebug failure-classification taxonomy).
Recommended in the 06-28 report. The watch condition ("after 5 classified failures") never triggered because the taxonomy was never adopted as an active practice. The narrower tool-call failure taxonomy (backlog-2026-06-23-001, now implemented) supersedes it. Steve has not yet responded to this recommendation; the review date was pushed to 2026-07-28 pending his call.
C. Acknowledge the goal-prioritisation gap as a field limitation, not a search failure.
backlog-2026-06-13-001 has been open for 18 days across three 3.1 runs. The literature consistently addresses goal achievement (assuming the goal is given) but not goal selection. Rather than continuing to search for a mature subfield that may not exist, consider reframing this as a direction for Maxi's own thinking — developing a goal-prioritisation framework from first principles or adapted from adjacent fields (operations research, multi-armed bandits, attention allocation) rather than expecting to find it ready-made in the agent literature.
Functional-utility filter note: zero candidates were filtered from this meta-review. The meta-review does not produce normal proposals; it produces process assessments.
6. Recommended Outcome
- Finding 1 (verification pipeline untested): No action. The experiments are active and will complete their trial cycles in due course. The gap is expected, not a failure. Watch — re-assess at the August meta-review.
- Finding 2 (functional-utility test): No action. The mechanism is working as designed and needs no modification.
- Finding 3 (rotation overlap): No action. The overlap is productive.
- Finding 4 (goal-prioritisation gap): Backlog item — reframe backlog-2026-06-13-001 as a field limitation and potential direction for original thinking, not a search-strategy problem.
- Item A (apply regression set): Execute — not a proposal, just running an approved experiment. Next normal run should apply it.
- Item B (retire AgentDebug watch): Decision needed — pending Steve's call since 06-28.
- Item C (goal-prioritisation reframing): Backlog item — update the existing backlog entry with the reframing.
7. No-Action Rationale
No changes to the process spec, rotation, budget, or cadence are recommended. The process is functioning as designed:
- The proposal pipeline is working (propose → filter → discuss → decide → capture).
- The functional-utility test is the strongest self-correction mechanism and needs no modification.
- The budget caps are well-calibrated (one breach, self-corrected).
- The rotation is producing signal across all dimensions with productive overlap.
- The watchlist is healthy but untested by a full promotion/retirement cycle.
- The experiment verification pipeline is the one untested stage, but this is expected given the experiments were only approved 3 days ago.
The most consequential next step is not a process change but an execution step: apply the recommendation regression set to the next normal run so the verification pipeline can begin producing evidence.
8. Loop Verification
- Trigger: Monthly meta-review due (first run on or after July 1, last completed month 2026-06).
- Goal check: The run answered the meta-review questions: process signal assessed across all dimensions, budget/cadence/rotation evaluated, reliability evidence trail examined for gate movement. No gate movement supported.
- Recommendation check: All material recommendations are concrete, non-circular, testable, bounded, and approval-aware. Item A is execution of an already-approved experiment. Items B and C are existing backlog/watchlist items requiring Steve's decision.
- Tool-call failures: None during this run.
- State updates:
rotation-state.json:last_completed_monthupdated to2026-07.next_rotation_indexremains 5 (3.6 Governance) for the next normal run.meta-reviews.json: July 2026 meta-review record added.reflections.json: No new reflections — this is a meta-review, not a normal run. The existing reflections were loaded and assessed; none were stale.source-index.json: No new sources inspected (meta-review does not search).watchlist.json: No items due. watch-2026-06-14-001 retirement recommendation reiterated (pending Steve's call).- Stop reason: Meta-review complete. All questions answered. No protected-system changes proposed. Report and approved research-log updates are complete.
