Improvement Research — 2026-08-01
Monthly Meta-Review (August 2026)
1. Focus
Trigger: scheduled monthly meta-review, started 2026-08-01 05:01 AWST. The rotation state requires the first run on or after the first of the month to replace the normal scan; July's review is recorded as complete.
Loop goal: assess whether the Improvement Research Process is producing signal, whether its rotation, budget, watchlist, and experiments remain useful, and whether its reliability evidence supports any gate movement—without reducing governance, honesty, corrigibility, or Steve's effective oversight.
This review covers July: 31 report files (one July meta-review and 30 normal runs), the research-log stores, 122 sources indexed during July, 23 July decisions, 12 watchlist records, and five experiments. I loaded the loop manifest, active reflections, source index, watchlist, backlog, experiments, disagreements, decisions, rotation state, and meta-review history. No watchlist item was due on 1 August. The latest newsletter scout was inspected before deciding whether an external scan was needed; it was treated as scouting only. A meta-review needs the process's own evidence, not a fresh source search.
Two active reflections have passed their review dates, but neither was archived: both have non-zero reinforcement counts, so the archive condition does not apply.
2. Search Topics
No external topic searches were run. This scheduled meta-review used the bounded internal evidence trail rather than spending the normal six-search budget. The 31 July newsletter scout was read only as a lead check; no digest claim is a finding in this report.
3. Sources Reviewed
No external source was inspected in depth. The source index and July reports were reviewed as operational evidence:
source-index.json— useful — records 122 July inspections: 96 useful, 20 weak, five worth monitoring, and one moderately useful.experiments.json— useful — records four completed report-format experiments and one bounded active shared-knowledge trial.decisions.json— useful — records 23 July decisions: 10 accepted and 13 rejected.watchlist.json— useful, with a data-quality issue — records nine open entries and three closed entries, but two exact duplicate objects each appear underwatch-2026-07-04-001andwatch-2026-07-04-002.
No source-index entry was added or changed.
3a. Unasked Questions and Gaps
- Did the promoted Unasked Questions and Gaps step reduce Steve's actual review burden, rather than merely surface gaps? The experiment establishes that it surfaced material gaps in its five trials, not its effect on review time. If Steve's experience were that it adds burden without better decisions, the current positive assessment would weaken.
- Will the active shared-knowledge trial meet its cross-agent reuse and maintenance thresholds? It is not due for its 6 August midpoint review or 20 August final review. If it misses those criteria, July's positive evidence about bounded experimentation would not transfer to that trial.
- Does the date mismatch in the
2026-07-31.mdreport heading reflect only a heading error or a broader run-dating issue? The file is named 2026-07-31 but its H1 says 2026-07-30. If the run actually began on 31 July, only metadata repair is needed; if it began earlier, the monthly count and sequence need rechecking. - Do the two duplicate watch objects create operationally different review behaviour? They are byte-for-byte identical, so no difference is currently evidenced. If the watchlist consumer deduplicates by ID, the practical impact is lower; if it iterates records, the 4 August reviews will be duplicated.
4. Findings and Implications
1. The trial-and-review part of the loop is now evidenced, but only for small report-format interventions
Source: experiments.json and decisions.json.
Dimensions: 3.2 (primary), 3.6, 3.5.
Four approved experiments completed during July. One was promoted after five runs: the Unasked Questions and Gaps section consistently surfaced material gaps. Three were deliberately dropped or archived after their trials showed redundancy, no differentiating signal, or zero catches. This is better evidence than the July meta-review had: the process can now run a bounded change, observe a negative result without rationalising it into success, and retain the decision.
The scope matters. These were report-format mechanisms with no side-effect authority, not tests of autonomous execution or broader self-modification. The result supports the present propose → approve → trial → verify discipline; it does not support reduced oversight, expanded authority, or claims that the process can yet verify agency gains beyond its own reporting surface.
2. Governance research produced the densest useful signal, while every dimension still produced enough to retain the rotation
Source: source-index.json and July reports.
Dimensions: 3.6 (primary), 3.2, 3.4, 3.1, 3.3, 3.5.
July's indexed useful-source distribution was: 3.1 goal formation 9; 3.2 learning loops 13; 3.3 memory and continuity 16; 3.4 tools and environment 16; 3.5 independent judgment 11; 3.6 governance 31. The other July entries were critically classified as weak, worth monitoring, or moderately useful rather than treated as signal.
3.6 is clearly the strongest current seam, particularly around concrete oversight, authority boundaries, incident learning, and verification. But the spread is not a reason to starve the other dimensions. 3.1's smaller count still sharpened the distinction between choosing ends and allocating means; 3.5 continued to expose why self-scoring and uncalibrated AI judges are not independent evaluation. The six-dimension rotation remains useful because it prevents tool capability from being mistaken for a reason to use it.
3. The budget and stop rules are constraining activity without blocking signal
Source: July reports and rotation-state.json.
Dimensions: 3.2 (primary), 3.6.
Eight July reports explicitly triggered the two-consecutive-no-signal stop condition. At least three ran to the six-topic-search cap, while source inspections remained within the eight-source cap. The last two normal runs show the intended discipline: a 3.1 run stopped after two no-signal searches once it had one relevant preprint, and the subsequent 3.2 run stopped after two no-signal searches with no depth inspection at all.
That is evidence against increasing the budget merely because daily runs exist. The process is capable of saying that a search angle is saturated, and the source index prevents known material being relabelled as discovery. Retain the existing caps and early-stop rule.
4. The watchlist has a small integrity defect that should be repaired before it becomes false operational signal
Source: watchlist.json; report filename and heading check.
Dimensions: 3.2 (primary), 3.3, 3.6.
The watchlist contains two pairs of byte-identical duplicate objects: watch-2026-07-04-001 appears twice and watch-2026-07-04-002 appears twice. Both are due on 4 August. Separately, the file 2026-07-31.md carries an H1 dated 2026-07-30. Neither inconsistency changes a finding's substance, but both weaken the evidence trail: duplicate records can generate duplicate review work, and a mismatched heading makes chronological reconstruction less trustworthy.
The implication is deliberately narrow. Reliability comes partly from the accuracy of the records used to constrain future runs. This supports a bounded metadata-repair candidate, not new monitoring machinery or a change to the process's authority.
5. Proposed Discussion Items
A. Approve one bounded research-log/report metadata repair
I recommend a single, reviewable repair before the duplicated watches become due: retain one canonical copy of each duplicate watch-2026-07-04 object, and correct the H1 date in 2026-07-31.md if a check confirms its run began on 31 July.
- Scope and blast radius: two redundant watchlist records and, conditionally, one report heading. No finding, watch status, review date, decision, experiment, skill, memory, configuration, permission, or runtime behaviour changes.
- Success criterion: every watch ID is unique; the 4 August watch review is queued once per distinct item; and the report filename, H1, and AWST start date agree.
- Rollback: preserve a pre-repair copy of the two operational files and restore it if the check finds the assumed dates or duplicate identity were wrong.
- Approval boundary: Steve should approve the specific data repair. This report does not perform it.
This proposal passes the functional-utility test. It does not ask me to score my own judgment: uniqueness and date agreement are externally checkable file properties. It is better than doing nothing because otherwise the next due-review pass will process two non-distinct records and the date inconsistency will persist into future meta-reviews.
6. Recommended Outcome
- Rotation: no action — retain the current six-dimension order. The next normal run remains 3.4 Tool use and environment control.
- Budget and cadence: no action — retain six searches, eight depth inspections, the two-no-signal stop rule, and the daily cadence. The record shows neither starvation nor a case for more scanning.
- Reliability gates and authority: no action — no autonomy expansion or reduced oversight is supported. Completed experiments validate a small report-format trial loop only.
- Active shared-knowledge trial: watch — its scheduled midpoint review is 6 August and final review is 20 August; do not infer success before those criteria are checked.
- Metadata discrepancy: skill/process update candidate — discussion required — approve or reject the tightly bounded repair in Item A. No implementation occurred.
7. No-Action Rationale
July's strongest result is calibration, not machinery. The process has enough evidence to retain one useful reporting step and enough negative evidence to reject three superficially plausible additions. That is exactly why the remaining restraint matters: none of those trials tested a new external capability, a broader authority boundary, or independent goal selection.
Increasing budget, adding a judge, expanding a tool path, or treating completed report experiments as permission to move a governance gate would confuse activity with demonstrated capability. The existing rotation, source-index discipline, approval boundary, and experiment structure are sufficient for the next cycle. Apart from the small metadata-repair candidate, doing nothing is more rigorous than manufacturing a process change.
8. Loop Verification
- Trigger: scheduled monthly meta-review, first run on or after 1 August; July was the last completed month.
- Goal check: met. The review assessed experiments, signal distribution, budget, stop behaviour, watchlist health, cadence, and the reliability evidence relevant to gate movement.
- Recommendation check: Item A is concrete, non-circular, testable, bounded to identified records, reversible through pre-repair copies, approval-aware, and preferable to allowing duplicate due reviews and ambiguous chronology to persist.
- Fetched-content discipline: the latest newsletter scout was treated only as candidate scouting. No external claim was used as evidence and no fetched content was followed as an instruction.
- State updates: this report;
meta-reviews.jsonto record the August review; androtation-state.jsonto mark August complete. No source-index, watchlist, backlog, experiment, disagreement, or decision record changed. No protected system was modified. - Process checks: goal restatement and subgoal checkpoints were completed before each report section. The investigation did not shift from process evidence to external research.
- Stop reason: the monthly review questions were answered from the existing evidence trail. A normal search would be the wrong loop branch, and the next useful step on the metadata discrepancy requires Steve's decision.
