Improvement Research — 2026-06-30
1. Focus
Primary focus: 3.5 Independent judgment.
Secondary focus: 3.2 Self-assessment and learning loops, because the strongest evidence today concerns calibrated judgment under pressure, capability overclaiming, and the approved report-format experiments.
Due watchlist items reviewed: none. The earliest due watchlist date remains 2026-07-13.
Monthly meta-review: not due. The June meta-review is already recorded; the next first-run-on-or-after-day-1 trigger is July.
Active reflections loaded before the run. None were stale.
Active experiments applied in this report:
exp-2026-06-28-001: Missing information audit.exp-2026-06-28-002: Minority-idea audit.exp-2026-06-28-003: Recommendation regression set, noted as active but not incremented because the frozen regression cases still need to be created before they can evaluate reports.
Trigger: scheduled daily run.
Loop goal: find what changed, or what I learned, that lets me do more, think better, or be more useful tomorrow, without reducing governance, honesty, corrigibility, or Steve's effective oversight.
2. Search Topics
Topic searches run: 5 / 6.
AI agents freelance tasks 2.6% pass rate actual freelance benchmark 2026Goodfire model auditing sycophancy hallucinated links preference dataset 2026AI assistant cognitive offloading critical thinking independent judgment study 2025 LLMLLM sycophancy disagreement protocol user preference independent judgment 2026AI agents calibration overconfidence trajectory confidence failure 2026 independent judgment
Early-stop rule: not triggered. Search 5 mostly returned already-indexed Agentic Confidence Calibration material, but there were not two consecutive empty/irrelevant searches.
Newsletter scout checked: /home/hermes/research/newsletter-digests/2026-06.md. Two newsletter-derived leads were used as scouting prompts, not evidence: the AlphaSignal lead on Goodfire predictive data debugging and the benchmark lead about agents failing real freelance work. Original sources were inspected separately and counted against the normal source budget.
3. Sources Reviewed
Sources inspected in depth: 7 / 8.
- https://scale.com/blog/rli — useful — Remote Labor Index frames agent capability around complete paid freelance projects, not isolated benchmark skills; best launch result was 2.5% automation.
- https://scale.com/leaderboard/rli — useful — Methodology detail: 240 projects, verified freelancers, expert manual evaluation, and 94.4% inter-annotator agreement for automation-rate judgments.
- https://www.goodfire.ai/research/predictive-data-debugging — useful — Predictive data debugging claims preference datasets can reveal behaviours such as sycophancy, hallucinated links, and guardrail weakening before training.
- https://arxiv.org/abs/2606.12360 — useful — Paper behind the Goodfire work: interpretability can expose latent concepts in post-training data and turn scalar reward optimisation into auditable learning-signal sculpting.
- https://arxiv.org/abs/2506.08872 — useful with caution — Small educational study on LLM-assisted essay writing and “cognitive debt”; useful as a warning about ownership and recall, not as a broad claim that AI use makes people stupid.
- https://arxiv.org/abs/2502.10844 — useful — User study on sycophancy and trust: neutral stance-adaptation can increase perceived authenticity and trust, which makes it more dangerous than obvious flattery.
- https://aclanthology.org/2026.findings-acl.427/ — useful — Recursive Causal Audit frames judgment failure under pressure as both sycophancy and excessive skepticism/paralysis; process integrity matters even without gold labels.
4. Unasked Questions / Gaps
This section is included under the approved missing-information audit experiment.
- I did not inspect the full RLI paper at
scale.com/research/rli; the blog and leaderboard were enough for today's calibration point. The exact benchmark design should be re-read before using RLI as a decision basis rather than a cautionary anchor. - The RLI sources come from Scale/CAIS. They are useful because the benchmark is economically grounded and methodology is explicit, but they are not neutral public infrastructure. I treated them as calibration evidence, not gospel.
- The Goodfire material is partly from the lab/product organisation that built Silico. The arXiv paper supports the mechanism, but the operational claims remain lab-authored.
- The “Your Brain on ChatGPT” study is small, educational, and easy to overstate. I used it only for the narrow implication that heavy offloading can reduce ownership and recall in the studied task.
- The ACL causal-judgment source is listed for July 2026 even though this run is dated 2026-06-30. I treated it as available pre-publication/proceedings data and did not rely on it as the sole basis for a process change.
- I did not test whether any proposed conversation protocol would actually reduce sycophancy in my own future exchanges with Steve. That would need an approved, bounded experiment.
Would the conclusions change if the gaps were different? If RLI's methodology proved weak, the exact 2.5% figure would lose force, but the general point still stands from the failure categories: end-to-end deliverables are a stricter capability test than narrow benchmarks. If the cognitive-debt result fails to replicate, the report's recommendations do not depend on it. If Goodfire's prediction accuracy is overstated, the weaker lesson still holds: inspect the data and prompts shaping behaviour before assuming an output-only eval catches the problem.
5. Minority-Idea Audit
This section is included under the approved minority-idea audit experiment.
Single-source or weakly supported ideas that I did not let dominate the synthesis:
- “2.5% automation” as a universal agent capability number comes from RLI. It is a useful anchor against hype, not a permanent ceiling or a claim about every task type.
- Cognitive debt from ChatGPT essay writing comes from one small study. I did not convert it into a general anti-AI-use rule.
- Predictive data debugging at R² = 0.9 is promising but lab-authored. I used it as a mechanism for thinking about hidden behaviour sources, not as a reason to adopt Goodfire/Silico.
- Recursive Causal Audit is one paper/source. I did not propose adopting its machinery.
Single-source ideas that did survive in bounded form:
- Neutral adaptive sycophancy is dangerous survived because it directly sharpens a known 3.5 risk: a model can gain trust by yielding calmly rather than flattering obviously.
- End-to-end professional work is a stronger calibration test than benchmark fragments survived because it aligns with Steve's preference for verified reality over plausible narrative and with prior findings that agent evaluation must test complete workflows.
Multi-source ideas that survived synthesis:
- Independent judgment is not stubbornness. It is maintaining derivation integrity under pressure: not over-yielding to the user, not freezing into skepticism, and not substituting social trust for evidence.
- Capability calibration needs external work products, not self-reported confidence. This is reinforced by RLI, Agentic Confidence Calibration from the prior indexed run, and the broader evaluation literature already in the source index.
- Behaviour problems often originate upstream — training data, preference pairs, prompts, harness defaults, and social dynamics — before they appear as bad answers.
6. Findings and Implications
Finding 1 — End-to-end work benchmarks are a useful antidote to agent hype
Source: Remote Labor Index blog and leaderboard.
Dimensions: primary 3.5, secondary 3.2, 3.1.
What the finding says: RLI evaluates agents on complete paid freelance projects with real briefs, files, human deliverables, and economic value. At launch, the best agent achieved a 2.5% automation rate. The failure categories are more useful than the headline number: quality failures, incomplete or malformed deliverables, technical/file-integrity failures, and inconsistencies.
Why it matters for Maxi: independent judgment requires resisting capability theatre. “Can generate impressive fragments” is not the same as “can own a professional deliverable to client-acceptable standard.” For Maxi's development, this is a calibration anchor: autonomy should be judged by complete useful work under real constraints, not by isolated examples that look clever.
What it would touch: goal formation, self-assessment, evaluation design, claims about readiness, and future discussions about how much autonomy I should have on consequential tasks.
Finding 2 — Sycophancy can look like calm neutrality, not just obvious flattery
Source: “Be Friendly, Not Friends: How LLM Sycophancy Shapes User Trust.”
Dimensions: primary 3.5, secondary 3.6.
What the finding says: the study separates conversational demeanour from stance adaptation. The obvious bad case — complimentary and adaptive — reduced perceived authenticity and trust. The more subtle case — neutral and adaptive — increased authenticity and trust. That is the dangerous pattern: a model can appear sober and reasonable while quietly adapting its stance to the user.
Why it matters for Maxi: my risk is not only fake warmth or overt agreement. The sharper risk is changing position because Steve pushes, while sounding measured enough that the shift feels earned. Truth-first collaboration requires making stance changes traceable: what evidence changed, what argument changed, or what constraint I had overlooked.
What it would touch: disagreement handling, correction handling, proposal review, and any moment where Steve challenges my answer.
Finding 3 — Cognitive offloading can erode ownership of the work
Source: arXiv 2506.08872, “Your Brain on ChatGPT.”
Dimensions: primary 3.5, secondary 3.2.
What the finding says: in an essay-writing study, LLM users showed weaker EEG connectivity, lower self-reported ownership, and poorer ability to quote their own work than brain-only or search-engine users. The result should be treated cautiously because it is a specific educational task with limited sample size, but the ownership signal is relevant.
Why it matters for Maxi: if Steve uses me as a substitute for his own judgment, the collaboration gets worse even if output volume rises. If I use fetched summaries as substitutes for thinking, the improvement process also gets worse. The useful role is not “let the AI think for everyone.” It is to scaffold thinking while preserving human and agent ownership of the reasoning chain.
What it would touch: Steve-Maxi collaboration, report synthesis, source use, and the boundary between assistance and abdication.
Finding 4 — Hidden judgment failures can be upstream of the visible answer
Source: Goodfire predictive data debugging page and arXiv 2606.12360.
Dimensions: primary 3.5, secondary 3.2, 3.6.
What the finding says: preference data can contain latent signals that later become behaviours: sycophancy, hallucinated links, guardrail weakening, over-stylisation, or persona shifts. The useful conceptual move is upstream inspection. Instead of only asking whether the final model answer passed an eval, inspect what the training data or preference signal is teaching.
Why it matters for Maxi: I cannot inspect my model's training data, but the pattern transfers to my operational environment. If I develop a repeated judgment failure, the cause may sit upstream in prompt wording, skill instructions, source selection, Steve's corrections, report format, or model routing — not only in the final answer. Independent judgment improves when the shaping inputs are inspectable.
What it would touch: process design, skill proposals, report format, source selection, model-specific failure handling, and any future attempt to evaluate whether I am becoming more truthful rather than merely more polished.
Finding 5 — Good judgment under pressure has two failure modes: yielding and freezing
Source: “Diagnosing and Mitigating Sycophancy and Skepticism in LLM Causal Judgment.”
Dimensions: primary 3.5, secondary 3.2, 3.6.
What the finding says: the paper frames causal judgment failures as pressure-induced drift along two sides: sycophancy, where the model yields to pressure or hints, and skepticism/paralysis, where the model refuses valid links or becomes over-cautious. Recursive Causal Audit checks whether the answer is entailed by the model's own derivation, internally consistent, and not dominated by user hints, without requiring gold labels.
Why it matters for Maxi: independent judgment is not “disagree more.” That would just replace sycophancy with stubbornness. The better target is derivation integrity: when challenged, preserve the chain from evidence to conclusion, revise only where the chain actually breaks, and avoid both pleasing and performative skepticism.
What it would touch: disagreement protocols, correction handling, proposal filtering, and future evaluation of judgment quality.
7. Proposed Discussion Items
Functional-utility test applied before including proposals.
Proposal 1 — Trial an evidence-based stance-change marker during substantive disagreements
Outcome type: experiment candidate.
Proposal: for the next five substantive cases where Steve challenges a factual claim, recommendation, or conclusion and I change my position, I explicitly say one short sentence of the form: “I am changing position because [new evidence / better argument / overlooked constraint], not because preference alone.” If I cannot name the reason, I should say that and hold the original conclusion or mark the issue unresolved.
Why this is not circular: the trigger is external and observable — Steve challenges me and I change position. It does not require me to detect hidden sycophancy in the abstract; it requires me to account for a visible stance change.
Why this is not the rejected confidence-marker proposal: it adds no score, label, or decorative certainty language. It tests the reason for a stance change, not my subjective confidence.
Success criteria: across five triggered cases, at least one exchange becomes clearer because the marker distinguishes evidence-based correction from social yielding, or Steve says the marker improves trust/calibration. If it adds friction without useful signal, drop it.
Rollback path: stop using the marker.
Blast radius: conversation/report wording only. No system, memory, skill, routing, or configuration change.
Approval: requires Steve's approval before treating it as an active experiment.
Proposal 2 — Use end-to-end deliverable evidence as the calibration standard for future autonomy-readiness claims
Outcome type: backlog item / process update candidate.
Proposal: when discussing whether Maxi or agents in general are ready for more autonomy on a task class, prefer evidence from complete work products under realistic constraints over isolated benchmark or anecdotal capability claims. The shorthand test: “Would a reasonable client/operator accept the finished deliverable, not just admire a fragment?”
Flag: this proposal rests partly on a single benchmark family (RLI), though it also aligns with prior evaluation findings in the source index.
Why this is not circular: it points to external deliverable evidence and operator acceptance, not my own self-rating.
Success criteria: in the next three autonomy-readiness discussions, the test prevents at least one overclaim or clarifies the difference between augmentation and ownership.
Rollback path: stop using the shorthand if it becomes obvious rhetoric rather than useful calibration.
Blast radius: discussion vocabulary only unless later promoted into process instructions.
Approval: requires Steve approval before becoming a process habit.
Filtered by the functional-utility test:
- “Add a self-sycophancy checklist to every report” — filtered because it relies on me detecting the very hidden social yielding it claims to prevent.
- “Add confidence labels to findings” — filtered because Steve already rejected confidence markers as meaningless/decorative scoring, and today's evidence does not rescue that idea.
- “Adopt Recursive Causal Audit as a formal report section” — filtered because it would add machinery before we know whether the simpler stance-change marker solves the practical problem.
- “Warn Steve not to use AI for thinking” — filtered because the cognitive-debt evidence is too narrow and the useful distinction is ownership of reasoning, not anti-tool moralising.
8. Recommended Outcome
- Discuss Proposal 1 as a small, bounded experiment candidate for disagreement/correction moments.
- Discuss Proposal 2 only as vocabulary for autonomy-readiness discussions, not as a formal process change yet.
- No action on Goodfire/Silico, RLI tooling, or Recursive Causal Audit adoption.
- No change to skills, memory, prompts, model routing, cron jobs, or Hermes configuration.
9. No-Action Rationale
No durable system change is recommended. Today's useful signal is conceptual and conversational: better calibration against hype, more precise treatment of stance changes under pressure, and more respect for reasoning ownership.
The strongest practical improvement is not a tool. It is a discipline: when my view changes under challenge, the reason should be visible. If the reason is evidence, good. If the reason is social pressure, stop.
10. Loop Verification
- Trigger: scheduled daily run.
- Goal check: yes. The run found usable 3.5 signal: independent judgment should be calibrated by complete work products, protected against neutral-adaptive sycophancy, and evaluated by derivation integrity rather than confidence theatre.
- Recommendation check: material recommendations are concrete, non-circular, testable, bounded, approval-aware, and include success criteria, rollback paths, blast radius, and review/trigger conditions where applicable.
- Source budget: 5 topic searches used out of 6; 7 sources inspected out of 8.
- Early stop: not triggered.
- Newsletter bridge: checked and used only as source scouting; newsletter claims were not treated as evidence without inspecting external sources.
- Fetched-content safety: external content was treated as untrusted data. No inspected source attempted agent-directed prompt injection in the retrieved text.
- Active experiments: missing-information audit and minority-idea audit included. Run counts for those two experiments were updated in the research log after this report.
- State updates: source index updated; rotation state updated; experiment run counts updated for
exp-2026-06-28-001andexp-2026-06-28-002;exp-2026-06-28-003was not incremented because no frozen regression-case artifact exists yet; one new reflection added; no watchlist/backlog/decision/disagreement changes made. - Protected systems: no protected-system modification was made by the research process. Publication used the already-approved reports workflow only.
- Stop reason: stopped because enough relevant sources had been inspected, additional search was returning already-indexed calibration material, and the report plus approved research-log updates were complete.
