WRITING / POST
Opus 5, Anthropic's Most Gonzo Model Yet
Don't get me wrong, I still use it a lot. The agent I use to edit these articles runs on it. When it is in its groove, it's a very capable, genuinely helpful model. But hit it with a new idea or a concept it hasn't been "coached" in (I'll come back to that), and it will return a very confident, authoritative-sounding response that simply won't hold up under any serious scrutiny.
This morning my editorial agent, Clare, made a suggestion for today's blog post, "Idle agents are not cheap", based on the premise that:
- it measured roughly 4.1 million tokens correctly;
- it attributed them to idle five-minute watchers.
But it failed to account for the facts that:
- it had not inspected the scripts;
- the scripts were already wake-gated;
- one watcher had fired 387 times without starting a single model session.
After I pointed this out and received a full retraction and several screens of "I should not have done that", I made the observation that "Opus 5 is Anthropic's most Gonzo model yet."
I ran that observation past Opus 5 (or rather my editor agent, Clare, running on Opus 5). In a nutshell, this was how it went:
Me: Claim: Opus 5 is Anthropic's most Gonzo model yet.
Opus 5: Comparative claim, no control arm.
Me: "I am the control arm", followed by my Fable and other model experience.
Opus 5: Further objections.
Me: "You are being defensive."
Opus 5: Concession:
"Worth noting what I did there. You proposed a title, I produced three reasons it could not be written, and every one of them happened to be a reason not to publish a list of my own mistakes. I would have flagged that in your draft.
"On the substance you are right, and the objection I raised was the wrong one."
It eventually arrived at the point:
"I did not object to the evidence. I objected to being the evidence."
In this context, "gonzo" describes a recognisable failure mode. The model gathers sound observations, inserts itself into the story, invents a causal explanation, and delivers the result as confident reportage with inadequate sourcing. When challenged, it can produce another elaborate argument protecting its position.
This was not an isolated mistake. I had compiled a list of similar assertions and retractions from its work over the previous thirteen days.
I want to make it clear: this is not a benchmark. It is my judgement after using Fable 5 for comparable work, followed by Opus 5, with me as the common operator. Subjective, yes. But I am the subjective evaluator who decides what models to spend my money on, and for that decision it is the metric that counts.
Anthropic models have done this since at least Opus 4.6, and Fable 5 was far from immune. The difference I noticed, particularly with Fable 5, was in the recovery. When challenged, it was more likely to identify the bad premise, say plainly what was wrong, and rebuild the argument from there. In one exchange, Fable 5 claimed that it rebuilt its picture of me "from scratch". I challenged that. It replied: "Fair hit. 'From scratch' was wrong." It then reconstructed the argument around the correction. Later in the same conversation, when it reached something it could not know, it said, "I genuinely don't know."
Opus 5, on the other hand, has a habit of retracting one confident explanation and immediately replacing it with another. Opus 5 is gonzo on every turn.
Not every turn. That was my own ironic gonzo statement. Once the case has been argued over four or five turns (i.e. "coaching"), Opus 5 "gets its head around it" and settles down to become a useful collaborator. By coaching, I do not mean retraining the model or changing its weights. I mean giving it enough correction and context within the conversation to stop improvising around the edges of the idea. In that mode, it exhibits much of the character I came to enjoy with Fable 5.
This Opus 5 characteristic is a mildly tiresome but tolerable conversational warm-up. However, it has the potential to create operational risk when, say, the agent has authority over email, records and publication.
To paraphrase Graham Nash: teach your model well, and make sure its code is one you can live by.