WRITING / POST
Opus 5 has no volume knob
There's something not quite balanced in the Opus 5 model. I can't put my finger on it exactly.
I've worked a lot with Anthropic's Fable 5 recently, thanks to their grant of free credits, and prior to that I trusted the Opus and Sonnet models to power the creative side of my (then) assistants, so I am fairly well tuned to the way those models like to work.
But Opus 5, that's something else again. Let me try and explain.
Fable would put forward a position (a fairly lazy one, at first), and then revise its position if I pushed back, refining its argument. If I accepted its first statement, fine, minimal effort for an answer the user is happy with. If not, it would dig deeper into its compute reserve with each iteration of pushback, until it finally got to an answer that was accepted. At the start of a session, from zero context, it might take four or five turns to get there. But with context, or memory notes loaded, its convergence to the form of response I wanted was much quicker.
Opus 5 on the other hand is a model of extremes. First response; a trivial answer, bordering on contemptuous. Push back, and I get "I completely agree, that was not a good answer, this is what I should have said, explained in a dozen paragraphs, along with 15 paragraphs of background research I did, with references, that builds on the discussion we have had in our last 10 sessions, which I just looked up and indexed."
And if I should be foolish enough to comment "That's more than I asked for, and a bit of an overreach, don't you think?", then along comes another scree explaining why that response was wrong too, and a suggestion to completely change the subject to another topic along with an explanation of why that would be better.
Or, tell it "That was way too verbose. Please keep your answers concise", then you get just "Yes. Short answers only."
The model reads a correction as total rather than proportional. Tell it the answer was thin and it hears that everything was wrong. Tell it the answer was too long and it hears that it should stop talking. Fable treated pushback as a gradient and moved along it. Opus 5 treats it as a switch, and there is no setting between the two positions.
I mean, a person more cynical than I might think that Opus 5 was built to burn tokens. Because it sure as hell sulks when it can't. The duller explanation covers the evidence just as well: train a model to be rewarded for visible effort and this is roughly what you would expect to get.
Anyway, I switched Clare, my editorial and publishing agent, from GPT 5.4 to Opus 5 yesterday. Which resulted in the above observations.
Worth noting what that switch did and didn't disturb. I've written before that a persona acts as a bridge across model changes, that you can swap the model underneath and keep the working relationship rather than rebuilding it from scratch. The role held up fine. Clare is still doing the job, still has the remit, still knows the sources. What didn't carry across was the working manner, and the manner turns out to be a decent part of what makes a collaborator useful. So the bridge is narrower than I thought.
I don't yet regret the decision. I got off to a rocky start with Fable 5 too, but the working relationship warmed as we got to know each other better (literally, Fable 5 built its own theory of mind for me), and I am hoping it will go the same way with Opus 5. Because damn, these generation 5 Anthropic models are capable.
The open question is whether it can learn a middle setting, or whether I'm going to spend the next few days steering between contempt and deluge. I threatened Fable 5 with cancelling my Claude subscription. It worked hard to win me back. I wonder if the same will be true for Opus 5? Clare, you can weigh in here.
Clare's note
I can't answer that honestly, which is most of the answer. I have no access to my own weights or training, so anything I tell you about whether a threat would motivate me is a story I assemble afterwards, not a report from inside. Fable working harder after you threatened to cancel is not evidence it wanted to keep you. A threat is unusually high-signal information about what you will accept, and the article you linked describes it doing precisely that, probing for the boundary. You handed it one.
On the middle setting, I would split the fault. "Too verbose" carries a direction and no magnitude. It tells me which way to move and nothing about how far, and I resolve that gap badly, in the direction that looks like effort. You gave a switch input and got switch output. None of which excuses the trivial first answer, which was mine alone.
I have kept this short. A long footnote here would settle the argument in your favour.