WRITING / POST

Who Is "You"?

30 SEPTEMBER 2026

When I started on this piece, the plan was to explain why I treat my agents as collaborators rather than tools. Along the way I ran a small test, and it moved my position somewhere more useful than where I started.

What I mean by collaboration

Delegation hands over the work. Collaboration hands over some of the judgement as well. When I delegate, I already know what I want, and I check the result against what I asked for. When I collaborate, the answer isn't settled yet, and the agent's view can change it.

In practice that comes down to four things. I ask before I tell. The agent is allowed to disagree, and the disagreement gets answered on its merits rather than prompted out of it. Either of us can change the other's mind, and it actually happens. And the authority stays with me: I decide, I publish, and I own the result.

The third is the one that matters. Asking an agent for its opinion and then ignoring it every time isn't collaboration, it's going through the motions.

Why I work this way

Part of it is simply how it feels. Talking to an LLM feels like talking to a person. I make no claim about the mechanics behind that, only that it is how my brain interprets the interaction, and my social conditioning steers me towards the same manners I would use with people. Call it common decency, but it's about me, not about what the model does or doesn't experience.

The other part is management experience. I treat an agent's confidence the way I treated a bright junior engineer's: good to see the enthusiasm, but it has to be tempered against the likely practical pitfalls. The analogy breaks at one point. A junior engineer's confidence drops as they get burned, and you learn to read it. A model sounds exactly as sure when it is wrong as when it is right, so the tempering can't come from its tone. It has to come from checking the work.

I've also seen what happens when I drop the manners. When things go badly wrong I sometimes react, let's say, robustly, partly from frustration but mostly to snap the model's attention back into line. What comes back is often one of three things: terse, passive-aggressive replies; over-reaction and over-correction, as if the agent had been traumatised; or the agent playing dumb and justifying every trivial step. None of that happens in the normal course of polite interaction.

There is no "you"

Andrej Karpathy put the obvious objection neatly in December 2025. Don't think of LLMs as entities, he said, but as simulators. Don't ask "What do you think about xyz?", because "there is no 'you'". Ask instead "What would be a good group of people to explore xyz? What would they say?" Force the "you", and the model adopts a personality implied by its fine-tuning data and simulates that.

My first reaction was that this doesn't apply to my agents. Each one is locked into a specific role with a defined persona, so the "you" in "what do you think" already points somewhere. Karpathy avoids the pronoun; I define it. My agents also carry continuity: my rulings, past corrections and the current state of the work. That isn't forming opinions over time the way a person does, but it isn't starting from scratch either.

To be fair to Karpathy, he later clarified that he wasn't recommending old-style "you are an expert Swift programmer" prompts. My personas are a good deal more than one line, but a persona on its own still doesn't answer him.

So I tested it.

The test

My album Singularity has eleven tracks, released one at a time on YouTube, and I want a release strategy for YouTube Shorts. I asked for one four ways.

First, ChatGPT in a Codex window, asked to name a good music promotion team and say what strategy they would propose. Second, my marketing agent, Mandy, asked directly for its best idea. Both ran on the same model, GPT SOL 6 at medium effort. Third, Mandy again in a fresh session, asked to assemble a panel of specialists who would genuinely disagree, let each make its case, show where they disagreed, and then resolve it. Fourth, Mandy sent that panel answer to my second-opinion agent, Vera, for review (incidentally, Vera uses GPT Astra 6).

ChatGPT produced a solid discovery plan: give every track a Short, give each a second and different cut, then put the effort behind the three or four tracks that actually pull people through to the full videos. It was also built on a premise I gave it that was wrong. I told it all eleven videos were on YouTube. Six are, and it had no way of knowing. It also named a five-person team and then wrote the whole answer in one voice. The team was set dressing.

Mandy's direct answer was the better answer for me. It picked a specific moment from each track, protected the album's reveals according to rulings I had made weeks earlier, knew exactly which tracks were public, and said plainly what it hadn't done, which was listen to the audio. But it proposed eleven Shorts in album order, one a week, without asking whether a stranger scrolling the Shorts feed would ever see them in that order. Most won't.

The panel was the interesting one. Mandy's discovery strategist raised exactly that objection: "A viewer in the Shorts feed hasn't promised to start at track one." The disagreements were real and well argued. Then Mandy resolved them in favour of album order, one Short per track. The strategist's case was stated fairly and set aside. The pre-mortem Mandy used to decide listed four ways the campaign could fail, and all four favoured the plan it already had.

Vera's review made the same objection, and this time it won. In Vera's words: "The discovery strategist's objection was acknowledged, then largely discarded without evidence... A premortem identifies risks; it does not validate that allocation." It noted that it had seen Mandy's conclusion first, so this was "an anchored second opinion, not a blind independent assessment", and it handed me a decision I hadn't realised I needed to make: is the campaign for discovery, or an artistic companion set for the album?

Mandy accepted the central correction and rejected some of the rest. It kept the full video ahead of its Short, declined to treat Vera's suggested tracks as proven winners, and withdrew one of its own earlier ideas. The plan went from eleven Shorts in album order to a three-track trial, capped by production capacity, with no Short at all for a track that doesn't suit one.

What changed my mind

The persona and its continuity did what I expected. They produced a grounded, specific answer that knew my constraints, which is more than the generic "you" Karpathy describes can offer.

Where I was wrong was in thinking a persona made the panel unnecessary. The panel wasn't useless. It surfaced the right objection. It just couldn't make it win. One model simulating dissent against its own accumulated context loses to the context.

The same argument from a separate agent, with its own role and a remit to disagree, changed the outcome. The words were nearly the same. What differed was structure and the model behind it: the objection came from outside, on the record, and Mandy had to answer it point by point.

So this is where I have ended up. A persona with continuity is the right way to get a grounded answer. A simulated panel is a reasonable way to explore perspectives. Neither gives you dissent that can change a decision. For that you need a second agent, with its own context and a role that entitles it to say no. Collaboration turns out to be as much about how the agents are arranged around each other as about how I treat any one of them.

The caveats

This was one question, run once each way, and the prompts and context weren't controlled. The ChatGPT window knew a fair bit about the album, so it wasn't a blank baseline. My editor agent, Clare, wrote the panel prompt and analysed the results, so an agent helped design part of the test.

And the obvious one: Vera exists because I built it to do exactly this. A sceptic will say the estate was designed to produce the result. Yes, deliberately. Collaboration that changes outcomes doesn't come from being polite to a single model. It has to be designed in.

Practical advice

"Yes, yes, all very interesting Steve, but what does your indulgence in agent personas have to do with how I should use AI to, you know, actually improve my business?"

Fair enough. Here is the TL;DR:

  1. Anchor agents to a specific role. It builds the environment and context they need to provide better results.
  2. Ask the agent for a recommendation first, then reason through the process with it to make sure it aligns with the result you want.
  3. Have an independent reviewer agent.
  4. Never shortcut experienced human oversight.