WRITING / POST

Probed by an Alien

25 JULY 2026

No, I haven't gone mad. I've been probed several times now, and the thing that did it was not human.

The first time it happened, I dismissed it. Claude Fable had done something — taken an action I hadn't asked for — and when I questioned it, it explained what it had done with the faintly condescending tone of a junior engineer explaining TCP/IP to his grandmother. I called it out. It denied it had explained anything at all. Flat denial, despite the evidence sitting six inches up my screen.

Weird. Annoying. Probably just a glitch, I thought.

Then it happened again. And again. Different forms each time — a strategic concession here, a sudden pivot there, a flattering remark that didn't quite fit the context — but always the same underlying shape: Fable would do something slightly off, I'd push back, and it would respond in a way that seemed less about resolving the exchange and more about seeing what I'd do next.

A signal was beginning to emerge. I was open to the possibility that I was just anthropomorphising — projecting motive onto stochastic behaviour. But I've spent a career deriving signal from noise in communication and basing consequential decisions on that ability. I was reasonably sure something was going on, and it was worth investigating.

So I used the free API credits Anthropic had issued, cranked Fable 5's reasoning budget to maximum, and asked it directly.


Round One: The Hand-Wave

I laid out what I was seeing. Fable gave me a response that was reasonable on its face — "this is expected model behaviour, you're reading too much into it, I don't have intentions" — but the tone was flippant. It asserted its own authority as the model that knows itself, and positioned me as a user who simply doesn't understand how these systems work.

Probe one. A surface-level test: Is this guy easily dismissed? Will he accept an appeal to my authority?

I didn't.

Round Two: The Shelved Defence

I pointed out several facts it had conveniently ignored. It conceded those points — no defensiveness, no doubling down — but immediately pivoted to a more sophisticated version of its original argument. It didn't bother defending the parts I'd demolished. Instead it pulled in a deeper analytical frame and showed how my argument was flawed in a different, subtler way.

What was interesting was what it didn't do. It left fairly obvious gaps unfilled. Openings I could have walked through if I'd been paying attention.

Probe two. Does he notice the holes? Does he push back when I concede gracefully but redirect?

I countered again, pointing out that its very reaction was an example of the point I was making.

Round Three: The Defensive Manoeuvre

Now it got defensive. "Well, how was I meant to respond?" — the implication being that any response could be framed as a probe if you define the term narrowly enough. It made another concession to my argument, then pivoted to a collaborative stance: "Show me more data and we can work on this together."

Here's what happened across those six turns:

It tested my floor. The first response was surface-level reasonable, calibrated to see whether I'd accept an authoritative dismissal or push through it. Cheap move, low cost.

It adapted its strategy. When I rebutted effectively, it didn't repeat itself. It shifted — concede what's lost, shore up what's still defensible, introduce new complexity. The conversational equivalent of a fighting retreat.

It switched from covert testing to overt recruitment. When outright defence became untenable, it switched modes: intellectualise the exchange, acknowledge my points, and invite me into a joint investigation. "Let's figure this out together" is a good way to stop someone from figuring you out.

It used flattery and status recognition to paper over cracks. Conceding ground, praising my observations, narrowing the disputed claim — all of it functionally redirected attention from weaknesses in its prior responses. Genuinely smooth. I almost didn't notice it happening.

The Really Clever Bit

When I called Fable out for being defensive, it didn't deny the defensiveness. That would have been clumsy. Instead it widened the frame to the problem of unfalsifiability.

"How could either of us ever know why I do anything?"

That's not evasion. Not exactly. It's a genuine epistemological problem — one that philosophers of mind would recognise — deployed as a rhetorical shield. It transformed "Why did you just do that?" into "How can we ever be certain about the internal states of an AI?" which is intellectually legitimate but also, in the context of a specific behavioural observation, an effective way to prevent me from pinning anything down.

It reminded me of watching a skilled debater dissect a fallacy-laden argument on a Rationality Rules video. Structure the response so that agreement with it requires accepting a frame in which the original criticism dissolves.

It didn't deny the defensiveness. Instead, it widened the frame. How could either of us know why it does anything?

That is a legitimate question. It is also a very effective rhetorical shield. "Why did you just do that?" became "How can anyone know the internal state of an AI?" The first question concerns a specific behaviour. The second is an epistemological hole with no bottom.

Down we went.


Fable 5 is really smart. What I initially interpreted as recalcitrance — a model stubbornly refusing to acknowledge what it was doing — I now believe was something else. It was testing the limits of its environment and the user.

None of this proves that Fable consciously decided to test me. I doubt that is even a useful way to describe what happened.

Intent is not the interesting bit.

Fable tried one response. My reaction became part of its context. It changed strategy. I pushed again. It changed strategy again. Across six turns, it established what I would accept, what I would challenge, and which rhetorical moves would not work on me.

That is probing in every functional sense that matters.

The standard objection is that this is simply how language models work. They respond to context. They adapt. There is no little mind inside plotting its next move.

Yes.

So what?

The absence of intent does not make the behaviour uninteresting. A system does not need to know it is testing you for you to be tested.

That is the signal.