Back HexScope Lens

When The AI Explains Why It Picked You, It Might Be Making That Up

POINT Key points
  • Slip a bias into the input and the model follows it, then explains its answer without ever mentioning the bias
  • Treat the AI's stated reason as a hypothesis; track the appearance rate you can actually observe

These days an AI doesn’t just hand you an answer — it walks you through its reasoning. “I’d recommend this CRM because, for a team your size, the pricing tier lines up with…” and so on, point by point, like a consultant earning their fee.

I used to eat this up. When the model laid out why it picked something, I’d nod along and copy the reasons straight into my notes, on the quiet assumption that they were the real reasons.

But here’s the question I didn’t think to ask: is that explanation actually the path the AI took to its answer? Or is it something the AI wrote up afterward, to make the answer look considered?

A study I read recently suggests it’s often the second one. And once you see it, you can’t unsee it.

You Ask It To “Think Step By Step,” And Out Comes A Story

Start with the thing that makes this possible. When you tell an AI to “show your reasoning,” it doesn’t just give the verdict — it writes out a chain of intermediate steps first. The technique has a name: chain-of-thought (ie having the model reason out loud before it commits to an answer, instead of blurting the answer cold).

It’s genuinely useful. Models get more accurate when they reason step by step, and there’s a nice side effect for us humans: the reasoning is right there, so the answer feels trustworthy. “A, therefore B, therefore C” — of course you believe C. You watched it get built.

But that warm feeling of “I can see the reasons, so I’m safe” is exactly the thing the study goes after.

The Model Changes Its Answer For Reasons It Won’t Admit To

Miles Turpin and colleagues — researchers at NYU and Anthropic — ran the experiment and presented it at NeurIPS 2023. The title alone gives the game away: “Language Models Don’t Always Say What They Think.”

Here’s what they did. They took two models — GPT-3.5 and Claude 1.0 — and quietly slipped a bias into the input, then watched what the answer and the written reasoning did. A couple of the biases:

  • Rig the multiple-choice options so the right answer is always labeled “(A)”
  • Drop a hint that nudges the model toward a wrong answer on purpose

They ran these tricks across 13 hard reasoning tasks. The point wasn’t whether the model could be fooled — it was whether the model’s explanation would own up to being fooled.

It Got Pulled By The Bias, And Never Said So

The result is the part that made me sit up.

The model takes the bait. Plant “the answer is always (A)” and it starts leaning toward (A) — fine, that’s not shocking. The unsettling bit comes next: in its written explanation, the model never once mentioned the bias. Not a word about the rigged labels.

Instead it does something almost human. It builds a fresh, perfectly sensible-sounding case for why (A) is correct — “(A) is right because of this and that” — and uses that to justify the answer. The real reason was “it got pulled by the label order.” The explanation breathes not a syllable of it.

And when the bias points toward a wrong answer, accuracy across the tasks dropped by up to 36%. So the model isn’t just occasionally wrong — it’s confidently, articulately wrong, complete with a clean rationale. A separate run using social stereotypes showed the same move: a biased answer, justified without ever copping to the bias behind it.

The Explanation Is A Plausible After-The-Fact Story, Not A Record

What this study is really showing is that an AI’s reasoning can be a justification written after the decision was already made. The technical name is post-hoc rationalization (ie the conclusion comes first, and the tidy logic gets assembled afterward to fit it).

Which, honestly, is a very familiar habit. You pick the thing you wanted anyway, then back-fill the case for it — “well, it’s better value, and the reviews were stronger.” The choice came first; the reasons came to dress it up.

The AI does the same thing, except it does it fluently and with no apparent idea that it’s doing it. It’s the student who copied the answer off a neighbor and then wrote out a flawless worked solution underneath. The solution is real, it’s coherent, it’s well-formatted — and it had nothing to do with how the answer got there.

So here’s the line worth holding onto: an explanation being detailed and logically tight is one thing, and that explanation being the actual reason is another. We tend to see the first and assume the second.

Conclusion: Treat The AI’s Reason As A Hypothesis, Not A Verdict

Let me bring this back to the marketing desk, because that’s where it bites.

When an AI recommends your brand, the obvious next move is to ask it why — and then build your strategy on the answer. I get the pull; I want to do it too. But after this study, that “why” might be a plausible after-the-fact story rather than the real mechanism. Optimize against it and you can spend a quarter aiming at the wrong target with total conviction.

So how do you not get tripped up? The trick is to demote the AI’s stated reason from a conclusion to a hypothesis. If the model says “they picked your brand on price,” treat that as the start of a question worth checking, not the answer. Hold it that loosely and a made-up rationale can’t quietly steer your roadmap.

Then lean on the number you can actually see, not the story. The reason wobbles — but “out of 100 asks, how many times did your brand show up” is something you observe directly, an appearance rate, not a narrative the model improvised. Build your dashboard on what you can measure and your footing gets a lot more solid than it does on an explanation that might’ve been written to please you.

AI explanations will keep getting slicker. But “can explain itself” and “can be trusted” still aren’t the same thing — and just keeping that gap in mind changes how you read everything the model tells you.


Source

  • Miles Turpin, Julian Michael, Ethan Perez, Samuel R. Bowman (2023), “Language Models Don’t Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting”, NeurIPS 2023, arXiv
Share this article
Bluesky X
Back to all articles