Back HexScope Lens

Why Won't ChatGPT Just Say "I Don't Know"? Blame The Grading

POINT Key points
  • Marking "I don't know" as zero taught the models to bluff
  • A confident tone tells you nothing about whether the answer is right

“Can’t ChatGPT Just Tell Me When It Doesn’t Know?”

You type your own company name into ChatGPT and get back an award you never won, sitting next to a number nobody in the building recognizes. Delivered without a flicker of doubt. You’ve had this happen, right?

I tried it on a friend’s company once and got a whole founding story that never happened. I read it twice.

But the odd part isn’t that the model gets things wrong; it’s the “just say you don’t know” part. A person would shrug and say “sorry, not sure”. The model writes fiction instead, and sounds great doing it.

In September 2025 a research team at OpenAI took that exact question seriously and put out a paper on it (the paper). So let’s look at what they say.

Is A Hallucination Some Mysterious AI Malfunction?

First, the word. A is when an AI states something that isn’t true, fluently and with total confidence.

Most people file this under “weird bug that shows up sometimes”; the assumption underneath is that smarter models will eventually grow out of it.

But the paper says no. Hallucinations aren’t a glitch; they’re the ordinary, predictable output of how these models are trained and, more importantly, how they’re graded.

The authors ran a small demonstration on themselves: they asked a model for one author’s birthday, three times. Three different dates came back, all of them wrong, none of them hedged. Three tries, three misses, zero hesitation.

The Student Who Never Leaves A Blank

So why can’t the thing stay quiet? The clearest explanation in the paper is an analogy: a student sitting an exam.

Picture it. You hit a question you can’t answer. Leave it blank and you get zero, guaranteed. Write something down and you might get lucky. So you write something down.

The paper’s read is that a model faces exactly the same incentive. Saying “I don’t know” pays worse than producing a plausible-looking answer, so the plausible-looking answer is what you get.

Which means the model isn’t lying to you, exactly. It’s going for points.

How Did Guessing Become The Winning Move?

This happens in two stages.

The first is the training itself: the model reads an enormous pile of text and learns how words hang together, and that pile contains true statements and false ones, mixed in together. With no perfect way to sort one from the other, some amount of error is statistically unavoidable. That part isn’t going away.

The second stage is the one that really bites: how the model’s ability gets measured afterwards.

Model performance is scored on (basically the industry’s standardized exams). Big question sets, one score, a leaderboard at the end.

And most of that grading is binary: right answer, one point; wrong answer, zero. Under those rules, an honest “I don’t know” and a confident falsehood earn exactly the same thing. Nothing.

The Honest Model Loses The Leaderboard

So what’s actually wrong with binary scoring?

If “I don’t know” scores zero, guessing is free upside — you only ever gain from the lucky hits, and a model that flags its own uncertainty leaves all of those points on the table and slides down the ranking.

It’s the same thing that happens to a student who’s been told to fill in every blank before handing the paper in. Honesty gets no credit, confident guessing does, and every model in the field has been competing on that pitch.

So the paper’s proposal isn’t “add another hallucination test”. It’s to change the scoring on the existing, mainstream benchmarks: penalize confident errors more heavily, and give partial credit for an appropriate “I don’t know”. Make honesty stop costing you.

I think that’s the sharpest part of the whole thing. Hallucination isn’t a problem that dissolves once the models get clever enough. As long as the scoring rewards guessing, smarter models keep guessing.

Conclusion: Don’t Read Confidence As Evidence

So what does any of this mean if you’re the one running a brand?

The main takeaway: how sure the model sounds and how right it is have nothing to do with each other. Saying “I don’t know” is structurally hard for these systems, and the less they know, the more smoothly they tend to invent. A fluent, unhedged answer is the most convincing thing on the page — and the fluency proves nothing.

And this isn’t something to wait out. “The next model will fix it” doesn’t follow — not while the exams still pay for guesses. So treat anything an AI says about your company as a claim that a human still has to check at least once, no matter how new the model is.

Which leads to the one habit worth building: don’t settle for checking once. Answers wobble between sessions, and a model update can move them wholesale. What the AI says about you, and how much of it is true, is worth watching on a schedule rather than assuming you already know. Dull, but it’s what keeps you from being pushed around by a confident tone.

Share this article
Bluesky X
Back to all articles