Back HexScope Lens

Let's Take "The AI Is Just Predicting" To A Brain Scanner

POINT Key points
  • Models better at guessing the next word matched brain activity better
  • Fluent isn't sourced, so read AI answers beside primary data

“It’s Just Predicting The Next Word, Right?”

Somebody in the meeting always says it, usually with a shrug. It’s just predicting the next word.

And I get why the line sticks around. You use ChatGPT all day, the writing comes out clean, it picks up what you actually meant — and then the explainer tells you the thing is doing autocomplete with a bigger engine.

I’ve bounced between those two feelings for a while now. Impressive, if that’s all it takes. Also: if that’s all it is, where does it break?

But here’s what made me stop and reread. Neuroscientists have been checking how much of your language processing is prediction. The answer isn’t “none”.

So let’s look at what they found — and, more useful for your week, where the resemblance runs out.

The Models That Guess Best Also Guess Your Brain Best

Start with the paper that put this on the map: Schrimpf and colleagues at MIT, in PNAS (the study).

They took dozens of neural network language models and ran each one against three neural datasets — recordings of human brains handling language — plus one behavioral dataset.

The question was blunt: which AI model best predicts what a human brain does with a sentence?

So here’s what came out. The better a model was at next-word prediction, the better it predicted human brain activity. And the best model got close to the noise ceiling (roughly: the point where whatever error is left is measurement noise, not a flaw in the model).

Next-word prediction is what it sounds like — given everything so far, guess the word that comes next. That’s the training objective behind an LLM (large language model, the machinery under ChatGPT).

So if people also run ahead of a sentence while reading it, some overlap between the model’s math and your brain’s language processing isn’t shocking. It’s about what you’d expect.

Which makes it very tempting to say the machine understands language the way you do. Hold that thought.

Your Brain Was Already Ahead Of The Sentence

That’s all measured on controlled texts and tasks. What about something closer to everyday listening — the way a customer half-listens to a podcast ad?

That’s the setup in a Scientific Reports paper from Kölbl and colleagues (the study). 29 people listened to a German audiobook while their brains were recorded two ways at once: EEG (electrical activity picked up at the scalp) and MEG (the faint magnetic fields the brain gives off).

The team then lined up those brain patterns against BERT’s predictability scores across four parts of speech — nouns, verbs, adjectives, proper nouns. (BERT is a language model that predates ChatGPT, and it’s still a standard tool for this kind of analysis.)

Two things stand out:

  • For nouns, significant pre-activation showed up before the word had started.
  • The more predictable BERT judged a noun to be, the smaller the N400 — a brain response that swells when a word is hard to square with what came before.

So even in ordinary listening, your brain is running ahead of the speaker. Prediction isn’t the cheap trick the AI does instead of understanding; it looks like a core piece of how people handle language too.

Genuinely cool. Also exactly where I’d stop before turning it into a business conclusion.

Where The Overlap Runs Out

Because if prediction were the whole story, one number from a model ought to explain the brain data. It doesn’t.

Krieger and colleagues went at that directly in Brain Research (the study). Their subject is surprisal — a model’s measure of how unexpected a word is. Low surprisal means “yeah, saw that coming”; high means “wait, that word?”

What they report: LLM surprisal didn’t consistently explain the N400 or the P600 (a later brain response tied to going back and reanalyzing a sentence). Small models got yanked around by plain word association. And the large models still didn’t capture graded plausibility or event knowledge — the everyday sense that a customer might return a jacket, is unlikely to eat one, and would never ask one for feedback.

That gap is the whole point. A model can produce “our customers care most about delivery speed” with perfect fluency and zero contact with your actual customers. Smoothness is a property of the text. Grounding is a property of where the claim came from.

Maybe I’m wrong about how far this generalizes — three studies and a handful of brain measures. But the direction is consistent, and it matches what anyone who’s ever fact-checked an AI summary already suspected.

Conclusion: Judge An AI Answer By Its Sourcing, Not By How Smoothly It Reads

So the usable version of all this: fluent and sourced are different things, and we’re unusually bad at keeping them apart — partly because prediction is what our own reading runs on too.

One change worth making this week. When an AI summary is about your market or your customers, read it next to a primary source: interview transcripts, support tickets, reviews, purchase data. Let the model summarize, cluster, and float hypotheses. Let real data settle the call.

And the corollary people skip: what the AI says about your brand is a signal you can measure, not a verdict on your category. The model’s account of who leads the category is a fluent guess assembled from whatever it happened to read — which makes it worth tracking, because that guess is often the first thing a buyer sees of you, and because it moves.

If you want to see what’s going on inside the model while it makes that guess, Does The AI Actually Understand, Or Just Sound Like It? picks the thread up from the inside.


Sources

  • Schrimpf, M. et al., “The neural architecture of language: Integrative modeling converges on predictive processing”, PNAS, 2021, pnas.org
  • Kölbl, M. et al., “Prediction, syntax and semantic grounding in the brain and large language models”, Scientific Reports, 2026, nature.com
  • Krieger, K. et al., “On the limits of LLM surprisal as a functional explanation of the N400 and P600”, Brain Research, 2025, sciencedirect.com
Share this article
Bluesky X
Back to all articles