“I Hadn’t Touched It In Months, And It Was A Different Thing”
About six months ago you tried some AI tool, decided “not for our workflow yet”, and left it alone. The other day you opened it again — and it was sharp in a way it hadn’t been, sharp enough to be slightly disorienting.
That’s happening to a lot of people lately, isn’t it.
It’s happened to me more than once: I write something off, come back later, and end up reversing my own verdict.
We all quietly assume that an AI’s smarts climb in proportion to how much it has studied — more training, a bit better; more training, a bit better again, in a nice orderly line you could draw with a ruler.
But reality, it turns out, is nowhere near that well-behaved, and there’s research saying so. So let’s dig into it.
So What Is This “Grokking” Thing?
The starting point is a phenomenon that OpenAI researchers reported in 2022, called grokking.
Grokking is roughly this: an AI memorizes its training data, looks for a long time like it has hit a ceiling, and then — much, much later — starts behaving as if it had understood the underlying thing all along. The English verb “grok” means something like “to finally get it, all the way down” (which is exactly what it looks like from outside: one day the thing just clicks).
The part that matters is the “one day”. Not a gradual climb, but a long flat stretch and then a lurch. And that’s precisely where it collides with the tidy picture most of us are carrying around.
Memorize First. Understand Much Later.
The experiments used modular arithmetic — the kind of counting that wraps around at a fixed number, the way a clock does. Nothing exotic.
At first the AI does what a lazy student would do: it memorizes the problems and their answers, so it’s near-perfect on the questions it has already seen and hopeless on anything new. Past exams learned by heart, zero transfer.
But if you don’t stop there — if you keep training long past the point where continuing looks pointless — a moment eventually arrives when the unseen problems start going down easily too. That’s the shift from memorizing to grasping the rule itself: generalization (ie the ability to apply what was learned to something it had never seen), arriving late.
Nothing Was Happening. Except Inside.
Here’s the question I find genuinely interesting: what was the AI doing through all that flat time?
Later researchers (the analysis by Nanda and colleagues) opened the model up and looked, and behind what had looked like a standstill they found the circuitry for generalization quietly assembling itself the whole way through. From outside, nothing. Inside, steady construction.
It’s the houseplant thing. You water it every day for weeks and nothing at all happens above the soil — while the roots spread out below it, on schedule, ignoring you completely.
Out of this came a way of framing the whole phenomenon: as a phase transition, which is what water does at zero degrees when it sits there as a liquid, sits there as a liquid, and then crosses a line and snaps into ice. Capability, on this reading, doesn’t accumulate smoothly; it sits, and then it jumps once a threshold is crossed (a deeply annoying property to plan around, if you’re the one doing the planning).
Now the honest caveat, and maybe I’m being too cautious here. This was observed in small, simple experimental setups, and you shouldn’t just assume it carries over unchanged to the enormous AI systems people actually use at work. But the suggestion itself — that capability sometimes jumps — seems too heavy to shrug at.
Conclusion: Don’t Price An AI Off A Single Snapshot
Pull this back to the day job and the lesson is pretty simple: “here’s how good the AI is” isn’t a verdict you hand down once and then file away as settled. Because if capability can move discontinuously, then reading a single snapshot and concluding “not usable yet” or “it’s plateaued” is a genuinely shaky call — the flat stretch might be exactly where the next jump is getting built.
So the thing worth doing isn’t a one-off appraisal; it’s a fixed observation point. Take the AI you care about, re-run the same task on it once a quarter, and watch for the early signs that something has moved. Hold the evaluation as something you keep updating (rather than a conclusion you reached once and never revisited) and you won’t be the last person in the room to notice.
And this isn’t only about capability. What an AI says about your company, and which brands it puts on the shortlist, can change on you just as abruptly. Checking once and feeling reassured is the risky version; watching the change from a fixed point is the useful one — and honestly, that habit is probably the best preparation available right now.