Back HexScope Lens

Why Bigger AI Gets Smarter Turns Out To Be A Packing Problem

POINT Key points
  • Double the model's width and it misses roughly half as often
  • Gives you a footing for how far scaling keeps paying off

“Does AI Just Keep Getting Smarter If You Keep Making It Bigger?”

You’ve been hearing it for years: make the AI bigger and it gets smarter. Every launch leads with a parameter count — hundreds of billions, trillions — and somewhere in the back of your head you file it as “so scale is what matters.”

I do this too. Every new model, the first thing I want to know is how much bigger it got this time.

But stop for a second and it’s a strange thing to nod along to. Why would bigger make it smarter? And is that going to keep holding?

But for a long time nobody could really say. It was a rule of thumb — a very reliable one, but still a rule of thumb.

Then in 2025 a paper turned up that answers the “why” from the inside, so let’s dig into it.

What “Bigger Is Smarter” Actually Says

Quick recap first. AI has a thing called the scaling law, which says roughly this: pile on more training data, more compute, and more model, and performance climbs smoothly.

“Performance” gets measured as how badly the predictions miss. The technical name is (ie how far the prediction sits from the right answer) — smaller is better, and that’s all you need from it.

What’s striking is how clean the curve is. Add this much scale, get this much less loss, over and over. Which is exactly why everyone’s been comfortable pouring money into “just make it bigger.”

And every year somebody announces the ceiling is finally here, and every year that prediction ages badly — the curve just keeps going (the “it’s over” genre never quite delivers).

But there’s a nagging gap in all this. Clean isn’t the same as explained; nobody could say why the curve behaves like that, and it stayed at “we tried it, and that’s what happened.”

The Trick Is How Much An AI Crams Into Itself

The key to the “why” is a phenomenon called .

Superposition is when a model holds far more features than it has parts to hold them in. A feature is one of the little handles an AI uses to read the world (eg “this looks like polite writing”, “this is a cooking topic”) — thousands upon thousands of small angles like that.

You’d assume 100 parts means 100 features, tops. But a real model packs thousands of features into those hundred parts, tilting each one slightly so they can sit on top of each other. Picture a filing cabinet with more documents than slots, slid in at an angle so you can still see the edge of every one.

Anthropic worked this out back in 2022 on small models (R), and it became the starting point for interpretability research — the business of reading what’s actually happening inside a model.

Cramming Is What Produces The Scaling Law

So here’s the new part. A study out of MIT (Liu, Liu, Gore and colleagues, picked as a Best Paper Runner-Up at NeurIPS 2025) argues that superposition is what the scaling law is made of (R).

The team worked the math out for the regime where features are strongly superposed — crammed in far past the number of parts available. And a clean relationship falls out of it:

Increase the number of parts (the model’s width), and the loss drops in inverse proportion.

Double the width, miss roughly half as much. That inverse curve is the inside of “bigger is smarter.”

The impressive bit is how sturdy it is. The relationship holds across a wide range of conditions, it doesn’t hinge on the fine detail of how the features are spread out, and if you nudge the assumptions the shape of the curve survives anyway.

At that point it stops looking like a coincidence and starts looking structural.

What’s New Is That A Rule Of Thumb Got A Reason

This is the part I’d underline.

Scaling laws were always reproducible — measure again, same shape. What was missing was the mechanism. The “why” was left hanging.

Now the curve can be derived from the geometry of how an AI crams features together. Anthropic’s superposition and the scaling law every lab has been leaning on turn out to be the same line, drawn from two ends.

So the scale-up everyone’s been doing because it works has a spine under it now: why it works. Knowing the reason doesn’t make anything smarter, obviously. But it changes what you can see from here.

Conclusion: You Can Say “Scale Still Pays” And Point At Why

So what’s in this for a marketer or someone signing off on budget?

One thing is reassurance. The industry’s “just make it bigger” looks likely to keep paying off for a while yet. At least as things stand, scale working isn’t a lucky streak; it looks rooted in how models are built on the inside. So I don’t think you need to plan around big models hitting a wall any time soon.

The other thing matters more, and it’s the one I’d actually watch. Once you know why scale works, you’ve got a footing for how far it works and where it starts to flatten. Cramming has geometric limits somewhere, and as models approach them that inverse curve should eventually bend. Maybe I’m wrong about how soon — but that’s a call you can now make from structure instead of vibes.

Where AI capability and cost are heading feeds straight into what you budget and which tools you buy. So don’t stop at “bigger seems smarter.” Carry the next question with you: why, and how far. Unglamorous, but it’s the thing that keeps you from swinging between hype and doom.

Share this article
Bluesky X
Back to all articles