Back HexScope Lens

Does The AI Know Your Numbers? 197,000 Questions Found The Dividing Line

POINT Key points
  • Accuracy tracked company size and how readable the filings were
  • The bigger and more recent the data, the more the model invents

“I asked an AI for our company overview and it got our numbers wrong.”

Ever had that land in your inbox?

I used to wave it away: the model’s training stopped before that, so of course it doesn’t know. Tidy explanation.

Except the figure it fumbled that day was three years old and had been public the whole time.

If older facts are the ones a model should have nailed down, why does accuracy swing so hard from company to company?

Somebody put more than 197,000 questions about US-listed companies to a set of models and checked every answer against the official filings.

How Accurately Do AI Models Know A Company’s Financials?

It comes down to size and to how readable the disclosures are.

Agam Shah and four co-authors at Georgia Institute of Technology asked models over 197,000 questions about the financial data of US-listed companies and matched each answer against official figures (published 2025, accepted at COLM 2025, arXiv:2504.00042).

Larger companies, companies with more investor attention, and companies whose filings read more plainly were answered more accurately.

The same study found that for larger companies and more recent fiscal years, the rate at which models confidently produced wrong numbers went up.

Being known accurately and not being invented about are two separate things.

Re-Measuring The “Anything Before The Cutoff Is Known” Assumption

A training cutoff is just the line marking how recent the material a model learned from is. Everyone knows a model is shaky on events after that line.

What this study measured is what sits before it.

The method is unglamorous. Build over 197,000 questions about financial line items for US-listed companies, then check each answer against the official data, one at a time.

Then see how four signals — market capitalisation, retail investor attention, institutional investor attention, and the readability of the filings — track with accuracy.

The first thing to break was the assumption that everything before the line is known. The further back the fiscal year, the lower the accuracy.

What Separated Accuracy Was Size — And Readability

Bigger companies get their numbers right more often. Most people would guess that much.

The interesting part is that readability of the disclosures pulled its weight alongside market cap. Same size, different writing, different result.

The readability measure is essentially a score for whether a document is written in plain sentences with its figures placed where they can be read off.

And that one is a variable a company can actually move. You can’t raise your market cap today. You can rewrite the key points of your results in plain text and legible figures today.

That’s the part of this study I found genuinely encouraging.

The investor-attention finding points the same way. Information people read and quote often is information that survives on the model’s side too.

The Paradox: Bigger Companies Get More Invented About

Here’s what stops this being pure good news. For larger companies and more recent fiscal years, hallucinations rose.

Hallucination here means the model answering with a fact that doesn’t exist while sounding entirely sure of itself. When it happens to your numbers, outsiders quote it without checking.

With more material lying around, the model stops declining to answer and assembles something plausible instead. It’s the exam candidate who studied hardest for a subject and therefore fills every blank with a confident wrong answer rather than leaving it empty.

So a large company is more likely to be both known accurately and misquoted freely.

“The models are precise about us, so we’re fine” doesn’t follow.

Reading This From Outside The US

The scope is the financial data of US-listed companies. Whether it extends to other language markets, private companies, or non-financial company information is not something this study establishes.

The readability measure was also built around the characteristics of filings submitted to the US securities regulator. Another country’s disclosure regime won’t necessarily take the same yardstick.

The models evaluated and the question set are as of publication, so whether the same pattern holds after model updates needs checking again.

With all three caveats in place, one thing still travels: the lever the company itself controls turned out to be the readability of its disclosures.

Put Your Numbers Somewhere A Model Can Read Them

One action is enough to start. Open the page carrying your own results and figures, and look at whether a model could read it.

Are the amounts and counts trapped inside images or charts? Are the fiscal year and the units written as text in the body? Fixing those two changes what a model has to work with.

Alongside that, get in the habit of checking answers more than once and across more than one model. Getting your numbers right on a single run doesn’t rule out the paradox above still being in play.

Known accurately, and not invented about. Two separate checks — that’s the takeaway.


Sources

  • Agam Shah and four co-authors (Georgia Institute of Technology), “Beyond the Reported Cutoff: Where Large Language Models Fall Short on Financial Knowledge”, arXiv:2504.00042, published 2025, accepted at COLM 2025, arxiv.org (over 197,000 questions on the financial data of US-listed companies, each answer matched against official filings. Accuracy rose with market capitalisation, retail and institutional investor attention, and the readability of disclosure documents, and fell for older fiscal years. Hallucination rates rose for larger companies and more recent fiscal years. Scope is limited to US-listed companies’ financial data; the readability measure is built around US securities filings; the evaluated models and question set are as of publication.)
Share this article
Bluesky X
Back to all articles