“Wait, That’s Not What Mine Said”
You’re comparing notes with a colleague about something you both asked ChatGPT, and the answers don’t line up. Roughly the same question. Different recommendations, different level of detail, a different tone entirely.
I did this a while back with a friend, holding up two travel plans we’d each had the model draft. The itineraries had almost nothing in common, and we both laughed.
Same tool. So where does the gap come from?
But this isn’t the model getting suddenly smarter. It’s the model getting to know you. OpenAI shipped a big memory overhaul in June 2026, and once you look at what’s inside it, this stops being a party trick and turns into a measurement problem for anyone who owns a brand.
ChatGPT Is Being Raised To Your Spec
The update in question is a memory revamp OpenAI announced in June 2026, called “Dreaming” (R). Memory is the part that keeps hold of your preferences and your situation from past conversations, and then bends the next answer around them.
The numbers OpenAI gives are a real jump. Factual recall (ie whether it pulls back something you told it weeks ago) went from 67.9% in 2025 to 82.8%. On preference adherence, how well it matches what you actually like, task success sits at 71.3%.
Now, this is a company announcing its own product (I couldn’t find an outside replication), and being a bit suspicious of the numbers is the healthy reaction. Discount them a little.
But the direction — models that remember users and personalize the answer — is the part worth holding onto.
”Everyone Gets The Same Answer” Isn’t The Baseline Anymore
OpenAI’s own example makes this concrete, and it’s clearer than most demos.
With memory on, the model leans on things you mentioned before — the camera you own, the trip you were planning last month — and shapes the recommendation around them. With memory off, everybody gets the same generic list, and you’re left doing the research yourself afterwards.
So the old assumption — ask ChatGPT, get roughly what everyone else gets — is on its way out.
Your colleague’s mismatched answer is just the front door.
Handy, And Nobody Can Check The Receipts
There’s an awkward flip side. A 2026 paper at ACM CHI (the international conference on how people and computers get along) puts its finger on it.
They call it the personalization-convenience paradox: the features users value most are exactly the ones that are hardest to audit and hardest to control. The model nudges its picture of you with every conversation, and how that picture got built — and how it fed into today’s recommendation — isn’t something even you can fully trace.
It’s a bit like handing your household finances to someone who genuinely does a good job of it every month, but never shows you the statement. The convenience is real.
The “why did this come out this way” is gone.
Convenience and unauditability sitting in the same box is an uncomfortable combination. And reporting on the memory revamp has already flagged the same thing: the basis for personalization is getting harder to follow.
Your Brand’s “AI Visibility” Isn’t One Number
Here’s where it becomes your problem rather than a curiosity.
If recommendations are personalized per person, then how your product looks to the AI is personalized too. You’re a regular on one person’s shortlist and completely absent from another’s.
That’s an ordinary outcome now, not a glitch.
Which breaks the usual way people measure. AI visibility is basically how readily the model recommends you on its own (no prodding, no prompt engineering).
The common check has been: ask ChatGPT about your company, see whether you come up. One shot, one account. But once personalization kicks in, that answer wobbles depending on whose ChatGPT you used. A number from a blank account might not match what a heavy user actually sees.
Conclusion: Measure With “Whose AI Is This?” In Mind
The more ChatGPT personalizes, the less your brand’s AI visibility can be told as a single correct story. Same question, different history, different recommendation. So the question itself — “how does the AI see us?” — needs a bit of rework.
The first thing I’d change in practice is to stop treating AI visibility as one number. A one-shot check on a blank account is an averaged picture of everyone, and it drifts from what personalized users actually get. Measure it more than once, vary the phrasing and the kind of person you’re imagining behind the prompt, and read the spread rather than the point.
The second is smaller and you can do it this week: run your own logged-in account and a fresh, memory-free one against the same question about your category, and put the two answers side by side. The gap between them tells you roughly how much personalization is moving your brand around, and it’s more informative than either answer alone.
Maybe I’m wrong about how big that gap turns out to be — personalization is still a black box, and this is early. But “the picture of us inside the AI isn’t a single flat image” is a safe assumption to start carrying around now.