When you write a target-customer persona, the age and the job title are the easy part.
It’s the next box down that gets awkward — the one labelled “values”, or “what’s keeping them up at night”. That one you fill in from your own head.
I’ve stalled in front of that box more than once. Not for lack of things to write, but because whatever I wrote was going to be me in a costume.
Lately there’s a shortcut. Ask ChatGPT to answer as a thirty-something urban mother and it’ll oblige, fluently, on the first try. So of course somebody’s wondered whether that could stand in for an actual interview.
But the question underneath won’t go away: how close is that answer to a real human’s?
And a team went at it with something close to brute force — they rebuilt a thousand real people, one at a time. So let’s look at what came out.
They Interviewed 1,052 People, Then Rebuilt Them
The study came out of Stanford and Google DeepMind in 2024 (R).
What they did was less clever than laborious. They recruited 1,052 real participants, put each one through a long, unhurried, in-depth interview, fed that whole account to a model, and built one agent per person.
A is basically a stand-in — an AI that fields questions on someone’s behalf. Which, for you, means the interesting bit isn’t the tech, it’s the input. A persona built from demographics is a sketch of a type of person. These were built from what the actual person said.
So How Close Did The Copies Get?
Here’s where it gets genuinely fun.
The team ran a big survey — the , a long-running staple of social research — past both the real participants and their AI stand-ins, then compared the answers.
The stand-ins reproduced their person’s answers with about 85% accuracy.
But the clever part is the denominator. 85% of what? The ceiling isn’t perfection — it’s the same person’s agreement with themselves when they retook the same survey two weeks later.
Because people, it turns out, drift. Two weeks is enough to disagree with your own answers, which is a slightly humbling thing to learn about yourself (I’d like to say I’d be the consistent one; I would not be). That “even the real person only matches the real person this much” line is the 100% mark, and the AI copy got 85% of the way there. Grading against live human consistency is a hard test, not a generous one.
What It’s Good At, And What It Isn’t
The match wasn’t just on the survey. On personality-trait prediction, and on the experimental games behavioural economists like to run (the split-the-money, trust-a-stranger sort of setup), the stand-ins scored about as well.
And a quieter result matters more than it looks: the gap between groups was small. Agents built from demographics alone tend to be more accurate for some racial or ideological groups than others; agents built from the person’s own account held their accuracy steadier across those groups.
Still, nothing here is a clean sweep. The team says plainly that some attitude-type predictions came up short. Which I’d file as: good at reproducing the average tendency, not at photocopying one person’s private reasons.
The Questions Need Checking Too
Push synthetic answers into actual marketing work and a second problem shows up, one that gets talked about far less: who writes the questions.
There’s research on handing survey-question drafting to AI as well. It’s good at producing questions that fit the context — but it also produces the classic human failures, like cramming two separate issues into one question, or laying the jargon on thick (R).
So neither the synthetic questions nor the synthetic answers are usable until a human has read them over.
Which is, I think, the honest summary: AI synthesis makes a pretty good first draft, and the moment you skip the verification it turns into confident, plausible fiction.
Conclusion: Let AI Do The Homework, Keep Measurement For The Real Thing
So what does a marketer or brand lead take away from this?
The line I’d draw is this: synthetic answers are strong for generating hypotheses and designing research, but they’re not a substitute for knowing how your brand is actually seen. Getting to 85% is impressive, but it’s 85% of a live human benchmark — the real people were the answer key. That’s not the same as replacing the measurement.
So in practice, label everything. This came from asking an AI; that came from measuring real people or real data. The first is what you argue over, the second is what you decide on. Blend them and you’re back in the same hole as the persona you invented and then believed.
And on your own brand specifically, don’t settle for the AI’s plausible summary of you — check, on a regular cadence, how you’re actually described and where you actually come up. Maybe I’m wrong about how long that distinction stays useful. But keeping synthesis and measurement in separate columns is a dull little habit that keeps paying, and it’s the difference between using this stuff and being used by it.