Back HexScope Lens

1,052 People, Cloned By AI, Matched 85% Of Their Own Answers

POINT Key points
  • Two-hour interviews, not demographic labels, drove the 85% match
  • Best for front-loading hypotheses and survey questions, not replacing real research

“Can’t we just run the research through AI first?” If you do marketing, you’ve probably heard that line in a meeting — or said it yourself. Research is necessary, but it’s heavy: recruiting drags, sample size hits the budget ceiling, and every tweak to the questionnaire pushes the timeline back.

So the obvious move is to ask the AI first. And plenty of teams already do it — letting a model play the respondent, or draft the questions. The catch is accuracy, and where in the workflow this actually belongs.

Because the answer changes everything downstream. If AI can replace the real study, you run things one way. If it can only speed up the front end, you run them another. So let’s look at two studies that decide which it is.

A “Pretty Convincing” AI Person Got To 85%

Start with the work from Park and colleagues at Stanford and Google DeepMind (paper).

The clever bit is the input. They interviewed 1,052 people for two hours each, then built an AI agent of every single person from that transcript. Not a demographic label, but life history and values — so the information density is nothing like the usual three-line persona.

To test the agents, they used the . And the result? The AI agents reproduced each person’s own answers about 85% of the time.

That 85% isn’t a raw accuracy score. It’s measured against the same person’s answers when they retook the survey two weeks later. People aren’t perfectly consistent with themselves either, so the right reading is: the agent got close to that human-vs-human ceiling.

There’s one more thing that matters more than the headline. Agents built from interviews showed less bias across groups than agents built from demographics alone. So the lever isn’t some model magic — it’s the depth of customer understanding you feed in.

The Question Writer Is Fast. Ship It Raw And You’ll Get Hurt.

The second study, from Mburu and colleagues, flips the roles (paper). Here the AI isn’t the respondent — it’s the survey designer. Their SQRA procedure has the model draft questions, then pre-tests those questions against synthetic answers before any human sees them.

What showed up was both the speed and the snag. The model is quick at producing context-fit questions. But it also slips in wordy phrasing and double-barreled items (one question secretly asking two things). So the drafting is strong, while the quality control stays on the human side of the desk.

For “get a first draft fast”, that’s genuinely useful. But shipping those questions live without a review pass is asking for trouble. That’s a line worth drawing in bold.

Two Opposite Setups, Tripping On The Same Spot

One study makes the AI the respondent; the other makes it the question writer. Opposite approaches — and yet the thing that decided accuracy lines up almost perfectly. In both, the result rode on how rich the input was and how much the human kept, not on how clever the model is.

Park’s 85% only exists because of a dense two-hour interview. Thin the input and the accuracy drops, and the cross-group bias creeps back up. Mburu’s question generator spits out drafts fast, but catching the redundancy and the double-barreled items still landed on a person.

Flip that around and you get the real lesson: what AI is bad at is being handed a heavy decision on thin material. Dodging that one spot is the whole game when you put this into practice.

Conclusion: Use AI To Compress The Front End, Not To Replace The Study

The value of AI consumer research isn’t swapping out the human study wholesale — it’s speeding up the front end, the hypotheses and the questionnaire. It’s strong when dense interview material is feeding it, and middling when the input is thin. Forget that limit and it’s easy to mistake a handy draft for a basis for a decision.

So the first move is to inventory your old customer interviews and voice-of-customer notes, and fatten up what you feed the model. Then set the boundary: AI output goes up to hypothesis-testing, and the real calls — budget allocation, what you actually ship — run on human data. Pin that down and you’ll dodge most of the “we grabbed the speed and dropped the quality” accidents.

If you want to keep going on how shaky AI measurement itself can be, the appearance-rate piece on AI visibility pulls at the same thread from the measurement side.


Sources

  • Park, J. S. et al. (Stanford University, Google DeepMind), “Generative Agent Simulations of 1,000 People”, arXiv:2411.10109, 2024, arXiv
  • Mburu, T. K. et al., “Methodological foundations for artificial intelligence-driven survey question generation”, Journal of Engineering Education, 114(3), e70012, 2025, DOI: 10.1002/jee.70012
Share this article
Bluesky X
Back to all articles