“Can’t we just run the consumer study through AI?” If you do marketing or product, somebody in the room has said this, probably last quarter, probably while looking at the research quote.
I get the pull. Research is slow and it’s expensive, and the front half of it — the hypotheses, the questionnaire drafts — feels like exactly the kind of grunt work a model should eat.
But the question hides two very different asks. “Replace the study with AI” and “compress the front end with AI” sound like the same sentence. They’re not. They need different accuracy, different rules, and they break in different places.
So let’s take the three approaches people actually mean by that one question, put them side by side, and figure out where each one is safe to point.
Approach One: Let The AI Be The Respondent
Park and colleagues at Stanford and Google DeepMind ran the version everyone quotes (R1). They interviewed 1,052 people, built an AI agent from each person’s transcript, and then checked how well the agent could answer as that person.
The agents reproduced their own human’s answers about 85% of the time. On the number alone, that’s a strong result.
But look at what fed it. That 85% sits on top of a dense interview — a life history, not a label. Hand the model a persona that says “34, works in retail” and expect the same fidelity? Come on. You wouldn’t do a convincing impression of your best friend from three lines of profile either.
So the safe use is upstream of the decision, not at it: screening which angles are worth testing, pressure-checking a hypothesis before you spend real money finding out.
Approach Two: Let The AI Write The Questions
Mburu and colleagues flipped the roles (R2). Their framework — they call it SQRA — has the model draft survey questions, then runs those drafts through a validation step before anyone fields them.
The speed is real, and honestly it’s the part I’d miss most — a model fits questions to your context faster than any human drafter. But the same work reports the snag: the drafts come out wordy, and they slip in double-barreled items (one question quietly asking two things at once, so you can’t tell which one the answer belongs to).
Which means AI question generation is a way to cut hours, not a way to skip quality control. It works if a human reads the draft before it ships. It doesn’t work otherwise.
Approach Three: Treat Real Prompts As The Research Itself
The third one isn’t simulation at all. Otterly AI went and looked at what people actually type into ChatGPT (R3), and the real prompts came back longer, more personal, and more problem-shaped than the ones marketers guess at.
That’s a different kind of asset. Nobody had to imagine how customers phrase their problem; the phrasing was sitting there in the log.
The catch is — basically, whether the people in your data look like the people you sell to. Which, for you, means a thin or lopsided sample doesn’t give you a small truth; it gives you a confident wrong answer. So I’d treat this as an extra signal alongside your existing research, not as the ground truth that retires it.
Conclusion: Pick The Stage First, Then The Method
Three approaches, and none of them is a candidate to replace the human study. They’re tools that bite at three different points in it: widening the hypothesis space, drafting the questionnaire, reading how customers actually talk.
So the first decision isn’t “which method do we adopt”. It’s “which stage are we handing over”. Screening hypotheses can go to AI respondents. First-draft questions can go to the question generator. The call that moves budget goes back to human data, every time — and if you only remember one line from this, make it that one.
Then put one gate in front of all three, and staff it with a person: is the input material thick enough, are the questions clean, is the sample representative of anyone you care about? Maybe I’m being too conservative about that gate. But every AI-research accident I’ve watched came from skipping it, so I’ll take the trade.
Sources
- [R1] Park et al. (Stanford / Google DeepMind), “Generative Agent Simulations of 1,000 People”, arXiv:2411.10109, 2024, link
- [R2] Mburu et al., “Methodological foundations for artificial intelligence-driven survey question generation”, Journal of Engineering Education, 2025, link
- [R3] Thomas Peham (Otterly AI), “Real vs Estimated Prompts: I Analyzed 100s of Real ChatGPT Queries”, 2026-02-03, link