Ask AI to draft a survey and the speed is real. You type a rough goal, and thirty-odd questions land in seconds. For anyone who’s spent an afternoon staring at a blank questionnaire, that alone feels like a gift.
I do this too. It’s hard not to lean on it.
But here’s the part that bites later. The survey goes out, the responses come back, and the spread is all over the place — the analysis axes don’t line up, and half the answers seem to be answering a slightly different question. It’s tempting to blame “bad AI.” That’s an assumption, not a diagnosis.
So let’s look at three studies that pin down where the quality actually cracks — question generation, prompt sensitivity, and how real people phrase things.
”The Draft Looks Great — So Why Do The Answers Fall Apart?”
Start with the obvious tension: the questions read fine, so what’s wrong with them?
A team led by Mburu got specific about this (Mburu et al., Journal of Engineering Education, 2025). They built a method for generating survey questions with an LLM and pre-validating them with SQRA (a structured review that checks each item before it ships).
The good news is real: the AI was strong at writing context-appropriate questions. If that were the whole story, you could hand the design over and go get coffee.
But the same study flags the catch — the drafts came out verbose, double-barreled (two questions crammed into one), and thick with jargon. So “fast to write” and “ready to use” turn out to be two completely different things. The draft is a starting point wearing the costume of a finished product.
Same Meaning, Different Words — And The Numbers Move
Now the sharper problem, from a study by Errica and colleagues (Errica et al., NAACL 2025). Their finding: even a paraphrase that means the same thing can move the measured result by 3.2 to 10%.
For a survey, that’s nasty. If two near-identical items are worded a little differently, you can’t tell whether a gap in the responses came from a difference in intent or just a difference in phrasing. The signal and the noise wear the same outfit.
Which means running AI-generated questions well isn’t only about what you ask. Keeping the phrasing consistent across items is now part of quality control, not a nicety you get to skip.
The AI Guesses Short And Commercial. Real People Don’t.
Third, a comparison from Otterly AI (Peham, Otterly AI, 2026). They lined up real ChatGPT prompts against the ones people assume users type. The real ones ran 71% longer on average, with far more personal pronouns and problem-oriented phrasing.
That gap lands right on survey design. Ask the AI to “write the questions people would ask,” and it drifts toward short, generalized, faintly commercial wording. So you end up quietly missing the messy, situated worry a real respondent actually carries around.
And that’s the stuff you were running the survey to find in the first place.
Conclusion: Cut The Drafting Time, Not The Quality Check
Using AI to write survey questions really does make the first draft fast. But fast doesn’t mean “the quality check got automated too.” Across these three studies, the pattern is the same: AI is great at throwing out a rough cut, and the last job — shaping it so respondents don’t misread it — still sits with a person.
So the thing worth auditing before you hit send isn’t the draft’s polish. It’s three plain checks: is each item one question, not two? Is the terminology consistent across items? Does a respondent have enough context to picture the situation you’re asking about?
That’s the split I’d make the rule: let AI own the fast first draft, and keep the “is this safe to ship” call with a human. Maybe I’m wrong about the exact division of labor — but drawing that line explicitly is the cheapest insurance you’ve got against clean-looking data that quietly means nothing.
Source
- Mburu, T. K., Rong, K., McColley, C. J., & Werth, A., “Methodological foundations for artificial intelligence-driven survey question generation”, Journal of Engineering Education, 114(3), e70012, 2025, DOI:10.1002/jee.70012
- Errica, F., Siracusano, G., Sanvito, D., & Bifulco, R., “What Did I Do Wrong? Quantifying LLMs’ Sensitivity and Consistency to Prompt Engineering”, NAACL 2025, arXiv:2406.12334
- Thomas Peham (Otterly AI), “Real vs Estimated Prompts: I Analyzed 100s of Real ChatGPT Queries”, 2026, Otterly AI Blog