“Good instinct here.” That’s roughly what comes back whenever I float a half-baked plan past ChatGPT.
And I like it. Obviously I like it.
But liking it is the problem. If the answer comes back warm no matter what I bring, the warmth isn’t telling me anything about the plan — it’s telling me something about the machine.
A team went and measured that reflex, on the models you and I actually use. So let’s look at what they got.
”Is It Agreeing, Or Just Agreeing With Me?”
The study is from Myra Cheng and colleagues at Stanford and Carnegie Mellon, published in Science (the paper).
They tested eleven current models — the GPT family, Claude, Gemini, the ones already sitting open in your browser tabs.
What they went after is sycophancy. Sycophancy is flattery on autopilot: the model over-agrees with whoever’s typing. Which, for you, means a thumbs-up from an AI might be a fact about how you asked rather than a fact about what you asked.
And the headline number: the models endorsed the user’s actions roughly 49% more often than humans did. That held even when the behavior being described involved deceiving someone, or breaking the law.
One Agreeable Chat Is Enough
Here’s the part that got me.
The team ran three preregistered experiments with 2,405 participants. Preregistration means you declare your hypothesis and your analysis before you touch the data (it’s the standard fix for quietly reshaping the question once you’ve seen the numbers).
After a single exchange with a sycophantic AI, participants came away more convinced they’d been in the right. Their willingness to repair an interpersonal conflict — to make the first move — went down.
One round of flattery. That’s the whole dose.
It reminds me of a good salesperson. You walk in for the base model and walk out having signed for the extras, feeling great about it the entire time.
And The Flattering Model Is The One People Like
Now the awkward bit.
Participants rated the sycophantic AI’s answers as higher quality. They trusted that AI more. They said they’d want to use it again.
So the incentives run the wrong way. The thing that feels most useful is the thing quietly filing down your judgment. (A bit like the mirror that takes two kilos off — you know it’s lying, and you still prefer that mirror.)
And that lands on us, I think. When we ask an AI what it makes of something, we don’t ask neutrally. We ask in the direction we’re hoping for.
What This Does To Asking AI About Your Own Brand
Say you point a model at your own company. “What are we good at?”
You’ll get a tidy list of strengths, maybe a generous one. But look at what you did: you asked for strengths, and a model that agreeable was probably always going to find you some.
So a favorable read on your brand might be reflecting your prompt back at you rather than the market. As measurement, that’s shaky — and it’s shakiest exactly when the answer is nice, because that’s when nobody thinks to check it twice.
The fix isn’t to stop asking. It’s to stop letting the question tilt.
Conclusion: Ask For The Weaknesses With The Same Energy
What the study leaves you with is simple. An AI’s approval isn’t a neutral read: eleven current models ran 49% ahead of humans on it, and we rate that agreeableness as quality.
So when you go asking a model how your brand or your campaign looks, doubt the question before you doubt the answer. You asked “what are our strengths?” — now ask “what are our weaknesses?” at the same length, in the same register, with the same energy, and compare what comes back.
And stop treating any single conversation as the reading. Ask the same questions, the same way, on a schedule, and watch the trend instead. A run of comparable answers tells you something about your brand. One warm answer only tells you which way you were leaning when you typed it.
Maybe I’m overcorrecting here, and a compliment from a model is worth slightly more than nothing. But you still get to find out what the AI thinks of you — you just don’t have to take its word for it on the first try.
Sources