Synthetic Users Are Influencing Your Design Decisions. New Research Says They're Right About as Often as a Coin Flip.
Back in March I wrote up the largest systematic review of synthetic participants ever conducted. It came out of UXtweak Research and the Slovak University of Technology in Bratislava, it covered 182 studies across nine domains, and its conclusion was blunt: synthetic users don't work. The thing I kept circling back to was that the authors had every commercial reason to find the opposite and the evidence just wouldn't let them.
Same group. Kuric, Demcak, Krajcovic. They just dropped a new preprint, and it's a different kind of paper. The review was a survey of everyone else's experiments. This one is their own experiment. They collected 29 real preference tests run by 29 different organizations on their platform, 2073 real participants, 78 tasks, and they used that as ground truth to see whether an LLM steered to simulate those exact people would land on the same design preferences the humans did.
A quick note on the word preprint, because I had to say this last time too and I'll keep saying it. Not peer reviewed yet. The methodology here is rigorous enough that I'd expect it to survive review, but the discussion framing hasn't had external pushback. Read it as strong evidence, not settled science. The people quoting a preprint on LinkedIn like it came down from a mountain are the same people who'll quote a vendor deck the same way.
Here's why this paper matters more than a fourth "LLMs are shallow" study. It isolates the single most common way a design team actually reaches for synthetic feedback. Not political polling. Not consumer surveys. The A/B. The "which of these two screens do people prefer, and why." The most innocent, most frequent, most load-bearing decision a design team makes on a Tuesday. And it tests that use case with real design stimuli, given to the model in full, with the model told exactly who the audience was.
The AI still couldn't pick the winner. But the interesting part is how it failed.
It didn't guess wrong. It refused to choose.
You'd expect the failure mode to be "the model picks the option humans didn't." That happens sometimes. Chi-squared tests found the synthetic and real preference distributions genuinely different in 44% of tasks, and the model agreed with the humans on the single most popular option only 53% of the time. A coin. Rank-order agreement averaged 0.53. So yes, a fair amount of straightforwardly wrong.
But the finding that should actually change how you think about this is the entropy result. The synthetic preferences were consistently flatter than the human ones. Normalized entropy was 0.93 for the AI versus 0.86 for real people, and that gap is statistically real. In plain terms: when humans had a clear favorite, the model spread its vote around and made the options look closer to equally good than they were. The paper's phrase is that LLMs "inflate indecisiveness." I'd put it harder. The model manufactures a tie.
Occasionally it ran the other way, humans split evenly while the model locked onto one option with no variation at all. But the dominant distortion was false balance. The AI took a question with a real answer and handed back "they're all pretty good, each with its strengths."
Sit with why that's worse than a wrong answer.
A wrong answer is falsifiable. If the AI says users prefer variant B and you have any instinct, any prior, any residual respect for your own eyes, you can push back. A wrong answer picks a fight. A tie doesn't. A tie is the most agreeable, least confrontational, most quietly corrosive thing you can hand a designer, because it launders whatever they already wanted to do. You built variant B. You like variant B. The synthetic panel says all three variants are roughly comparable. So there's no evidence against B. Ship B. The AI ratified your bias instead of correcting it, and dressed the ratification up as data.
This connects to something the design literature has known for years and the paper cites directly. Designers evaluate their own concepts with a confirmation bias (Nikander and colleagues called it the preference effect back in 2014). The whole point of testing preferences with real users is to introduce friction against that bias. A synthetic panel that flattens every distribution into a soft tie doesn't reduce the bias. It removes the friction and keeps the confidence. You walk out more sure of the thing you walked in wanting, and now you can say you tested it.
I've written before that the danger with synthetic users isn't that they're obviously fake, it's that they're believable enough to launder a bad decision. This is the mechanism, made specific. It's shallow in exactly the direction that tells you what you want to hear.
Every fix, already tried
I know the reflex, because it's the reflex in every comment section under every one of these posts. "You used the wrong model." "You didn't prompt it right." "Rich personas fix this." The paper went through the list. This is the part I'd tape to the fridge.
A reasoning model doesn't help. They ran the baseline on GPT 4.1 and then swapped in GPT 5.2, the current flagship reasoning model. Significant distribution differences went from 44% to 38%, which sounds like movement until you see it isn't statistically significant. First-choice agreement nudged to 65%. The chain-of-thought machinery that's supposed to make the model reason its way to a human answer bought essentially nothing on the preferences, and in the justifications it actually made things worse, compounding the model's biases and flattening semantic diversity further. "The illusion of thinking" is doing real work here. More visible reasoning, same blind spot.
Sampling parameters don't help either. Dropping temperature to 0.2 left the difference at 41%. Dropping top_p to 0.2 left it at 38%. And crucially, turning down the randomness did not fix the false-balance problem. Entropy stayed pinned around 0.94. You cannot tune your way out of this with the knobs everyone reaches for first.
Specific, detailed personas don't help. This is the one that should end a particular genre of LinkedIn advice. They compared rich personas (modeled demographics, personalities, the actual screening and pre-study data from the real samples) against generic featureless personas with none of that. The results were nearly identical. 46% divergence with rich personas, 44% at baseline, not a significant difference. All the persona craft, all the "give it a name and a backstory and a Myers-Briggs type," added nothing measurable to whether it could predict the design people preferred.
And single-persona simulation is actively worse. When they simulated one detailed individual at a time instead of an aggregated "mega-persona" of the whole audience, divergence jumped to 91% of tasks. Entropy collapsed to 0.28, meaning the model got rigidly deterministic, and it explored far fewer of the available options. The mega-persona approach looked better only in comparison to this, and the authors are honest about why: aggregating a population probably just triggers the model's encoded expectation that a crowd should be diverse, so it performs diversity. That's not understanding a population. That's pattern-matching the word "population" to more spread. When you ask it to actually be one specific person, the machinery falls apart.
So: not the model, not the temperature, not the personas, not the individuation. The thing people assume is a prompting problem is not a prompting problem.
The justifications get embarrassing
The preference votes are half the story. The other half is the "why," the open-ended follow-up where a real participant tells you what made them choose. This is the stuff design teams actually want from a synthetic panel, the qualitative color. It's also where the model's lack of a mind is hardest to hide.
Lexical similarity between synthetic and real justifications sat around 0.25. Semantic similarity around 0.70. So the AI is using different words to gesture at vaguely overlapping themes, with meaningful gaps underneath. The qualitative read is worse than those numbers suggest, and the patterns are worth naming because they're the same ones you'll see in any AI-generated feedback if you know to look.
It praised things that were identical across every option, complimenting the visibility of a button that didn't change between variants, which tells you it isn't comparing anything, it's just talking. It hallucinated the task itself and then praised elements for serving a purpose that didn't exist. It leaned on empty abstractions: clarity, vibe. Real people anchored on specific differences, why an element belongs on the left and not the right. The model reached for the generic because it has no actual perception to report. It produced the tell-tale shape of a generic positive followed by a generic negative with no logical link between them. The paper's example: "Visibility is great, but makes me hesitant to recommend." That's not a person. That's a model balancing a sentence. It faked prior experience, claiming familiarity with a design "from other websites" with zero specifics, because a real person would have an actual memory and it's approximating the shape of one. And it handed back impersonal guidelines instead of preferences, suggestions to improve accessibility for non-native speakers, that kind of thing. RLHF again. The model is constitutionally helpful, harmless, and honest, so it will not do the things real testers constantly do. Openly dislike everything. Refuse to engage. Admit a pure gut aversion. Decline to share information because it doesn't feel like it. Real preference data is full of that friction. The model sands it all off.
On complex stimuli like dashboards, it got worse, latching onto a random visual element and building its whole argument around it even when that element was irrelevant. On copywriting variants that said the same thing in different tones, it missed the nuance entirely, which is exactly the subjective register where preference lives.
The one useful artifact
If you take one practical thing from this, take the fakery checklist. Because the flip side of "the model produces recognizable garbage" is that the garbage is detectable, and that matters right now for a reason that has nothing to do with synthetic users as a research method: your real studies are getting polluted by bots and LLM-assisted survey farmers, and the fake open-text answers have the same fingerprints as the synthetic ones in this study.
The authors list the tells. Disregard for the key differences between the options being compared. Hyperfocus on one or a few specific visual elements. Generic properties assigned to a design with no implication drawn from them. Elaboration that adds length but no depth. No narrative thread between statements. Overpraising, or a suspiciously precise balance between praise and criticism. An absence of genuine subjective evaluation. Assumptions that don't fit the actual context. Off-topic suggestions. Internal inconsistency.
That's a usable screen for your next survey's open-text field, whether the pollution came from a vendor's synthetic panel or from a person running your incentive through ChatGPT. Print it.
Where I'd push back
I try to be fair to the thing even when I like its conclusion, so: the primary experiment is GPT only. GPT 4.1 and GPT 5.2. They reference Claude and LLaMA but didn't run them here. On its own that's a real limitation, and I'd want to see it before I said "all models." The reason I'm comfortable generalizing anyway is that this doesn't sit on its own. It sits on top of the 182-study review from the same group that spanned GPT-2 through GPT-4o, Claude, Gemini, Mistral, Qwen, DeepSeek, and found the same failure shape across all of them. The narrow model set in the experiment is buttressed by the wide one in the review.
Preference testing also carries its own baggage as a self-report method. Showing people multiple options changes behavior, and the authors know it. They chose ecological validity over control on purpose, real studies from real organizations with all their messy heterogeneity, and I think that was the right call for a paper about what happens in practice. But it's a tradeoff, and a cleaner lab version might sharpen or soften specific numbers. And because the underlying studies are proprietary, the results are aggregated in a way that limits reproducibility. They're transparent about that.
One more thing worth naming, because I did the same thing with their earlier review. Two of these three authors are UXtweak Research, and every one of the 29 studies here ran on UXtweak's own platform, through the tool they sell. The declared competing interest is none. I don't think that's dishonest, and the methodology doesn't read like it was built toward a foregone conclusion. But sit with the incentive for a second: a finding that AI can competently reproduce what your own paid preference-testing tool measures would have been a genuinely valuable feature to ship. They tested it on their own clients' real data, and it said no. That's the same shape of against-interest evidence that made the review credible, and it applies here too.
None of this touches the core finding. If anything the ecological messiness makes the consistency of the failure more damning, not less.
What this actually changes
For a while the honest position on synthetic users has been "unreliable, use with extreme caution, know what you're losing." I've held versions of that myself, including a piece where I walked back some earlier certainty because the augmentation case deserved a fairer hearing. This paper sharpens the caution into something more specific and more useful.
The failure isn't random noise you can average out. It's a directional distortion. LLMs flatten preference distributions toward false balance, and false balance is precisely the output that does the most damage in the one moment design teams reach for them, because it disguises a non-answer as a green light and hands it to someone already biased toward their own work. The believability isn't a bonus that offsets the shallowness. The believability is what makes the shallowness dangerous, because it gets the flat tie taken seriously.
If your process uses synthetic preference feedback anywhere near a real decision, the question to ask isn't "is it accurate." It's "what does it do to the confidence in the room." And the answer this paper gives is that it raises the confidence while lowering the information. That's the worst possible trade for anyone whose job is supposed to be introducing evidence into a decision, which is to say, the whole job.
Telling your team which design people actually prefer, and why, against what they were hoping to hear, is the hard part. Producing something that looks like preference data has never been the hard part. This is the study that shows the model doesn't just miss the hard part. It gives you the opposite and makes it look like agreement.
🎯More like this at The Voice of User. No vendor pitches. No hype cycles. Just what the evidence actually says. Subscribe.