Democratization is happening whether research approves or not. Who decides what counts as evidence?
Part five of UXR for AI-native teams, a six-post series.
Your PM's six Tuesday interviews already shape the spec, and nobody has written down whether they count
Your PM ran six customer conversations on Tuesday, with decent questions and notes. On Wednesday her spec says "users want to complete verification later" and cites the conversations.
Do they count?
Watch what your organization does with that question. Almost certainly, nothing is written down anywhere. One camp says it isn't research, because there was no method and the questions were probably leading, so it doesn't count. The other camp says six customers beat zero, so of course it counts. Both camps are guessing. The argument replays every quarter with new people in it, and whoever wins on a given Thursday wins by seniority or because the other side got tired.
The question got urgent for one reason. AI tools have cut the cost of doing research to almost nothing for people who are not researchers. An interview guide or a survey that is good enough for most decisions takes a prompt. With the connections that now let a model operate other software (the industry calls them MCP), that survey can be in the field in seconds, sent by someone who has never opened the survey tool. Transcription is free, synthesis is another prompt, and an unmoderated test runs overnight.
So more people are producing things that look like evidence than at any point in the field's history. PMs interview, designers run unmoderated tests, support has years of tickets, and agents summarize all of it into confident paragraphs.
This is democratization, and nobody decided to run it. When running a survey yourself costs less than asking the research team for one, people run the survey. That will keep happening whether the research function approves or not. So the useful question is not whether other people make evidence. They already do. The question is whether what they make gets labeled.
The volume went up and the policy stayed unwritten. An unwritten policy is still a policy: the evidence that wins is whatever sounds most confident. That is drift, statements about users with nothing behind them, treated as knowledge.
The earlier posts made all of this faster, with a memory of claims the organization can ask and a front door where every question gets a price before anyone answers it. The faster evidence moves, the less time anyone has to check it. So the organization needs a written answer to a question it has argued about, without writing anything down, for years: who decides what counts as evidence?
Ban it, allow everything, or publish the standard, and only the third one works
There are three ways an organization can respond to evidence made by people who are not researchers.
Ban it. Only evidence made by researchers counts. This feels like protecting quality, and I understand the feeling. People work around it within a quarter. The PM's interviews still happen, still shape her spec, and still swing the Thursday decision. They just do it without a label, so nobody can check them. A ban doesn't stop evidence from non-researchers. It stops the labeling, which was the only part protecting anyone.
Allow everything. There is no standard and no label, so six interviews become "users want" by Friday, an agent's summary of a dashboard becomes a finding, and whoever sounds most confident wins. Plenty of well-meaning democratization programs land here by accident, because they confused removing the gatekeeper with removing the standard.
Publish the standard. Write down, on one page, what something has to meet before it can be called evidence: what was observed or asked, of whom, how many, with the source attached and the confidence stated. Anyone who meets it can submit evidence. It counts at the weight its label says, with their name on it. Researchers hold the standard and audit against it. They stop being the only people allowed to make evidence.
The PM's six interviews count, as exactly what they are: stated preferences, six people, one interviewer, indicative at best, her name attached. That is not a consolation prize. That is the standard working.
Whether six interviews are enough is a separate question, and it has a plain answer. Evidence is good enough when its strength matches the cost of being wrong on the decision it serves. Six stated preferences are good enough for a copy change that ships Thursday and can be reverted Friday. They are not good enough for a verification redesign that costs a quarter of engineering. The label is what lets the team tell the difference: it says how strong the evidence is, and the decision says how expensive a mistake would be.
Labeling beats gatekeeping for a structural reason, not a diplomatic one. Labeled evidence can be weighed, challenged, and corroborated. Unlabeled evidence wins by sounding confident, and a statement with nothing behind it can sound exactly as confident as one with a study behind it.
The standard governs what may be called evidence, never what a team may decide
One distinction decides whether the standard holds up when a deadline arrives, so it goes in the policy's first paragraph. The standard governs what may be called evidence. It never governs what a team may decide. Whether evidence is good enough for a decision is the team's call, made against the cost of being wrong. Whether something is evidence at all is the standard's call.
Teams can ship on judgment alone, tonight, and the standard has nothing to say about it. What the standard refuses is the mislabel. The team may proceed on judgment, and it may not call judgment "user evidence" in the spec.
The moment a standard starts blocking decisions, it becomes the old two-week intake queue under a new name, and people work around it the same way. A standard that only checks labels never has to win an argument about a launch, which is why it survives every launch.
Machine-made evidence lands in one of three rows, and only generated respondents are out
The first kind is human evidence that a machine processed. Humans were observed or asked, and the machine transcribed, clustered, or summarized what they said and did. This counts as user evidence, with the human sessions cited as the source and the processing step named. It is where last week's speed came from. The machine processed the evidence; it didn't make it. Citing the sessions is what keeps that true.
The second kind is a pattern a machine found. A model noticed structure in behavioral data, and no human said or did anything new. This counts, labeled as inference, at reduced weight, until a human observation confirms it. It is useful but weak. A model noticing that older users abandon camera flows is a lead. A lead becomes evidence when a human observation confirms it, not when it is phrased confidently.
The third kind is a respondent a machine generated: synthetic users, simulated interviews, generated quotes. This does not count as user evidence at any weight, because no users were involved. There are no exceptions, for the reason from last week. Whatever a simulation returns, you can't tell a finding from a fluke without checking against humans, and at that point the check was the evidence and the simulation was overhead. A generated quote is a sentence nobody said. Evidence needs somebody who said or did something.
Enforcement is one section in the spec template: cite the evidence or sign the waiver
A published standard with no consequences is a PDF. The enforcement is small and lives where work already happens. The spec template gets an evidence section, and every claim about users in the spec either cites its evidence or carries a waiver, in the shape from last week.
Proceeding on judgment; risk accepted by [name]; watching [metric] for [window].
Nothing is blocked. A team with zero evidence ships anyway, tonight, with a named waiver. But every spec now shows what it knows and what it assumed, so a mislabel is visible. And "risk accepted by" does the same work it did last week: people read that phrase twice before typing their own name after it. The label is the enforcement.
Every decision records which evidence or which waiver it rested on
The last piece costs one sentence per decision. When a call gets made, the record says why: which evidence, or which waiver.
Two months later, when the post-mortem asks what the team believed and why, somebody reads the answer instead of reconstructing it from memory. An override that writes down "shipping against the abandonment finding; accepting the churn risk; revisit at 60 days" is a decision. The same override unwritten is something that happened to the team, and nobody will be able to say why.
Ask a PM, a designer, and someone in support what would make their evidence count
Send three people, a PM, a designer, and someone in support, the same question: "If you gathered evidence about users yourselves, what would make it officially count here?"
Three different answers, and you've found the argument nobody has written down. Three versions of "huh, I don't think that's written anywhere," and you've found it too.
Either way, the answers are the first draft of the standard's one page, written by the people who will live under it, which is the only way a standard gets followed rather than ignored.
Next week
Every mechanism in this series has assumed a person somewhere. Someone reads the answer at 11pm, someone types their name after a waiver, someone's interviews get labeled and admitted.
But the highest-stakes place research lands, the instruction files agents read while they build, has no person in it at all, on every task, at scale. Those files have pages on naming conventions and nothing on who the users are. The final post writes the missing chapter: what research looks like when the reader is a machine.
If you run the three-message test, I want the answers, and my email is open. Those answers are the data this series runs on.
See you next week.
đ¯ Part five of six. Subscribe and the last one arrives next week. Or don't, and find it in your feed sometime in November, between a hiring announcement and someone's thoughts on leadership.
đ If the fast-research layer underneath this series is the part you need first, that's my book: AI-Powered UX Research, the operating manual for running research at the speed your team actually needs. This series is what I've been thinking since.