Not every question deserves a study
Part four of UXR for AI-native teams, a six-post series.
A PM at 11pm has three options, and nobody has ever priced them for her
A PM is mid-spec at 11pm and hits a question she can't write around: do users actually read the verification explainer before they start uploading?
She has three options, and she doesn't know they're options, because nobody has ever priced them for her. The answer might already exist somewhere in the company's research, ninety seconds away. It might not exist yet, and then someone should decide, out loud, whether it's worth making. Or it might not matter enough to chase, and then the move is to say so and name the number that will catch a mistake.
What she does instead is what everyone does. She asks the company's ChatGPT, the one connected to the drive. It gives her a confident paragraph stitched from four old decks, with no dates, no scope, and no way to tell which parts are still true. It goes into the spec.
Nobody priced the question. So nobody knows whether the company already had a better answer, or whether this was the question the whole bet rests on. Those are the two expensive research mistakes, and neither is a bad study. A function without a front door makes both constantly, in opposite directions, while feeling busy.
This post is the front door: what happens when a question arrives. The three posts before it built what the front door stands on, the claim as the unit of knowledge, the organization's memory where claims live, and the map of the five doors. There's also a fight at the end, and I scheduled it on purpose.
Every question has one of three prices: minutes, days, or an owned risk
So every question gets a price before anyone answers it, and there are only three. The first price is minutes. The answer already exists in the organization's memory (1, 2), so someone looks it up and the asker gets it with its receipt attached, meaning the evidence it came from.
The second price is days. The answer doesn't exist and the decision matters, so someone makes new evidence on purpose.
The third price is nothing, which is not the same as free. The question wouldn't change the decision, so the team proceeds without evidence, and the risk gets written down with a name on it and a metric to watch.
Routing is the discipline of not paying days for a question that minutes could answer, and not paying minutes for a question that decides the bet.
Routing is five questions, asked in order, in about five minutes
The order matters, and the reason is in question two.
1. What decision does this serve, and when is it being made? If there is no decision, there is no route. Curiosity is welcome, and browsing the organization's memory costs nothing; the budget follows decisions.
2. What are the stakes? If the decision is reversible and cheap, the stakes are low. If it is costly to reverse, touches user trust, or defines the bet, the stakes are high, and a bet-defining question gets sessions with humans, full stop. Stakes come before the memory check on purpose. If you check memory first, you will be tempted to accept a lookup as the answer to a bet-defining question, because the lookup is right there and the study is not.
3. What's the gap? Now check the organization's memory. Is the distance between what the decision needs and what memory holds zero, small, or decisive? A zero gap routes to the lookup. A small gap routes to the lookup plus a note on what the claim doesn't cover. A decisive gap routes toward new evidence.
4. What answer would change the decision? This is the question that does the most work. If no conceivable answer changes what the team will do, the route is nothing, and the risk gets accepted in writing. (This route made me flinch for years. Then I noticed how often "we should test this" actually means "nobody wants to own this call," and the flinch became a tell.)
5. What can evidence made in that time claim? When is the decision really being made, not officially, really? The answer sets what the evidence can claim. A quick pass with a handful of users can eliminate weak options by Wednesday. It cannot produce a finding you'd bet the quarter on. Promising that it can means inflating the rigor because the deadline asked for it.
Every question leaves by one of three routes: the organization's memory, new evidence, or nowhere with the risk owned in writing. That is the entire front door.
Some questions deserve no research, and saying so takes one written sentence
Some questions deserve no research at all, and that is the part nobody says out loud. Shipping without evidence is fine. Shipping without evidence and pretending otherwise is not.
So the route that feels like doing nothing is the one that needs its artifact spelled out. The artifact is one sentence with three parts.
Proceeding without evidence on [the question]; risk accepted by [name]; watching [metric] for [window].
It takes thirty seconds to write, and it turns a shrug into a decision with an owner and a number someone is watching. Half the value is the watching. The other half is that people read "risk accepted by" twice before typing their own name after it.
Some questions get re-routed to a quick study at exactly that moment, which is the sentence doing its job. The reverse happens more often. A lot of studies get run so that nobody has to write that sentence.
About half of all questions get answered from the organization's memory, and those answers follow three rules
Roughly half of everything that hits a working front door gets answered from the organization's memory, which is why last week's post exists. But an answer from memory has to follow rules, or it becomes the drift it was supposed to replace: statements about users with nothing behind them.
First, every answer shows where it came from.
Second, the answer's confidence lives in its wording, not in a label. A well-evidenced claim is written as a plain statement. A thin one hedges in the sentence itself. A guess says it's a guess. That way, when someone pastes the answer into a spec, the hedge goes with it.
Third, the system may arrange what it knows, group it, and put it in the asker's words. It may never stretch a claim beyond what was studied, and it may never merge two claims into a third that nobody made. A system that does that is making things up, with citations.
And "no claim covers this" is a complete and useful answer. It hands the question straight back to the router.
A 72-hour study is possible because only the logistics compress
The evidence route survives at AI-native speed because the two-week study was always about twelve hours of evidence inside two weeks of logistics. Only the logistics compress.
Agents and standing infrastructure shrink the logistics from days to minutes: recruiting, scheduling, transcription, synthesis, and deck production. What never compresses is the method: people, their attention, their consent, and the thinking.
So the 72-hour study exists, and it's boring in the right ways. On day one the instrument gets designed, piloted on a colleague, and sessions get booked from a standing participant pool. Day two is ten or so moderated sessions, with a quantitative pull running alongside. Day three triangulates the two, writes the findings back into the organization's memory as claims, and the spec cites them by Thursday.
This assumes two pieces of infrastructure: humans reachable by tomorrow, and a test setup that can run whatever got built this morning. Both are buildable. Neither is this post.
One kind of moment inside these sessions is the whole argument for the human moderator. Mid-session, a participant finishes the task and goes off-script: she'd poke around for a while before putting her own money into a new app anyway. The instrument has no question about this. The agent taking live notes records it and moves on, because agents follow instruments. The human follows her for ninety seconds, and those ninety seconds produce the best evidence of the week. The agent runs the instrument. The human notices when the instrument is the wrong size for what's happening.
Synthetic users lose on arithmetic once evidence from humans is fast
The market will offer you a shortcut here, so here is the position, on the record. The pitch has several names: synthetic users, simulated respondents, digital twins of your customers. Under every name it is speed without humans, and I decline it, not out of purity but out of arithmetic.
The case for simulation rests on one premise: that evidence from humans is slow. Everything above exists to make that premise false. When a standing pool fills sessions overnight and a 72-hour study produces triangulated findings, the simulation has nothing left to sell. Its market was the two-week recruit, and the recruit is gone.
What remains is the cost side, and it doesn't improve with model quality. A simulation is a model of the users your data already describes. Anything new it appears to find was already in the organization's memory, and you're running a study precisely because memory came up empty.
Whatever a simulated study returns, you can't tell a finding from a fluke without checking it against humans. At that point the check was the study, and the simulation was overhead.
The failure is one-sided too. A false negative removes a good option and nobody notices. A false positive ships. Either way the drift carries your function's name on it. And notice what a simulation would have done with the off-script moment above: run the instrument perfectly. The instrument was the wrong size.
The machines do belong here, because this is not a purity position. Agents draft instruments, run screeners, schedule, transcribe, take live notes, and assemble synthesis, and that is where the speed comes from. Humans are the only source of the evidence. (I know a researcher arguing this reads as self-serving. Check the arithmetic anyway; it comes out the same whoever benefits.)
Log your next ten questions before answering any of them
Before you build any of this, run the cheap experiment: log your next ten questions before answering anything. One line each: who asked, what, what decision it serves. Then route them through the five questions, and count how many the organization's memory could have answered, how many deserved no research at all, and how many needed new evidence.
Most functions that run this land near half lookups, a good number that deserved nothing, and only a few that needed new evidence. That means the operating model where everything became a project was mispriced at the front door all along. The log is ten minutes of work and it's the most clarifying instrument in this series so far.
Next week
Routing, when it works, creates the series' least comfortable problem. The faster and more open this system gets, the more the organization has to trust what gets admitted as evidence, because there is less time than ever to check.
Your PM ran six interviews on Tuesday. Do they count? Your organization has no written answer, which means it has a policy, and the policy is drift. Next week: who gets to make evidence, the table that names what does not count, and the one distinction that lets a standard hold up when a deadline arrives.
If you route ten questions and the log surprises you, my email is open, and I mean that. Those logs are the data this series runs on.
See you next week.
đ¯ This is part four of UXR for AI-native teams, a six-post series. Subscribe, and the rest arrives as it publishes.
đ If the fast-research layer underneath this series is the part you need first, that's my book: AI-Powered UX Research, the operating manual for running research at the speed your team actually needs. This series is what I've been thinking since.