A wizard behind the curtain is not a coded probe
Wizard-of-Oz tests are trending again — especially for AI-shaped features that are expensive to wire end-to-end. A participant talks to what feels like a working system while someone offstage types the replies. That is a real instrument for interaction risk. It is a terrible instrument for feasibility risk. We have watched teams graduate a polished wizard session into a roadmap item that still had not touched the citation path, the secrecy model, or the sync budget. The curtain answers whether people can work with the loop. The coded probe answers whether the loop can be honest.
What the wizard answers
A Wizard-of-Oz session tests an interaction hypothesis: the participant makes a request, receives a response that looks systemic, and reacts. You learn whether the shape of the loop matches their mental model — what they ask for first, what they trust, what they do when the answer is partial, and when they abandon.
That evidence is hard to get from a survey or a static mock. People will say they want an assistant. They will not show you, until they are mid-task, that they expected a citation, a refuse, or a handoff to a human. The wizard can produce those moments without a model, a queue, or a production datastore.
We treat that as pretotyping for interaction, not as a prototype of the product. Same family as a fake door, different question. The door measures whether anyone reaches. The wizard measures whether the conversation or workflow still makes sense when responses feel real.
What it quietly skips
The wizard can be smarter, kinder, and more consistent than the system you will ship. Latency can be staged. Edge cases can be avoided. Every answer can cite a rule number even when retrieval would have missed. If you only record smiles and completions, you have validated the wizard's performance, not the product's.
RuleCaddie's load-bearing risk was never whether a golfer would ask a rules question. It was whether an answer could be forced to cite a real Rules of Golf number — and refuse when it could not. A skilled operator behind a chat UI can invent that experience in five minutes. A thin coded slice against the actual rulebook is the only thing that can fail the way production will fail.
Veto is worse for the wizard method. Secrecy is a publication and schema problem. If another player's veto can be reconstructed from what the client receives, the game is already broken. No amount of offstage typing proves last-writer-wins, RLS, or realtime channels. The hard part is invisible to the participant and fatal to the product.
When we still run one
We run a wizard when three things are true at once: the underlying system is expensive to stand up, the interaction is load-bearing on the value proposition, and we cannot predict user reaction from a spec review. A triage inbox, a draft-and-edit loop, a routing suggestion the operator must accept or override — those earn the curtain.
Even then, script failure on purpose. A correct response, a partial response, and a wrong response. The findings live in repair and abandonment, not in the reel where the wizard stuck the landing. Write the wizard's decision rules before session one and revise them between sessions; otherwise every participant met a different "system."
Timebox it and name the keep/kill signal up front, the same way we write kill criteria before the first commit. If participants cannot recover from a partial answer without a moderator rescue, that is evidence — not a reason to make the wizard smoother.
Prefer the thin coded slice when the seam is the risk
If vibe-coding or a weekend spike can put a dishonest-but-real vertical slice in someone's hands, prefer that when the open question is technical or domain-honest. Real rules text. Real Postgres policies. Real offline packages. Real webhook latency. The participant may see a uglier UI; you get failure modes the wizard would have papered over.
That is the difference from a fake door and from a wizard. The door is high fidelity for attention. The wizard is high fidelity for conversational shape. The coded probe is high fidelity for the seam that can lie — retrieval, permissions, sync, print acceptance, inventory truth.
On BuilderHelp-shaped AI work, we have used human review loops in production on purpose. That is not Wizard-of-Oz research; that is a designed gate because the cost of a wrong file is real. Do not confuse an operator in the loop as a product feature with an operator in the lab as a substitute for building the seam.
Name the instrument before the session
Before anyone sits behind the curtain, write four lines: the one question, which failure mode would kill the idea, whether the artifact is throwaway, and which instrument can produce the signal. If the question is "will they understand the loop," a wizard can earn its keep. If the question is "can we be correct under constraints," schedule a coded probe.
Keep the outputs in different notes. Wizard results update an interaction note: mental models, repair strategies, abandon points. Probe results update a decision note: keep, kill, or narrow. Mixing those documents is how a good session becomes three months of architecture with the hard proof still outstanding.
Call the curtain what it is. It is a way to rehearse an interaction with a human standing in for an expensive system. It is not feasibility evidence, not a keep, and not permission to skip the seam. When the session ends, you still owe the real risk a real instrument.