A higher fidelity is not a better answer
Fidelity ladders are trending again in prototyping playbooks — paper, wireframe, clickable, hi-fi, coded. Useful as a menu. Dangerous as a sequence. We keep seeing teams treat each rung as homework: finish the sketch, then the grayscale, then the polished Figma, then "build the real one." That is calendar progress, not learning. RuleCaddie, Veto, and Private Reserve only moved when the next artifact matched the risk still open. A prettier mock of a settled flow is not a finding. A thin coded probe of an unsettled seam is.
Fidelity is a menu, not a syllabus
A ladder diagram implies you start at the bottom and graduate upward. That framing quietly smuggles in the wrong goal: completeness of the artifact rather than sharpness of the question. Paper answers mental model. A mid-fi clickable answers whether someone can reach the goal without a narrator. Hi-fi answers feel and brand fit. A coded slice answers whether the stack, the data, or the offline path will hold. Those are different instruments.
The failure mode is sequential habit. The team finishes a clean wireframe, so the next ticket is "raise fidelity." Nobody asks whether the open risk still cares about layout. If the open risk is citation honesty, secrecy, or sync under a budget, climbing into polished frames burns days and still leaves the seam untested.
We treat the ladder as a picker. Name the risk, pick the lowest rung that can falsify it, timebox the work, and stop when the decision is made — even if the artifact still looks unfinished.
What each rung is allowed to prove
Paper and whiteboard are for argument shape. Can a person explain the product in one pass without inventing screens that paper cannot hold. If they cannot, no amount of Figma polish will fix the confusion — it will only decorate it.
Clickable mid-fi is for unaided pathfinding. Put five people in front of the flow and watch where they stall. That is the right tool when navigation and task structure are the risk. It is the wrong tool when the risk is whether Postgres row-level secrecy holds or whether a streaming answer can resume after a drop.
Coded probes are for honesty under real constraints. RuleCaddie's citation path, Veto's secrecy model, Private Reserve's ingest seam — those were never going to be settled by a hi-fi prototype that faked the hard part. The coded slice can be ugly. It cannot lie about the seam.
Timebox the rung, not the roadmap
A fidelity climb without a clock becomes a mini-product. The wireframe gains edge cases. The hi-fi gains empty states. The coded "spike" grows a design system. None of that is wrong if those details are the risk. Usually they are not — they are comfort work while the scary question waits.
We write the question and the kill or keep signal before the clock starts, the same way we write kill criteria before the first commit. Two days on a clickable flow to see if the task is findable. One day on a wizard session if the interaction is the unknown. Three days on a coded probe if the seam is. When the clock ends, the output is a decision record and a next risk — not a prettier file.
If the timebox expires without an answer, that is also a result. It usually means the question was too wide, the fidelity was mismatched, or the seam needs a different instrument. Extend only when the next hour would change the decision.
Do not promote polish into proof
Stakeholders react to fidelity. A glossy prototype reads as "almost done," even when every backend call is a hard-coded string. That social pressure is how teams skip the coded probe: the room already feels convinced.
We label the artifact with the question it answers and the questions it explicitly does not. A fake door measures curiosity. A wizard measures whether the loop is usable with a human behind the curtain. Neither measures feasibility. Saying that out loud in the review keeps a CTR or a polished click-through from becoming a schedule.
Promotion happens only when the open risk changes. If usability is settled and the seam is not, the next artifact is a coded probe — not a second hi-fi pass with better shadows.
Match the rung to the failure mode
Before we open Figma or Xcode, we ask which failure would waste the next month of build. Wrong mental model — stay on paper. Wrong task structure — mid-fi with users. Demand fog — fake door, not a prototype. Interaction fog for an AI loop — wizard, then stop. Domain-honest failure — thin coded probe against the real constraint.
That mapping is the whole practice. The ladder is useful because it names options. It becomes harmful when climbing it feels like progress. Progress is a smaller set of open risks, written down, with the next instrument already chosen.
If you need a default: spend the least fidelity that can embarrass the idea this week. Anything shinier before that embarrassment is decoration.
Questions
- When should we raise prototype fidelity?
- Only when the open risk has changed and the current rung cannot falsify it. Settled navigation does not justify a hi-fi pass if the seam is still untested — jump to a thin coded probe instead.
- How long should a fidelity rung take?
- Timebox to the decision, usually one to three days. Write the question and kill/keep signal before the clock starts. Extend only if the next hour would change the call.
- How is this different from a fake door or a wizard test?
- A fake door measures demand. A wizard measures interaction with a human simulating the system. Fidelity matching is the meta-rule: pick the lowest instrument that can answer the failure mode you actually fear.
Sources
- Software Prototyping 2026: fidelity ladder framing — Industry ladder language; we reject the implied sequential climb and the cost framing.
- Spike work as timeboxed knowledge, not delivery