An Accessibility Inspector pass is not a VoiceOver test
Xcode's Accessibility Inspector will give you a clean audit on a build that still fails the first real VoiceOver session. The audit is good at missing labels, contrast, and hit targets. It is weak at focus order, rotor paths, custom drawing surfaces, and the habit of leaving Screen Recognition on so the OS invents labels your code never shipped. Apple's App Store VoiceOver evaluation is explicit: common tasks have to work with VoiceOver alone, on the devices you support. That is a device test, not a green check in the inspector.
What the audit is allowed to prove
We run Accessibility Inspector audits early and often. Missing accessibilityLabel values, insufficient contrast, and controls that are too small are cheap to catch in CI with XCUITest's accessibility audit APIs, and catching them before TestFlight saves real review cycles. For Endeo, RuleCaddie, and Veto, that pass is part of the pull-request gate the same way a simulator smoke is.
The trap is treating a clean audit as VoiceOver support. The inspector walks an accessibility tree and scores properties. A VoiceOver user walks gestures, rotors, and custom actions on a live screen that may regroup, dismiss, or redraw under them. Those are different claims. A green audit proves the tree is annotated. It does not prove a sermon note can be opened, a rules answer can be followed to its citation, or a veto can be cast without sighted help.
Turn Screen Recognition off before you trust the tree
Screen Recognition is an OS feature that guesses labels from pixels when your app left the accessibility tree thin. It is a gift for users stuck in an unlabeled UI. It is poison for a test plan. If it is on, the inspector and VoiceOver can sound fine while your buttons still ship without labels — the machine is covering for you.
Before any audit we treat as a shipping signal, Screen Recognition is off. Then we listen again. The silence where a label should be is the bug. Keeping recognition on during QA is how teams ship "Supports VoiceOver" feelings that evaporate the moment a user has recognition disabled or the guess is wrong on a custom canvas.
Custom surfaces fail in ways buttons do not
Endeo's core loop is PencilKit on a canvas. Strokes are the source of truth; OCR text is a disposable index. That architecture is right for handwriting, and it is exactly where Accessibility Inspector gets polite. A canvas can pass as a single element with a generic label while the note list, session clock, and search results around it are the parts a VoiceOver user actually needs to complete the task.
RuleCaddie's citation cards and Veto's private reveal flow have the same shape: the interesting UI is not a stock Form. Grouping, custom actions, and whether focus jumps into a modal or leaks into the dimmed background are the failures. The audit may be green. The rotor path through "headings" or "containers" still strands someone mid-task.
So the device checklist names the custom surfaces: canvas adjacent chrome on Endeo, citation navigation on RuleCaddie, join-and-reveal on Veto. If those paths were only exercised with a mouse in the inspector's simulation mode, they were not exercised.
How we run a VoiceOver session test
Triple-click the side button into VoiceOver on a physical iPhone and iPad from the support matrix. Simulator VoiceOver simulation is useful for reading order while you iterate; it is not the gate. Gestures, braille display behavior, and real focus after interruptions belong on glass.
Pick the common tasks the App Store evaluation cares about — the ones a new user has to finish without sighted help. For each task: start from a cold screen, swipe through the full forward path, reverse it, try a rotor jump, trigger the primary action with double-tap, and dismiss any modal with the escape action. Watch whether a background refresh resets the VoiceOver cursor to the top of a list. That reset is a product bug, not a VoiceOver quirk.
Write fail criteria before you start. Examples we use: an unlabeled destructive control; focus trapped in a dismissed sheet; a canvas that steals the swipe chain with no way to reach Save; a status toast that never posts an accessibility announcement. A build that only fails one of those is still a failed VoiceOver test.
What we treat as done
A release candidate is VoiceOver-ready when Screen Recognition was off for the audit, the automated accessibility checks are green, and a named set of common tasks was completed on device with VoiceOver alone on each form factor you claim. The device note lists which tasks, which OS versions, and which failures were filed.
Accessibility Inspector stays in the loop for regressions on labels and contrast. It does not get to sign the VoiceOver claim. Apple's nutrition-label style evaluation is task completion under VoiceOver, not a property checklist. We treat it the same way we treat Pencil and thermal session tests: the tool that runs on the Mac proves structure; the session on hardware proves the product.
Questions
- Is Xcode's Accessibility Inspector enough before shipping?
- No. Use it to catch missing labels, contrast, and hit targets early — then complete common tasks with VoiceOver on a physical device with Screen Recognition off.
- Why turn Screen Recognition off during testing?
- Screen Recognition invents labels from pixels when your tree is incomplete. That masks missing accessibilityLabel values and makes an unlabeled build sound fine until a user runs without recognition.
- Does simulator VoiceOver count?
- It helps while iterating on reading order. The shipping gate is still a real VoiceOver session on the iPhone and iPad form factors you support, including rotors, modals, and focus after content refresh.
Sources
- Apple — VoiceOver evaluation criteria — App Store Connect accessibility evaluation: common tasks with VoiceOver alone.
- Apple — Support VoiceOver in your app — Platform guidance for labels, traits, and focus.