← All notes
Document & OCR Automation7 min

A confidence score is not a gate

OCR and extraction models return a confidence number with every field. It is tempting to treat that number as the gate: above 0.85, auto-post; below, send to a person. On Deerman's purchase-order pipeline and BuilderHelp's supplier-invoice OCR, that shortcut fails in both directions. A high score can still be a wrong part code, and a low score on a field that does not affect the order wastes review attention. The useful gate is a named condition you can explain — not a float.

What a confidence number actually is

A confidence score is the model's estimate of its own certainty on a field. It is not a measurement of correctness against a catalog, a price file, or an arithmetic check. It is also not calibrated the same way across engines, document types, or even fields on the same page. A 0.91 on a vendor name and a 0.91 on a line total are not the same kind of risk wearing the same number.

Production IDP write-ups in 2026 keep rediscovering the same failure: one global threshold looks disciplined until money-moving fields and cosmetic fields share it. Either you flood the queue with safe noise, or you auto-approve the fields that hurt when they are wrong. The number did not become a policy. It became a dial you turn until the wrong thing happens less often — which is not the same as a design.

Deerman: reconcile, don't score

Deerman staff upload vendor PDFs that have to become line items matched against manufacturer price files before anything emails out as a manufacturer PO. The parser can look confident on a code that is well-formed, plausible, and not a part Pfister sells. A confidence gate would let that line through. A reconciliation gate does not.

Lines reach review when they do not reconcile — the code did not resolve, or the match produced something inconsistent. That is a mechanical condition with a reason attached. Staff see the concrete problem next to the document, not a slider between 0.79 and 0.82. Recovery passes never promote a candidate on pattern strength alone; the manufacturer database has to confirm the part exists first. Confidence never enters that decision.

BuilderHelp: posting is the dangerous verb

BuilderHelp's Python Tesseract service reads supplier invoices and posts line items back into the web app, where they can be matched against budget lines. The builder used to do that matching with two windows open. Compressing that work is useful. Auto-filing a misread amount into the budget is not.

Here the gate is about what the automation is allowed to commit. The OCR can propose. Reconciliation and a person still decide what becomes a budget fact. A confidence threshold that auto-posts above some float would trade one saved click for a quietly corrupted number the builder makes decisions from. That is the wrong trade even when the score looks good.

If you must use scores, make them field-shaped

Some stacks only expose a score and no catalog to join against. Then a single document-level cutoff is still the wrong shape. Split by field and by consequence: totals, remit-to, quantities, and part identifiers need stricter bands than a description line. Layer deterministic checks on top — line items that do not sum to the total fail regardless of how sure the model claimed to be.

Route only the failing fields, not the whole document, or reviewers will start approving on pattern. Keep the crop of the source region beside the proposed value. And treat the score as a triage hint that feeds a named rule, not as the rule itself. When a correction lands, store the delta: that is the labeled example that tells you which fields deserve a tighter band next month.

What we ask before we wire a threshold

Which fields, if wrong, create an order, a payment, or a budget lie — and which are merely ugly? Is there a confirming source (catalog, price file, PO number, arithmetic) that can reject a confident wrong answer without a person? What should the review UI ask, in plain language, when that confirmation fails?

If you cannot answer those, you are not ready to auto-post on a float. Build the reconciliation and the review path first. Add scores later as a way to order the queue, if at all. A confidence score can help you prioritize attention. It should not be the thing that decides whether an order leaves the building.

Questions

Why not auto-approve fields above a confidence threshold?
Confidence is the model's self-estimate, not a check against a catalog, price file, or sum. A high score on a wrong part code still ships a bad order. Prefer a named condition — reconcile failed, totals do not add — over a global float.
When are confidence scores still useful?
As triage: stricter bands on money-moving fields, looser on descriptions, and only when layered with deterministic checks. Use them to order a review queue, not as the sole gate that posts into an ERP or budget.
What should a reviewer see instead of a score?
The concrete failure — unmatched code, transposed quantity and price, line items that do not sum — beside a crop of the source region. Ask a question they can answer from the document, not whether 0.82 feels high enough.

Sources

  1. CNCF — Observability for AI agents — Adjacent lesson: green systems can still be wrong; need decision-level checks.
  2. ExtractBee — Confidence scores in document extraction — Per-field thresholds vs document-level gates.
  3. Fathom — Deerman Sales case

Have something to build?

Tell us what you're working on and we'll tell you honestly whether we're the right fit.

Work with us