Skip to content

Setting a face-match threshold without guessing

A face-match threshold is a policy choice between wrongly accepting impostors and wrongly rejecting real customers. Set the pass mark from labelled pairs of your own traffic, add a bounded review band just below it, and revisit both as your users and cameras change.

Charles Archibong

, Co-founder

· 5 min read

Headline "Face-match thresholds" beside an illustration of a face outline mapped with landmark points, on a soft lavender gradient.

Key takeaways

  • Raising the pass mark trades fewer impostors accepted for more genuine customers rejected; there is no free setting.
  • A review band sends near-misses to a person instead of forcing a yes or no on an ambiguous score.
  • Bound the review band above by the pass mark, or clear matches will land in the review queue.
  • Calibrate on labelled pairs from your own traffic, and recheck when your users or devices change.

Choose a face-match threshold by deciding which mistake costs you more, then measuring where your own genuine and impostor scores actually fall. The pass mark is not a quality setting you turn up for "more security". It is a trade between two errors: accepting someone who is not the person on the ID, and rejecting someone who is.

Most teams need two numbers, not one. A pass mark above which a match is accepted automatically, and a review floor below which it is rejected, with the space between them sent to a person. The rest of this article is how to set both without guessing.

What is the threshold actually trading?

A face-matching system compares two images and returns a similarity score. Genuine pairs (the same person) tend to score high; impostor pairs tend to score low. The two groups overlap, and every threshold cuts through the overlap somewhere.

  • Set the pass mark higher, and fewer impostors get through (a lower false match rate), but more genuine customers fail (a higher false non-match rate).

  • Set it lower, and the reverse.

The US National Institute of Standards and Technology makes the same point in its authentication guideline, SP 800-63B-4 (opens in a new tab): a chosen threshold "usually emphasizes a low FMR to maximize security since false non-matches can often be mitigated by repeating the measurement". For remote identity proofing, SP 800-63A-4 (opens in a new tab) (July 2025) sets performance targets of a false match rate of 1:10,000 or better and a false non-match rate of 1:100 or better. These are US federal benchmarks rather than rules in African markets, but they are a useful statement of where a serious programme aims: impostors should be much rarer than rejected genuine users.

Why is a single pass mark not enough?

Because the scores near the line are the ones where the number tells you least.

A selfie that scores just below the pass mark could be an impostor who looks similar, or a genuine customer photographed in poor light against a ten-year-old record photo. Forcing a yes or no on that score means either declining real customers or waving through lookalikes.

A review band handles the ambiguity honestly: scores comfortably above the pass mark are accepted, scores well below the floor are declined, and the band between them goes to a person who can look at both photos. The face match API in Myaza Trust works this way: its verdict is three-way (match, review or no_match), judged against the organisation's own pass mark and review floor, and deliberately takes no per-request threshold, so every recorded verdict was judged by recorded policy.

Bound the band at the top

One configuration mistake is common enough to call out. A review rule written as "confidence at or above the floor, send to review" catches every score above the floor, including perfect matches. The band must be bounded on both sides:

review when: facialConfidence >= floor AND facialConfidence < passMark

If you build the band as a decision rule, check it against a known strong match before publishing.

How do you set the numbers from your own data?

Vendor defaults are a starting point, not an answer. Your users, their phones, the age of the record photos and your lighting conditions all move the score distributions. Here is a practical calibration method.

  1. Start from the default and collect. Run for a few weeks with a wide review band, so more borderline cases reach people than you will eventually want.

  2. Label the reviewed pairs. Have reviewers record a clear outcome for each: same person, different person, or unusable photo. Unusable photos are not identity evidence either way; keep them out of the calculation.

  3. Plot two histograms. Scores of pairs labelled "same person", and scores of pairs labelled "different person". You now see where your own overlap is.

  4. Set the pass mark from the impostor side. Choose the score above which labelled impostors almost never appear. That is where automatic acceptance is safe.

  5. Set the floor from the genuine side. Choose the score below which labelled genuine pairs almost never appear. Below that, decline or ask for a retake.

  6. Size the band to your review capacity. If the band sends more cases than your team can review within your service time, narrow it, and decide explicitly which error you are accepting more of.

  7. Repeat. Recheck after changing capture flows, adding markets or seeing a new fraud pattern.

An illustrative example: a lender finds, from several hundred labelled pairs, that impostors almost never score above the high 80s and genuine customers almost never score below the mid 60s. It sets the pass mark near the top of the impostor tail and the floor near the bottom of the genuine tail, and sends the gap to review. The numbers here are invented for illustration; yours will come from your own histograms.

Should the threshold differ by photo source or by group?

By photo source, route rather than retune. A selfie compared with a government record photo, an authenticated chip portrait and a photo printed on a document are not equally strong evidence, and a match against a printed photo proves less if the document itself is forged. In Myaza Trust the result reports which photo was used (facialMatchSource: gov_record, chip or document), and a decision rule can send document-photo matches to review regardless of score.

By demographic group, no. NIST SP 800-63A-4 says the system should be configured with a fixed threshold, not one per demographic group, and expects performance for each group to be no more than 25% worse than for the overall population. If your review data shows one group failing more often, the fix is better capture (light, guidance, camera quality) or a better model, not a separate line.

How do you put this into a workflow?

In Myaza Trust Workflows, a workflow can set its own facial pass mark and a review band. The band is a decision rule on verification.facialConfidence, a score from 0 to 100, bounded between the floor and the pass mark, that lands on the review outcome; the decisioning documentation lists the field. A few practical notes:

  • Keep the score out of the applicant's message. Myaza Trust's applicant-facing reason names what did not match and what to do next, and carries no score or threshold. Publishing the bar tells an impostor how close they came.

  • Record every override. When a reviewer approves a near-miss, the result keeps checkStatus: failed and reasonCode: selfie_mismatch beside the approval, which is the record an auditor asks for.

  • Test the band before going live. Sandbox test IDs include a selfie-mismatch scenario, so you can confirm your rule routes a failed match where you expect.

The decision rule

Write your policy down in one line before you touch a slider: "We accept automatically above X because labelled impostors almost never score there; we decline below Y because genuine customers almost never score there; a person reviews everything between, within Z hours." If you cannot fill in X, Y and Z from your own data yet, run with a wide band until you can.

Sources

Charles Archibong

About the author

Charles Archibong

Co-founder

Charles Archibong co-founded Myaza Trust. He writes about identity verification, financial technology, and the practical work of building trusted digital services.

  • Headline "Gesture or flash liveness" beside an illustration of concentric liveness rings around a face outline, on a soft lavender gradient.

    Identity Verification

    Gesture or screen-flash liveness: choosing the right challenge

    Gesture challenges suit broad consumer populations and defeat photos and pre-recorded video. Screen-flash challenges add a random colour sequence that a replayed or injected video cannot contain. Use both where the account is valuable enough to justify the extra seconds.

  • Headline "Not every pass is equal" beside an illustration of three rising tiers, on a vivid purple gradient.

    Product Updates

    Not every pass is equal: using assurance levels in onboarding decisions

    A chip pass, a government database pass and a document-only pass are different strengths of evidence. Decide in advance what each tier may do, route on the tier together with the photo the face was matched against, and let customers move up a tier rather than failing them.

  • Headline "How decision graphs decide" beside an illustration of a decision graph splitting into approve and decline, on a vivid purple gradient.

    Product Updates

    How decision graphs turn checks into approve, review or decline

    When a verification finishes, the workflow's decision graph gathers the results, walks branches over a closed set of fields, and lands on approve, review or decline. The outcome sets the customer's disposition without rewriting what the checks found.

Build your product.We'll handle the rest.

Identity and compliance, end to end, built to global standards, priced for founders.

How to set a face-match threshold and review band · Myaza Trust