Skip to content

Gesture or screen-flash liveness: choosing the right challenge

Gesture challenges suit broad consumer populations and defeat photos and pre-recorded video. Screen-flash challenges add a random colour sequence that a replayed or injected video cannot contain. Use both where the account is valuable enough to justify the extra seconds.

Charles Archibong

, Co-founder

· 5 min read

Headline "Gesture or flash liveness" beside an illustration of concentric liveness rings around a face outline, on a soft lavender gradient.

Key takeaways

  • Gesture challenges ask the user to nod, turn, blink or smile, in a random order the attacker cannot predict.
  • A screen-flash challenge reflects a random colour sequence off the face, which a replayed video cannot contain.
  • Flash depends on light: bright surroundings can make the reflection too faint to read.
  • Combining both costs a few seconds and closes the gaps each method leaves on its own.

Choose gesture challenges when you serve a broad consumer population on varied phones and your main worry is photos, masks and pre-recorded video. Choose a screen-flash challenge when your main worry is replayed or injected video, because the colour sequence is generated for that session and cannot exist in a recording made earlier. Choose both when the account is valuable enough that a few extra seconds of capture is cheap compared with the loss from one successful impersonation.

The choice is a trade between the attacks you expect and the conditions your users verify in. Neither method is universally better, and each has a failure mode worth designing around.

What does each method ask the user to do?

Gesture challenges ask the user to perform facial movements on request. In the Myaza Trust SDKs the default pool is nod, turn, blink and smile; two are drawn at random for each session, and each has 8 seconds to complete, according to the Flutter SDK liveness configuration. An animated avatar demonstrates each movement, and optional voice guidance reads the instruction aloud (text-to-speech output only, with no microphone access).

Screen-flash challenges ask the user to hold still while the screen shows a short, randomised sequence of colours. A real face in front of the screen reflects those colours; the camera records the reflection. The check asks whether the right colours appeared on the face in the right order.

In both cases the selfie is captured automatically once the challenges pass, rather than by the user pressing a button, and a short video is recorded so the session can be reviewed on the server.

Which attacks does each one resist?

Presentation attacks put something in front of the camera: a printed photo, a phone playing a video, a mask. Injection attacks skip the camera altogether and feed synthetic or recorded video into the app, often with a virtual camera or a tampered device. NIST's 2025 identity proofing guideline, SP 800-63A-4 (opens in a new tab), treats the two separately for this reason, and notes that injection attacks are increasingly paired with generative tools.

Attack

Gestures

Screen flash

Printed photo or cut-out

Strong: a photo cannot blink or turn on request

Strong: a photo reflects differently from skin, and does not move

Pre-recorded video of the real person

Moderate: the order is random, but a recording with many movements can be spliced

Strong: the recording cannot contain a colour sequence generated after it was made

Replay of a previous session

Moderate

Strong, for the same reason

Real-time face swap or animated puppet

Weaker: modern tools can nod and smile on command

Moderate: the swap must also render a plausible reflection of colours it did not know in advance

Injected video stream

Depends on detecting the injection itself

Strong where the sequence is checked against the recorded video

Read the table as relative, not absolute. No liveness method on its own stops every attack, and claims that one does should be treated with scepticism.

Where does each one struggle with real users?

The attacks matter, but so does the person on a cracked screen at a bus stop.

Gestures are demanding for some users: people with limited neck mobility, some older users, people who find timed instructions stressful. Poorly lit faces make movements harder to track. The random order also means a user cannot practise.

Screen flash depends on light. The reflection comes from the phone's screen, so it competes with everything else lighting the face. Outdoors in strong sun, or with a low screen brightness, the reflection can be too faint to read with confidence. A well-designed flash check treats that as inconclusive rather than as a failure, because a user standing in daylight is not evidence of fraud. Flashing colours also raise an accessibility question: the W3C's WCAG 2.3.1 (opens in a new tab) guidance says content should not flash more than three times in any one-second period, or should stay below its flash thresholds. Ask any vendor how fast their sequence changes, and tell users what is about to happen before it starts.

A practical comparison, for illustrative user groups:

Situation

Better fit

Mass-market wallet sign-up, many users outdoors

Gestures

Users likely to verify indoors, at a desk or at home

Flash works well

Older users or users with limited movement

Flash, or gestures with generous timing and voice guidance

High-value accounts, lending, crypto withdrawals

Both

Re-verification after a suspicious login

Both, because the attacker has already shown intent

When is it worth using both?

Using both runs the gesture challenges and then the flash sequence. The two close each other's gaps: gestures are indifferent to ambient light, while the flash sequence is hard for a pre-recorded or puppet video to reproduce.

The cost is time and some added drop-off. That trade is usually worth it where one successful impersonation is expensive, such as a loan disbursement, a high-limit wallet or a crypto account that can withdraw. It is rarely worth it for a low-limit account you can restrict later.

In Myaza Trust this is one setting. A workflow chooses gesture challenges, the screen-flash sequence, or both, and every embedded SDK and hosted link inherits it; the livenessMode prop (gestures, flash or both, gestures by default) is the code-side equivalent. The High-assurance biometric template combines gesture and flash liveness with a passport chip read.

What happens after the challenge passes?

The challenge on the phone is only half the check. A liveness verdict produced by the app can itself be forged by an attacker who controls the device, which is why the recorded video matters.

In Myaza Trust Identity Verification, the recorded liveness video is re-analysed on the server to confirm the flash sequence actually appears on the face. A mismatch is recorded as a risk signal for your rules and reviewers, rather than failing the verification by itself. The SDK also pauses the challenge if more than one face appears in frame, and resumes when only one remains.

How to decide

Work through four questions, in order:

  1. What does one successful impersonation cost you? If it is small and recoverable, gestures alone are usually proportionate.

  2. Where do your users verify? Mostly indoors favours flash; mostly outdoors favours gestures.

  3. Who are your users? If many have limited movement, prefer flash or configure gestures with voice guidance.

  4. Is this a first onboarding or a response to suspicion? Re-verification after a risk signal justifies both.

Then test it with real users on real phones before you commit. Measure completion rate and retake rate by method, and review a sample of failures by hand to see whether they failed because of an attack, the light, or the instructions.

Sources

Charles Archibong

About the author

Charles Archibong

Co-founder

Charles Archibong co-founded Myaza Trust. He writes about identity verification, financial technology, and the practical work of building trusted digital services.

  • Headline "Don't trust the phone's verdict" beside an illustration of a shield with a check mark, on a soft lavender gradient.

    Identity Verification

    Why a liveness result from the phone is not enough

    A liveness result produced on the phone is a claim made by a device the attacker may control. Treat it as input, keep the recorded evidence, and re-check that evidence on the server before the result counts.

  • Headline "One person behind many accounts" beside an illustration of a network of connected ownership nodes, on a warm cream gradient.

    Risk & Compliance

    Finding the same person behind many accounts

    Rank shared artefacts by what they prove. A repeated verified ID number means the same person. A repeated face is strong evidence for review. A shared device, phone or address is common in families, so treat it as corroboration and count distinct people before acting.

  • Headline "What a passport chip proves" beside an illustration of a passport booklet with a chip symbol and contactless waves, on a soft lavender gradient.

    Identity Verification

    What reading an ePassport chip proves, and what it can't

    Reading an ePassport chip can prove the data was signed by the issuing state and has not been altered, and with Active Authentication that the chip is not a copy. It cannot prove the person holding the passport is its owner; only a face match against the chip photo does that.

Build your product.We'll handle the rest.

Identity and compliance, end to end, built to global standards, priced for founders.

Gesture vs screen-flash liveness checks compared · Myaza Trust