Gesture or screen-flash liveness: choosing the right challenge
Gesture challenges suit broad consumer populations and defeat photos and pre-recorded video. Screen-flash challenges add a random colour sequence that a replayed or injected video cannot contain. Use both where the account is valuable enough to justify the extra seconds.

Charles Archibong, Co-founder
· 5 min read

Key takeaways
- Gesture challenges ask the user to nod, turn, blink or smile, in a random order the attacker cannot predict.
- A screen-flash challenge reflects a random colour sequence off the face, which a replayed video cannot contain.
- Flash depends on light: bright surroundings can make the reflection too faint to read.
- Combining both costs a few seconds and closes the gaps each method leaves on its own.
Choose gesture challenges when you serve a broad consumer population on varied phones and your main worry is photos, masks and pre-recorded video. Choose a screen-flash challenge when your main worry is replayed or injected video, because the colour sequence is generated for that session and cannot exist in a recording made earlier. Choose both when the account is valuable enough that a few extra seconds of capture is cheap compared with the loss from one successful impersonation.
The choice is a trade between the attacks you expect and the conditions your users verify in. Neither method is universally better, and each has a failure mode worth designing around.
What does each method ask the user to do?
Gesture challenges ask the user to perform facial movements on request. In the Myaza Trust SDKs the default pool is nod, turn, blink and smile; two are drawn at random for each session, and each has 8 seconds to complete, according to the Flutter SDK liveness configuration. An animated avatar demonstrates each movement, and optional voice guidance reads the instruction aloud (text-to-speech output only, with no microphone access).
Screen-flash challenges ask the user to hold still while the screen shows a short, randomised sequence of colours. A real face in front of the screen reflects those colours; the camera records the reflection. The check asks whether the right colours appeared on the face in the right order.
In both cases the selfie is captured automatically once the challenges pass, rather than by the user pressing a button, and a short video is recorded so the session can be reviewed on the server.
Which attacks does each one resist?
Presentation attacks put something in front of the camera: a printed photo, a phone playing a video, a mask. Injection attacks skip the camera altogether and feed synthetic or recorded video into the app, often with a virtual camera or a tampered device. NIST's 2025 identity proofing guideline, SP 800-63A-4 (opens in a new tab), treats the two separately for this reason, and notes that injection attacks are increasingly paired with generative tools.
Attack | Gestures | Screen flash |
|---|---|---|
Printed photo or cut-out | Strong: a photo cannot blink or turn on request | Strong: a photo reflects differently from skin, and does not move |
Pre-recorded video of the real person | Moderate: the order is random, but a recording with many movements can be spliced | Strong: the recording cannot contain a colour sequence generated after it was made |
Replay of a previous session | Moderate | Strong, for the same reason |
Real-time face swap or animated puppet | Weaker: modern tools can nod and smile on command | Moderate: the swap must also render a plausible reflection of colours it did not know in advance |
Injected video stream | Depends on detecting the injection itself | Strong where the sequence is checked against the recorded video |
Read the table as relative, not absolute. No liveness method on its own stops every attack, and claims that one does should be treated with scepticism.
Where does each one struggle with real users?
The attacks matter, but so does the person on a cracked screen at a bus stop.
Gestures are demanding for some users: people with limited neck mobility, some older users, people who find timed instructions stressful. Poorly lit faces make movements harder to track. The random order also means a user cannot practise.
Screen flash depends on light. The reflection comes from the phone's screen, so it competes with everything else lighting the face. Outdoors in strong sun, or with a low screen brightness, the reflection can be too faint to read with confidence. A well-designed flash check treats that as inconclusive rather than as a failure, because a user standing in daylight is not evidence of fraud. Flashing colours also raise an accessibility question: the W3C's WCAG 2.3.1 (opens in a new tab) guidance says content should not flash more than three times in any one-second period, or should stay below its flash thresholds. Ask any vendor how fast their sequence changes, and tell users what is about to happen before it starts.
A practical comparison, for illustrative user groups:
Situation | Better fit |
|---|---|
Mass-market wallet sign-up, many users outdoors | Gestures |
Users likely to verify indoors, at a desk or at home | Flash works well |
Older users or users with limited movement | Flash, or gestures with generous timing and voice guidance |
High-value accounts, lending, crypto withdrawals | Both |
Re-verification after a suspicious login | Both, because the attacker has already shown intent |
When is it worth using both?
Using both runs the gesture challenges and then the flash sequence. The two close each other's gaps: gestures are indifferent to ambient light, while the flash sequence is hard for a pre-recorded or puppet video to reproduce.
The cost is time and some added drop-off. That trade is usually worth it where one successful impersonation is expensive, such as a loan disbursement, a high-limit wallet or a crypto account that can withdraw. It is rarely worth it for a low-limit account you can restrict later.
In Myaza Trust this is one setting. A workflow chooses gesture challenges, the screen-flash sequence, or both, and every embedded SDK and hosted link inherits it; the livenessMode prop (gestures, flash or both, gestures by default) is the code-side equivalent. The High-assurance biometric template combines gesture and flash liveness with a passport chip read.
What happens after the challenge passes?
The challenge on the phone is only half the check. A liveness verdict produced by the app can itself be forged by an attacker who controls the device, which is why the recorded video matters.
In Myaza Trust Identity Verification, the recorded liveness video is re-analysed on the server to confirm the flash sequence actually appears on the face. A mismatch is recorded as a risk signal for your rules and reviewers, rather than failing the verification by itself. The SDK also pauses the challenge if more than one face appears in frame, and resumes when only one remains.
How to decide
Work through four questions, in order:
What does one successful impersonation cost you? If it is small and recoverable, gestures alone are usually proportionate.
Where do your users verify? Mostly indoors favours flash; mostly outdoors favours gestures.
Who are your users? If many have limited movement, prefer flash or configure gestures with voice guidance.
Is this a first onboarding or a response to suspicion? Re-verification after a risk signal justifies both.
Then test it with real users on real phones before you commit. Measure completion rate and retake rate by method, and review a sample of failures by hand to see whether they failed because of an attack, the light, or the instructions.
Sources

Charles Archibong
Co-founder
Charles Archibong co-founded Myaza Trust. He writes about identity verification, financial technology, and the practical work of building trusted digital services.


