Deepfakes and injected camera feeds: how identity checks are attacked
Attackers either hold fake media up to a real camera (a presentation attack) or feed it straight into the data stream (an injection attack). Each needs its own defence, and the strongest checks add evidence a camera cannot supply.

Charles Archibong, Co-founder
· 6 min read

Key takeaways
- Presentation attacks fool the camera; injection attacks skip it. They need different defences.
- NIST SP 800-63A-4 (July 2025) requires controls that detect virtual cameras, emulators and jailbroken devices.
- Liveness challenges raise the cost of an attack but are not proof on their own.
- Evidence the camera cannot fake, such as a signed passport chip or a government record photo, is the strongest counterweight.
Attackers beat remote identity checks in two ways. In a presentation attack, they hold something fake in front of a real camera: a printed photo, a replayed video on a second screen, a mask, or a deepfake playing on a phone. In an injection attack, they skip the camera entirely and feed forged images or video straight into the app or the network request, often through a virtual camera, an emulator or a modified app.
The defence follows from the split. Presentation attack detection (PAD) judges what the camera sees. Injection attack detection asks whether the camera saw anything at all. Neither is complete, so the most reliable programmes also rely on evidence a camera cannot supply: a cryptographically signed passport chip, or a comparison against a photo held by the issuing government.
What makes a deepfake dangerous to a verification flow?
A deepfake is a way of producing convincing fake media. On its own it has to reach the verifier somehow. NIST's July 2025 identity guidelines describe the combination directly: attackers "pair digital injection attacks with increasingly effective and available generative AI tools" to create images or videos "to defeat automated document validation processes, biometric operations, and visual comparisons" (NIST SP 800-63A-4, section 3.14 (opens in a new tab)).
NIST's glossary defines an injection attack as supplying "untrusted biometric information or media into a program or process", for example a falsified image of identity evidence, a forged video of a user, or a morphed image.
So there are two targets to defend: the selfie and liveness step, and the document capture step. A deepfaked face paired with a doctored document photo, both injected, attacks both at once.
How do the two attack types compare?
Presentation attack | Injection attack | |
|---|---|---|
Where the fake enters | In front of a real camera | Between the capture point and the server |
Typical instruments | Printed photo, screen replay, mask, deepfake on a second device | Virtual camera software, emulator, rooted or jailbroken phone, altered app, intercepted request |
What the camera records | The fake object | Nothing real; the stream is replaced |
Standards that test defences | ISO/IEC 30107-3:2023 (opens in a new tab) (PAD testing and reporting) | CEN/TS 18099:2024 (biometric data injection attack detection) |
First line of defence | Liveness challenges and passive PAD | Sensor and device integrity checks, media forensics |
The two standards are deliberately separate. ISO/IEC 30107-3:2023 sets out how to assess PAD mechanisms and report the results, and puts "overall system-level security or vulnerability assessment" out of its scope. CEN/TS 18099, published in the UK as PD CEN/TS 18099:2024 on 30 November 2024 (opens in a new tab), covers injection attack instruments and how to build and test injection attack detection, and places presentation attack testing out of its scope. A vendor citing one standard has told you nothing about the other attack.
What do the NIST guidelines require?
For remote identity proofing, NIST SP 800-63A-4 (opens in a new tab) (published July 2025) says credential service providers "SHALL implement technical controls to increase confidence that digital media is being produced by a genuine sensor", for example to "detect the presence of a virtual camera, device emulator, or a jailbroken device". They SHALL analyse all submitted media for signs of modification or forgery, and SHOULD analyse media for signatures of generative AI tools. NIST also notes that live document capture and PAD "do provide some protection" against injection but "are not sufficient to address all possible cases".
For biometric authentication, NIST SP 800-63B-4 (opens in a new tab) (also July 2025) says a biometric system "SHALL implement PAD for facial recognition" and SHOULD demonstrate an impostor attack presentation accept rate below 0.07, tested under Clause 13 of ISO/IEC 30107-3. It also asks verifiers to check the sensor and its endpoint, noting this "increases the likelihood of detecting injection attacks due to compromised endpoints, sensor emulators, and similar threats".
These are US federal guidelines, not law everywhere, but they are the clearest public description of the controls a remote verification flow should have. Requirements in your market may differ, and this article is general information, not legal advice.
Which defences make each attack harder?
No single control stops a determined attacker. A layered design raises the cost at each point.
Against presentation attacks
Active liveness with randomised challenges. Asking for a nod, a turn, a blink or a smile, in an order the attacker cannot predict, makes a pre-recorded clip much less useful.
A light challenge from the screen. Flashing a random colour sequence and checking that the colours reflect off the face, in order, tests whether a face was actually in front of the device while they played.
Automatic capture. If the user cannot choose the moment the selfie is taken, they cannot hold up a still image at the right instant.
Against injection attacks
Device and sensor integrity. Look for virtual camera names, emulator fingerprints and rooted or jailbroken devices, as NIST asks.
Live capture only for documents. Removing the "upload from gallery" option for documents forces a camera scan, which a stored doctored image cannot satisfy without also being injected.
Server-side analysis of the recording. Checks run on the server cannot be switched off by a modified app.
Against both
Evidence the camera cannot fabricate. An e-passport's chip carries data and a portrait signed by the issuing state. A chip read and verified server-side, with Active Authentication where the chip supports it, shows the document is genuine and not a copy. Comparing the selfie against the photo on the government record means the attacker must match a face they do not control.
How Myaza Trust applies these layers
Identity Verification uses Presence Intelligence (selfie and active liveness) with three modes a workflow can choose: randomised gesture challenges (nod, turn, blink, smile), a screen-flash colour sequence, or both. The selfie is captured automatically once the challenges pass, a challenge pauses if more than one face is in frame, and a short liveness video is recorded for server-side review (SDK documentation). The recorded video is re-analysed on the server for the flash sequence; a mismatch is recorded as a risk signal rather than failing the check by itself, so you decide how much weight it carries.
Decision rules can read verification.cameraSuspect, which is true when capture-integrity signals flag injection, and device.emulator (decisioning fields). A sensible starting rule sends either to manual review. Document capture can be set to camera scan only by turning off document upload.
In supported markets the selfie is compared against the government record photo. On mobile, the React Native and Flutter SDKs can read an e-passport or chip ID card; the chip is checked server-side and counts as authentic only when its signer chains to a trusted issuing-state certificate (NFC chip documentation). Browsers cannot read chips, so web flows rely on the other layers.
None of this makes a verification immune to deepfakes. It makes a successful attack need several things to go right at once.
A checklist for your own flow
Can a user upload a document photo instead of scanning it? If so, decide whether that route should carry a lower assurance level or go to review.
Is the selfie captured automatically, after challenges the user cannot predict?
Do you record and keep the liveness session so it can be re-examined?
Do your rules act on device integrity signals (emulator, virtual camera) rather than only logging them?
For your highest-risk products, do you require evidence outside the camera: a chip read or a government record match?
When you evaluate a vendor, ask which attack each test covered: ISO/IEC 30107-3 for presentation attacks, CEN/TS 18099 for injection.
Sources

Charles Archibong
Co-founder
Charles Archibong co-founded Myaza Trust. He writes about identity verification, financial technology, and the practical work of building trusted digital services.


