AI Safety & Security

To Catch a Deepfake Caller, Make Them Do Something Live

Human on a video call discovers an AI deepfake hiding behind a mask of an elderly woman.
Deepfake video calls can be harder to sustain when the person on screen is asked to perform an unexpected action live.

A deepfake caller may look and sound convincing until the person on the other end asks for an unexpected action in real time. A new preprint tested that idea by making suspected synthetic callers complete simple live challenges, such as changing their expression, moving their head or producing an unplanned sound. The researchers found that current real-time face and voice deepfake systems often struggled to satisfy the challenge while preserving a believable identity and responding quickly enough to look natural. The Deep-Fake CAPTCHA study was submitted on 10 September 2026.

The attraction is easy to understand. Most advice about deepfakes asks the target to inspect the picture or voice for flaws. That approach becomes weaker as generators improve. A live challenge changes the problem. Instead of asking whether a face looks slightly wrong, it asks whether the system can react correctly to a new instruction that an attacker could not fully prepare in advance.

Why a Deepfake Caller Struggles With an Unexpected Request

Real-time deepfakes have to do several things at once. They must capture the attacker’s performance, transform it into the target identity, maintain synchronisation with the conversation and return the result with low enough latency that the call still feels normal. The researchers evaluated four real-time face-replacement systems with 38 volunteers and five voice-cloning systems with 41 volunteers. Their defence judged whether a response stayed realistic, preserved identity, completed the requested task and arrived within an acceptable time. The paper’s evaluation reports that challenge-response testing improved detection across the systems examined.

This is different from passive detection. A passive detector receives whatever the attacker chooses to send. An active detector creates a new condition and waits to see whether the synthetic system can adapt. That idea resembles a CAPTCHA on a website, except the challenge is aimed at a live audio or video stream rather than a visitor clicking pictures or typing characters.

The principle also fits a broader change in impersonation fraud. LiveAIWire has previously reported how AI voice-cloning scams can imitate a relative in distress. In those situations, emotional pressure is part of the attack. A demand for money can arrive before the target has time to think about whether the voice itself is genuine. An active challenge gives the recipient something concrete to do instead of relying on intuition.

What This Means for You During a Suspicious Call

The study does not prove that one particular gesture will expose every fake. Its more useful message is behavioural: make the caller react to something they could not confidently predict. A video caller could be asked to turn their head in a particular direction, briefly cover part of their face or perform another harmless movement. A voice caller could be asked for an unusual sound or phrase. The important feature is unpredictability, not a fixed public checklist that criminals can train around.

For high-stakes calls, the live challenge should sit alongside an independent verification step. If someone claiming to be a relative, bank employee or colleague suddenly asks for money, credentials or sensitive information, ending the call and contacting them through a number already known to you remains stronger than trying to become a forensic deepfake analyst. LiveAIWire’s earlier coverage of AI-assisted phishing attacks makes the same point in another channel: attackers benefit when urgency prevents verification.

A challenge is especially useful when hanging up is socially difficult. Video meetings, recruitment calls and family conversations can all create pressure to continue. Asking for a spontaneous action introduces a small amount of friction at exactly the moment when an impersonator wants the interaction to remain smooth and unquestioned.

The Defence Is Promising, but Attackers Can Adapt

The research is a preprint, not a peer-reviewed field trial. It tested current systems under controlled conditions, and the participant groups were relatively small. The authors also focused on real-time generated calls rather than every form of manipulated media. A prerecorded clip, for example, creates a different detection problem because there may be no live system to challenge.

Attackers can also learn. If one challenge becomes widely recommended, future deepfake systems can be optimised for it. The strongest interpretation is therefore not that coughing, clapping or turning is a permanent magic test. It is that unpredictable live interaction creates a moving target that synthetic systems must satisfy under time pressure.

That is an important distinction because deepfake detection has often become a race to recognise visible artefacts. LiveAIWire’s coverage of AI political deepfakes shows why appearance alone is a poor long-term foundation for trust. As synthetic images and voices improve, security measures that depend on a specific visual mistake can age quickly.

Active Verification May Be More Durable Than Spotting Glitches

The study points towards a broader security principle. Authentication works better when the person being verified has to respond to a fresh challenge than when an observer simply studies a presented identity. Banks already use one-time codes, websites use challenge-response mechanisms and organisations use call-back procedures for unusual requests. Applying the same logic to a suspicious synthetic caller is a natural extension.

It is too early to know whether this approach will become a built-in feature of calling platforms or remain a human tactic. A platform could eventually generate random tasks automatically and measure response quality, timing and identity consistency. That would need careful testing because false alarms could interrupt legitimate calls and accessibility requirements would rule out some challenges for some users.

For now, the useful lesson is modest and practical. If a call matters enough that identity must be trusted, do not only stare harder at the face or listen harder to the voice. Make the person react to something new, then verify through a separate channel when the stakes are high. The deepfake may be prepared for the conversation. It is less likely to be prepared for your next move.

About the Author

Stuart Kerr is Technology Correspondent at LiveAIWire, covering artificial intelligence, cybersecurity and the social impact of emerging technology. LiveAIWire is an independent, human-led technology publication using AI-assisted research, editorial production and original AI-assisted editorial illustrations under his direction.