AI Ethics

Can AI Tell When You’re Lying? Inside the New Generation of Digital Lie Detectors

AI lie detection illustration of a face being scanned by analytical grid lines
AI lie detection is spreading fast, but the science behind it remains deeply disputed.

AI lie detection is now being marketed for police interviews, border checkpoints, job screening, insurance fraud investigations, and even marriage counselling, and the scientific consensus underneath nearly all of it has not changed in decades: there is no reliably measurable physiological signature of deception. The American Psychological Association’s own review, cited in reporting on the industry, states plainly that “most psychologists agree that there is little evidence that polygraph tests can accurately detect lies.” The AI-powered systems now being sold as the modern replacement for the polygraph inherit that same fundamental problem, dressed in more sophisticated packaging.

This piece looks at what these systems actually measure, what the companies selling them claim, what independent researchers have found when they test those claims, and where the technology is already being used to make decisions about real people’s lives. The pattern that emerges across every sector examined here is remarkably consistent: impressive-sounding accuracy figures from the vendor, and a considerably less flattering picture once anyone outside the company checks the work.

What AI Lie Detection Systems Actually Measure

No commercial AI lie detection system measures deception directly, because deception is not a physical state with a distinct signature. Every product on the market instead measures a proxy: eye movement and pupil dilation, voice pitch and micro-tremors, facial muscle movement, or mouse-cursor behaviour, and infers deception from patterns its developers believe correlate with the cognitive or emotional effort of lying. Converus’s EyeDetect, one of the most widely deployed systems, tracks pupil dilation and reading behaviour on the theory that lying requires more cognitive load than truth-telling, and the company reports accuracy figures around 86 to 88 percent for screening tests.

That figure comes almost entirely from research conducted by scientists affiliated with Converus itself. The Intercept’s investigation into the company’s lobbying efforts found that Charles Honts, a psychology professor who sits on Converus’s own advisory board, declines to use the product himself, telling the outlet the underlying database remains small, drawn mostly from one laboratory, and not yet independently replicated by researchers outside the company.

Why “Cognitive Load” Is Not the Same as “Lying”

The theoretical foundation nearly every AI lie detector relies on is that lying takes more mental effort than telling the truth, and that extra effort produces a measurable physiological trace. Ewout Meijer, a psychology professor at Maastricht University who studies deception detection, explained to The Intercept that measuring emotion and inferring deception from it are two different problems entirely: an elevated stress response could just as easily mean someone is telling the truth but anxious about not being believed.

That distinction is not academic hair-splitting. An innocent person facing a consequential interview, at a border crossing, in a job screening, in a police interrogation, experiences exactly the kind of cognitive load and stress these systems are built to detect, for entirely innocent reasons. Brandeis University social psychologist Leonard Saxe, a long-time polygraph critic, made a similar point in the same reporting: for almost anyone subjected to this kind of test, the experience itself is inherently anxiety-provoking, regardless of whether the person has anything to hide.

The EU’s Own Border Pilot Is the Clearest Cautionary Tale

The most detailed independent investigation of an AI lie detector in actual government use examined iBorderCtrl, an EU-funded pilot that tested a system called Silent Talker on volunteers at borders in Greece, Hungary, and Latvia. MIT Technology Review’s investigation found the system’s published accuracy never exceeded 80 percent even in its own developers’ studies, that it broke down entirely if a subject wore eyeglasses, and that the research supporting it came from training populations as small as 32 people, twice as many men as women, with no Black or Hispanic subjects represented at all.

Judee Burgoon, a researcher at a rival deception-detection company quoted in the same investigation, put the industry’s core weakness plainly: no device that claims to be a straightforward lie detector should be taken at its word. Vera Wilde, an academic and privacy activist who helped organise opposition to the EU pilot, made a related point in the same reporting: an AI system substitutes an algorithm’s own untested correlations for a human examiner’s theory, which makes the underlying assumptions harder to interrogate rather than more rigorous.

The Field Test That Should Worry Every Buyer

Laboratory accuracy claims routinely collapse once a system is tested outside controlled conditions with real stakes attached. The clearest documented example comes from voice stress analysis, a deception-detection technology built on similar theoretical foundations to newer AI systems. A study funded by the National Institute of Justice tested two commercial voice stress analysis programs on more than 300 real arrestees being questioned about recent drug use, with urine tests providing objective ground truth. The programs correctly identified only 15 percent of the arrestees who were actually lying about their drug use, and the researchers’ overall accuracy calculation put both programs at roughly 50 percent, no better than a coin flip.

That gap between laboratory claims and field performance is not incidental. Systems trained and validated in controlled lab settings, where subjects are asked to lie about trivial, low-stakes scenarios, are being deployed in settings with genuine consequences: a real interrogation, a real border crossing, a real insurance claim, where the jeopardy is entirely different from anything the training data captured.

Where This Technology Is Already Making Consequential Decisions

AI-based credibility assessment tools are marketed and deployed well beyond law enforcement. Converus documents its own technology being used for pre-employment screening at police departments and in Latin American operations of major US brands, for infidelity investigations at marriage counselling practices, and its executives have actively lobbied the CIA and Department of Defense to approve the technology for federal security clearance screening, an area currently restricted to traditional polygraph testing under existing federal directives.

This deployment pattern connects directly to a broader problem LiveAIWire has documented in how algorithmic systems get used against people with the least power to contest them. Our reporting on AI at the border making immigration decisions found that courts consistently place the burden of proving an algorithm caused a wrongful outcome on the person affected, using evidence they typically only see after the harm is already done. A traveller flagged by an AI lie detector faces exactly that same evidentiary wall, with even less established legal precedent to draw on than immigration applicants challenging a visa refusal.

The Legal System Has Already Answered This Question, Mostly

US courts have rejected polygraph evidence for decades under the same legal standards that govern novel scientific techniques, and AI lie detection systems face an even steeper climb to admissibility, since they typically cannot explain the specific reasoning behind an individual result the way even a flawed polygraph chart can be reviewed line by line. That legal skepticism has not stopped adoption outside the courtroom. Private employers, insurers, and government agencies operating outside strict evidentiary rules face far fewer barriers to using these tools during screening and investigation, even when a court would never admit the result as evidence.

That gap between courtroom inadmissibility and everyday deployment mirrors a pattern LiveAIWire has traced in AI jury selection tools and AI-assisted police report writing: technology that would face serious scrutiny if formally introduced as evidence continues to shape consequential decisions informally, upstream of any courtroom, where the standards protecting a person from an unreliable tool are considerably weaker.

The Accountability Gap Runs Through All of It

The same lack of independent verification that undermines EyeDetect’s accuracy claims and Silent Talker’s border pilot shows up across nearly every AI system deployed to make judgments about people rather than simply process information. LiveAIWire’s coverage of facial recognition and algorithmic tools used to police the police found the identical structural failure: agencies adopting technology largely on the strength of vendor-supplied accuracy claims, with genuine independent replication either absent or years behind commercial deployment.

AI lie detection is simply the starkest version of that pattern, since the underlying scientific premise, that deception has a detectable signature at all, remains unresolved regardless of how sophisticated the measurement technology gets. A facial recognition system can, in principle, be independently tested against a known correct answer, this is or is not the same face. A lie detector has no equivalent ground truth to test against outside a lab, since the only way to know for certain whether a real-world subject was lying is often the very question the test was brought in to answer.

The Version Aimed Directly at Consumers

The newest wrinkle in this market is not aimed at governments or employers at all. Converus released VerifEye, a mobile app that performs the same ocular deception test using a smartphone’s own camera, letting anyone administer a self-serve credibility check in about ten minutes without any specialised equipment. The company markets it for exactly the kind of personal disputes that used to require a private investigator or a counsellor’s referral: infidelity suspicions, a partner’s honesty about finances, or a teenager’s claims about their whereabouts.

That consumer packaging does nothing to resolve the underlying scientific problem; it simply removes the last layer of professional oversight that a licensed polygraph examiner or a trained counsellor might have provided. A worried partner running a ten-minute phone-camera test on a spouse is applying a technology whose own advisory board members will not personally rely on it, in exactly the kind of emotionally charged, high-stakes context where an innocent person’s stress response is most likely to be misread as guilt.

What This Means for You

If you are ever asked to submit to an AI-based credibility assessment, whether at a border crossing, in a job interview, or during an insurance investigation, the research summarised here supports a specific and practical response: you are entitled to ask what the system actually measures, what independent, non-vendor-funded research validates its accuracy, and whether the result can be appealed or explained if it flags you incorrectly. Asking those questions directly, and asking for the answers in writing, is itself a reasonable and defensible request in almost any screening context.

In jurisdictions and contexts where you can decline the test without automatic negative consequences, the weight of the evidence here suggests that is a defensible choice, not an admission of guilt. The mere existence of a lie detector, AI-powered or otherwise, has never been proof that it works, and six decades of research into the polygraph and its descendants have consistently found that the confidence surrounding these tools outpaces the science supporting them.

About the Author

Stuart Kerr is Technology Correspondent at LiveAIWire, covering artificial intelligence, cybersecurity, and the social impact of emerging technology. He publishes daily at LiveAIWire.com.