AI Ethics & Privacy

What Gives Away an AI Face? People Can Learn to See the Whole Picture

Guardian-style illustration of a person comparing two near-identical portraits, one human and one subtly AI-generated, for signs beyond the face.
People can become better at identifying AI-generated faces when they examine the wider image, not just facial details.

AI face detection training worked when people stopped hunting for broken teeth, mismatched earrings and other obvious mistakes. In a peer-reviewed PNAS study, all 45 participants in the main experiment improved and mean accuracy nearly doubled after they learned to judge six broad impressions of a face. The method focused on the whole picture: how distinctive, memorable, proportional, symmetrical, attractive and expressive a face appeared.

The result offers a different answer to a familiar problem. Synthetic faces are improving too quickly for lists of rendering glitches to remain dependable. The researchers instead taught participants to notice statistical qualities that image generators tend to produce across the face. High performers reached near-perfect detection in the study, and an online Canadian replication produced a similar improvement.

Why AI Face Detection Based on Glitches Keeps Ageing

Most public advice about synthetic images is built around visible defects. Count the fingers, inspect the hairline, check whether reflections agree and look for jewellery that changes shape. Those clues can catch weak generations, but every improvement in the generator removes another item from the checklist. Fraudsters can also discard obviously flawed images before anyone else sees them.

Faces pose an especially difficult case because generators are trained on huge numbers of portraits and can produce an identity that never existed without needing to alter footage of a real person. A clean profile image can then be attached to an AI-written biography, messages and fabricated documents. LiveAIWire’s reporting on AI voice cloning scams shows how synthetic media becomes dangerous when several plausible signals are combined into one coherent story.

The PNAS researchers argued that global impressions may be more durable than individual artefacts. Generative systems learn a statistical centre from their training data. The faces they produce can therefore look unusually balanced, regular and conventionally appealing while lacking some of the idiosyncrasy and emotional texture found in real portraits.

The Six Whole-Face Qualities Used in Training

Participants were directed towards distinctiveness, memorability, proportionality, symmetry, attractiveness and expressiveness. They were not given a mechanical formula that converted six ratings into a verdict. Instead, the training helped them build perceptual experience of how AI and human faces differ across those dimensions.

According to the Australian National University account, AI faces tended to appear more symmetrical, proportional and attractive. Without training, people often interpreted those same qualities as signs that a face was real. The intervention taught them to reconsider an impression that would otherwise work against accurate detection.

The other side of the pattern was reduced distinctiveness, memorability and expressiveness. A synthetic face can be visually flawless yet feel generic. That feeling is difficult to express as a checklist item, which is precisely why repeated examples and feedback may teach it better than a paragraph of instructions.

How the AI Face Detection Study Tested Improvement

The main study used a pre-training and post-training design with unseen test faces. That final detail is essential. If participants had merely memorised the training images, their score would not demonstrate a transferable perceptual skill. The PNAS article reports that all participants improved and that mean accuracy nearly doubled.

Participants also became better calibrated about their own confidence. Before training, a confident answer could still be wrong. Afterwards, confidence tracked accuracy more appropriately. In security settings, knowing when to pause and seek another check can be as valuable as getting more individual decisions right.

A test-retest control study was used to rule out simple practice as the explanation. Researchers at the University of Victoria then replicated the training online with a new Canadian sample and saw a similar gain. The peer-reviewed publication record identifies the paper, authors, data availability and publication dates, while the anonymised behavioural data is available through OSF.

The improvement is notable because the test faces were not the ones used to teach the pattern. That separates perceptual learning from remembering which image belonged in which category. The online replication also matters for deployment: a method that depends on specialist laboratory equipment would be hard to scale, while a validated lesson that works in a browser could reach reviewers in many organisations.

Confidence calibration provides a second potential benefit. Detection work often sends only uncertain cases to another reviewer, so a person who knows when they may be wrong can support a better escalation system. More correct answers are useful, but fewer confidently wrong answers can be equally important when the image belongs to a real applicant, customer or colleague.

What the Study Does Not Prove

The study used convincing StyleGAN faces. It did not establish that the same training works equally well for every newer diffusion or multimodal generator, for manipulated video, or for a face shown briefly inside a live call. The researchers themselves identify generalisation as the next question.

Durability is also unknown. A person can improve immediately after training and still lose the skill weeks later. The research team is working on shorter training and on whether gains last over time. A scalable online lesson is useful only if its effect survives long enough to matter in real verification work.

The sample of 45 in the main experiment supports a controlled proof of concept, not a universal accuracy figure for the public. High performers reached near-perfect detection, but that should not be turned into a promise that everyone can be trained to near perfection. Mean improvement and individual peak performance are different claims.

Seeing the Whole Picture Is Not the Same as Trusting a Hunch

Global impressions can sound subjective, and they are. The strength of the method is that the impressions were trained against labelled examples and then tested on new faces. A vague feeling that an image looks too perfect is not equivalent to completing the study’s intervention.

There is also a fairness risk. If people apply an unstructured idea of what looks expressive, memorable or normal to unfamiliar groups, bias can be mistaken for detection skill. Training materials must include appropriate diversity and be evaluated across demographic groups. LiveAIWire’s coverage of AI privacy repeatedly shows that biometric systems can appear accurate overall while distributing errors unevenly.

For any consequential decision, human perception should be one signal rather than the whole verdict. Image provenance, account history, reverse-image searches, identity documents, liveness checks and a second reviewer can all add evidence. No individual should be accused of using a synthetic identity solely because their photograph looks unusually symmetrical.

Where Human Training Fits Beside Automated Detection

Automated detectors can process media at scale, but they face their own generalisation problem. A detector trained on one generator may fail on another, and its reasoning can be opaque to the person making the decision. A human trained on global qualities offers an explainable layer, but humans are slower, inconsistent and vulnerable to fatigue.

The useful design is therefore layered. A platform can retain provenance information and run automated screening, while trained people review uncertain or high-risk cases. LiveAIWire’s article on Wi-Fi becoming a biometric illustrates why human oversight matters whenever an indirect signal is used to identify someone. The output must be understood as evidence with conditions, not a magical answer.

Training could be especially valuable for investigators, moderators, recruitment teams and financial staff who already examine suspicious identities. A short, validated module is easier to update than an intuition acquired through inconsistent exposure. It can also teach confidence discipline, helping reviewers recognise when their impression is weak.

What Gives Away an AI Face Today

In this study, the clue was not one impossible detail. It was a cluster of ordinary-looking qualities pushed towards a synthetic average: greater symmetry and proportion, conventional attractiveness, lower distinctiveness, weaker memorability and less expression. The pattern only became useful after people were trained to notice it.

Future generators may learn to vary those qualities, just as they learned to fix fingers and reflections. The research is therefore a method, not a permanent cheat sheet. Detection training has to evolve against current generators and be retested with unfamiliar images.

For now, the work provides rare evidence that people are not condemned to guess forever. When the training moved attention from tiny errors to the face as a whole, every participant improved. The next test is whether that lesson survives new generators, diverse faces and the pressure of real decisions where a convincing identity can cost far more than a wrong answer in a laboratory.

About the Author

Stuart Kerr is Technology Correspondent at LiveAIWire, covering artificial intelligence, cybersecurity and the social impact of emerging technology. LiveAIWire is an independent, human-led technology publication using AI-assisted research, editorial production and original AI-assisted editorial illustrations under his direction.