AI & Society

Four Major AI Models Gave Faces Higher Beauty Scores Than People

Three AI judges score contestants in a talent-show setting, illustrating how AI models rated the attractiveness of human faces.
Major AI models consistently gave human faces higher beauty ratings than people did, revealing a measurable difference between machine and human judgements of attractiveness.

AI beauty scores were systematically higher than human ratings in a pre-registered exploratory study comparing 2,513 people with Claude, Gemini, GPT and Grok. The four commercial models also used a narrower range of scores, even though their overall ranking of faces correlated strongly with human judgements. The preprint, submitted on 2 September 2026, concludes that the models could broadly track which faces people rated more attractive without reproducing human ratings in absolute terms.

That combination is more interesting than a simple claim that “AI thinks everyone is beautiful”. The models were not random. They broadly agreed with the ordering found in human judgements, but their scoring behaviour was shifted upwards and compressed. A system can therefore look aligned at the ranking level while still giving a meaningfully different answer to an individual user.

What Higher AI Beauty Scores Actually Mean

Imagine two judges who broadly agree that one set of portraits is more attractive than another, but one judge rarely gives a low mark. Their rankings can correlate even when the numbers mean different things. That is the pattern the researchers found across the tested models.

Claude, Gemini, GPT and Grok all rated faces more favourably than the human comparison group and used less of the available rating range. The paper says the models strongly agreed with one another, with Grok the main exception and also the model showing the lowest agreement with human ratings. The finding suggests a shared tendency among several commercial systems rather than an isolated quirk in one product.

This matters because beauty ratings are increasingly offered informally by general-purpose chatbots and more directly by services built around appearance, styling or aesthetics. A user may interpret a numerical score as if it were a measurement. The study shows why that interpretation is too strong. The output reflects the behaviour of a model, not an objective beauty scale.

AI Beauty Scores Can Agree on Rank but Disagree on Meaning

The difference between relative and absolute judgement is crucial. If a model tends to rank the same faces near the top and bottom as people do, it may appear highly human-like in aggregate. But if the model shifts almost everyone upwards, a score of eight from the model may not correspond to a score of eight from the human group.

That makes direct comparisons risky. A user who asks several models to “rate my face out of ten” may see similar flattering numbers and conclude that the systems have independently confirmed the same judgement. In reality, the systems may share training patterns, safety preferences or conversational tendencies that push them towards a narrower and more positive response style.

LiveAIWire has previously looked at how people can learn to detect AI-generated faces. That research focused on recognising whether an image was synthetic. The beauty study asks a different question: when models evaluate a real human face, how closely does their subjective judgement resemble human ratings?

What This Means for People Using AI to Judge Appearance

The safest interpretation is that an AI beauty score is feedback from a conversational system, not a measurement of personal worth, health or social outcome. The model may be influenced by instructions, safety tuning, training data and the way the prompt is framed. Even a stable answer across models does not turn a subjective judgement into an objective fact.

For consumer products, that means interfaces should be careful with precision. A number such as 8.3 can look scientific even when the underlying construct is culturally variable and the model’s scale does not match human use of the same numbers. A qualitative description may sometimes be more honest than a pseudo-exact rating.

For businesses using facial analysis, the stakes can be higher. Appearance-based scoring can affect marketing, casting, content ranking or recommendation systems. A model that systematically compresses ratings may reduce visible variation, but that does not automatically remove bias. It can simply reshape the distribution in a different way.

The Models Did Not Use Every Human Cue in the Same Way

The researchers examined whether age, ethnicity and gender predicted attractiveness ratings. Face age was the only predictor shared consistently between humans and the multimodal models, while ethnicity and gender showed inconsistent patterns across systems. That is another warning against assuming that a high overall correlation means the model reasons about faces in the same way people do.

Different cues can produce similar rankings. A model may arrive at a broadly human-like order through statistical patterns that do not match the factors people consciously or unconsciously use. That is important when organisations try to explain or audit a model’s judgement.

LiveAIWire’s report on AI political profiling from facial photographs raised a related issue. A model can find correlations in faces without those correlations becoming a reliable basis for consequential decisions. Attractiveness is different from political profiling, but both show why a confident facial judgement can carry more authority than the evidence warrants.

Flattering Outputs May Be a Product Behaviour, Not a Discovery

General-purpose chatbots are designed to interact helpfully and avoid unnecessarily hostile responses. It is plausible that some of the upward shift reflects conversational or moderation behaviour, although this study does not isolate the cause. The researchers report the scoring pattern; they do not prove which training or safety mechanism produced it.

That uncertainty matters because users often interpret consistency as objectivity. If several models all give relatively generous ratings, the shared tendency could arise from overlapping data, similar commercial incentives or similar alignment choices. Independent-looking outputs do not necessarily represent independent evidence.

This is part of a wider problem in AI evaluation. A model can imitate the surface pattern of human judgement while behaving differently at the edges, in calibration or across demographic groups. Those differences become important when a score is used for comparison rather than casual conversation.

The Study Has Important Limits

The paper is a preprint rather than a peer-reviewed final publication, and its findings apply to the tested model versions and study materials. Commercial models change frequently. A provider can alter behaviour through model updates, system prompts or safety policies without changing the product name visible to the user.

The human comparison group also has its own limits. Attractiveness is shaped by culture, age, social context and personal preference. The study cannot establish a universal human standard against which machines can be permanently calibrated. It can show how these model outputs differed from this large human sample under the study design.

That is especially important when discussing demographic effects. A result found in a particular face set and participant pool should not be expanded into claims about all populations. The paper itself presents an exploratory comparison, not a licence to use attractiveness models for high-stakes assessment.

Beauty Is a Poor Place to Hide Calibration Problems

The title of the paper plays on the phrase “beauty is in the eye of the beholder”. The deeper technical lesson is about calibration. A model can correlate with people and still use the scale differently. The same issue appears in risk scores, confidence estimates and other AI judgements where users may care about the number rather than only the ordering.

LiveAIWire’s wider coverage of AI surveillance and automated profiling has repeatedly shown why quantified outputs deserve scrutiny. Numbers create an impression of precision. Whether that precision is meaningful depends on what the model was trained to do, how the scale behaves and whether the result has been validated for the intended use.

For a casual chatbot conversation, a flattering beauty score may be harmless or even welcome. For anyone tempted to treat that score as evidence, the new study offers a useful correction. The models may broadly recognise the same ordering as people while still speaking a different numerical language.

Models Can Change Without the Question Changing

There is another practical reason not to treat an AI beauty score as stable evidence. Commercial models are continuously updated. The same prompt, image and product name can produce a different rating after a model or system-policy change. For subjective judgements, that means the apparent “measurement” can drift even though nothing about the person has changed.

A useful consumer interface would make that uncertainty visible instead of presenting a precise score as permanent. The study captures a snapshot of four systems at one point in 2026. Its strongest value is showing that the models shared a measurable tendency to rate more generously, not establishing an eternal property of those brands.

About the Author

Stuart Kerr is Technology Correspondent at LiveAIWire, covering artificial intelligence, cybersecurity and the social impact of emerging technology. LiveAIWire is an independent, human-led technology publication using AI-assisted research, editorial production and original AI-assisted editorial illustrations under his direction.