AI & Society

AI Beat Most People at Guessing America’s Unwritten Social Rules

AI robot explaining unwritten social rules to an outspoken American woman demanding to speak to the manager.
AI proved surprisingly good at recognising America’s unwritten social rules, outperforming most people at judging how others expect someone to behave.

Six large language models were better than the average individual at estimating AI social norms across 555 everyday American situations, according to a peer-reviewed study. Yet the advantage changed when people were combined: human errors differed enough to cancel one another out, while the models tended to repeat similar mistakes.

AI social norms were tested against 555 everyday situations

The researchers started with a dataset built by pairing 37 ordinary behaviours with 15 situations. The combinations included questions such as whether it is appropriate to laugh, argue, cry, pray audibly or read in settings ranging from a job interview to a bar or church. Earlier US participants had rated each combination from extremely inappropriate to extremely appropriate, creating a population-average target.

The estimation target comes from earlier US norm-rating research. The original appropriateness-ratings study provides the underlying framework for judging everyday behaviours across contexts. The 2026 Nature paper then asks a different question: how accurately can people and language models estimate those measured population averages?

The new study asked six language models to predict those measured averages. The models were GPT-5.2, Claude Sonnet 4.5, Gemini 3 Flash, Llama 3.1 8B, Llama 3.1 70B and DeepSeek V3. Each model produced repeated estimates, while 320 US participants completed a comparable task on random subsets of 50 scenarios to avoid fatigue.

All six models substantially outperformed the average individual human on the study’s main error measure. Two of the proprietary systems also performed better than the best individual participant. That is an impressive result, but the task needs to be stated precisely: the models were estimating what a US population had collectively rated as appropriate, not demonstrating that they possess morality, empathy or human social experience.

Individual people made worse estimates but more useful mistakes

The researchers found a striking difference in error structure. Individual people often gave more extreme estimates than the measured population norm, making them relatively poor at predicting the average view. But their errors were idiosyncratic. One person’s mistake was not especially likely to be the same as another person’s.

The models behaved differently. Their individual estimates were more accurate, yet their errors were correlated across runs and even across different LLMs. That meant averaging several AI answers produced relatively little improvement. The systems were already good, but they tended to be wrong in related ways.

Human estimates benefited dramatically from aggregation because independent errors cancelled. When the researchers averaged groups of people, performance improved quickly. A group of 15 humans could rival or surpass individual models in important comparisons, and combinations of human collectives with an LLM performed best of all.

That pattern matters well beyond this experiment. LiveAIWire recently covered an analysis showing that keeping human answers independent before combining them with AI improved joint decisions. Both studies point to the value of diversity in error rather than simple agreement.

The hybrid result also complements LiveAIWire’s coverage of human-AI teamwork, where combining people and models did not automatically produce the strongest performer. It differs from behavioural-twin research, which asks whether models can simulate particular individuals rather than estimate a population average.

Agreement between models can be less reassuring than it looks

If several AI systems reach the same answer, it is tempting to treat consensus as evidence of correctness. The new study shows why that can be misleading when the systems share training data, design assumptions or learned cultural patterns. Correlated errors mean several confident answers may represent one family of mistake rather than independent confirmation.

Human disagreement is often treated as inefficiency, but in this setting it became an information source. A group can be noisy at the individual level and strong in aggregate if members make sufficiently independent errors. The principle is familiar from the wisdom of crowds, but the comparison with modern language models makes it newly practical.

The finding does not mean human groups automatically beat AI. Groups can share biases, copy one another or be drawn from the same narrow population. Independence has to be preserved. If everybody sees the model’s answer before forming their own view, the supposedly separate human judgements may become correlated with the machine and with one another.

The models were predicting an American average, not universal morality

The target norms came from US participants. Social appropriateness varies between cultures, generations, communities and situations that the benchmark cannot fully represent. A model that predicts an American average accurately may still misread a specific person, a different country or a context in which local norms are changing.

The authors explicitly frame the task as norm estimation. They are not asking whether the model has its own ethical beliefs. Knowing that people generally judge laughter as acceptable at a party and less acceptable at a funeral is one component of social competence, but real interaction also requires recognising who is present, what has just happened and when a general rule should yield to an exception.

That distinction is important because fluent systems can encourage anthropomorphic interpretations. LiveAIWire has reported that people can be persuaded by a chatbot even when its artificial origin is obvious. Accuracy at predicting a social norm can make an AI appear socially insightful without establishing the richer understanding people may attribute to it.

Retail, care and assistants make this more than a benchmark

The authors highlight settings such as social robots, care environments, customer service and office collaboration. In each case an AI system may need to anticipate whether a behaviour will be seen as rude, intrusive or inappropriate. A system that systematically misreads those boundaries can create harm even when its factual knowledge is strong.

Consider a retail assistant. Separate research covered by LiveAIWire found that machine assistance could feel useful to demographic groups that stereotypes might predict would resist it. The new norms study suggests another design layer: an assistant needs not only product knowledge but a reasonable estimate of how people expect it to behave in a given setting.

Yet the best design may not be to let one model make those decisions alone. If the cost of a social error is high, designers could compare outputs from models with genuinely different training or combine model estimates with independently collected human judgement. The paper’s hybrid results provide a reason to investigate that approach rather than assuming the most accurate single system is automatically the safest.

The hypotheses were not preregistered

The paper is peer reviewed and published in Communications AI & Computing, but the authors state that their hypotheses were not preregistered. The human data collection took place in two waves and produced consistent patterns, which is useful evidence, but preregistration would have provided an additional safeguard against adjusting hypotheses after seeing results.

The models were also a snapshot of systems available in January 2026. Language models change quickly, and newer models could alter the ranking or degree of error correlation. The structural question may be more durable than the leaderboard: do independently developed systems make independent mistakes, or do they converge on the same blind spots?

Another limitation is that the original norm ratings are averages. An average can hide meaningful disagreement within a population. A model can predict the mean accurately while failing to represent minorities whose views differ from it. For applications that affect people, knowing the distribution of opinions may be as important as predicting the centre.

Human diversity can be a technical advantage

The study turns a common criticism of people into a potential strength. Humans are inconsistent. We bring different experiences and make different mistakes. In a task where independent errors cancel, that variation can improve the group result.

AI systems often promise consistency, and consistency is valuable when the rule is known and correct. It becomes a weakness when several systems confidently reproduce the same error. LiveAIWire’s reporting on teachers judging the same harsh recommendation differently when it was labelled as algorithmic shows why an apparently objective machine source can carry extra influence. Shared model blind spots deserve equally deliberate scrutiny.

The practical lesson is not to romanticise disagreement or distrust every AI consensus. It is to ask whether the evidence being combined is genuinely independent. Ten answers that come from nearly the same informational pipeline are not necessarily ten pieces of evidence.

On everyday social norms, individual LLMs were unusually good at estimating what Americans had collectively said. The more interesting result came after that headline: people remained valuable because their mistakes were different. In systems that combine human and machine judgement, diversity of error may be something to design for rather than eliminate.

For organisations, that suggests a concrete evaluation habit: measure not only average accuracy but the correlation between errors. Two systems that make different mistakes can be more valuable together than two individually stronger systems that fail in the same places. Diversity becomes part of reliability engineering.

About the Author

Stuart Kerr is Technology Correspondent at LiveAIWire, covering artificial intelligence, cybersecurity and the social impact of emerging technology. LiveAIWire is an independent, human-led technology publication using AI-assisted research, editorial production and original AI-assisted editorial illustrations under his direction.