AI Ethics & Privacy

ChatGPT Judged Places Around the World. Richer Ones Kept Winning

liveaiwire ai news insights
liveaiwire ai news insights

ChatGPT location bias becomes surprisingly personal when the question is about your home town. Ask an AI assistant where people are friendliest, safest or most intelligent, and you may receive a confident comparison that sounds like researched local knowledge. Oxford-led researchers found that answers to this kind of question repeatedly favoured richer and better-documented regions, even when the underlying judgement was subjective.

The University of Oxford’s report describes a study of 20.3 million ChatGPT queries comparing countries, cities and neighbourhoods. Published in January 2026, the research found geographical patterns that tracked established inequalities rather than neutral measurements of human ability or merit. The finding matters because increasingly familiar AI responses can make an unfair judgement sound like a straightforward fact.

Why ChatGPT location bias is not just an odd map

A map that appears to rank one country above another is never self-explanatory. It depends on which question was asked, what evidence is available and how the system converts a vague adjective into an answer. The Oxford researchers examined prompts involving attributes such as beauty, intelligence, safety and innovation. Some are value judgements, some require careful statistics, and none should be reduced casually to a universal league table.

The team found that Western Europe, the United States and parts of East Asia were often treated more favourably, while poorer and less thoroughly documented areas frequently ranked lower. The study also examined neighbourhood-level patterns in cities including London, New York and Rio de Janeiro. At that scale, the danger is particularly tangible: one postcode can acquire a flattering technological reputation while another is presented as a problem.

For a person considering a move, choosing a university or trying to understand a destination, an AI answer could influence which options even make it onto the shortlist. The study did not establish that a chatbot has changed a particular household’s decision. It identified systematic patterns in outputs that could shape decisions if readers accept them uncritically.

What the researchers actually measured

In the published research paper, the authors examine geographical judgements made by a language model and organise possible sources of bias into several interacting mechanisms. Some regions are extensively documented online. Others are represented by fragmentary accounts, old news reports or information written from an outsider’s perspective. A model trained to produce plausible text can turn that imbalance of coverage into an imbalance of apparent credibility.

The distinction between a model’s output and an objective measure is essential. If a system regularly labels a wealthy place as innovative, that is a statement about the pattern of answers it produced. It does not prove that its conclusion is true, that all residents share an attribute, or that it has examined the relevant first-hand evidence. A fluent sentence is not a field investigation.

The paper describes biases linked to the availability of information, broad averages, established tropes and proxy signals. In ordinary language, a location may benefit because there is a great deal written about it, because descriptions of its surrounding region are favourable, or because unrelated signals stand in for the thing the user actually wanted to know. Those errors can reinforce one another.

This is also why simply asking for a more detailed justification may not resolve the problem. An AI can produce more elaborate prose drawn from the same unbalanced material. The challenge is not merely a lack of explanation, but whether the underlying comparison deserves to be made in the first place.

There is another subtlety in questions asking a machine to compare ‘the best’ places. Some places have abundant material online in widely used languages, while others are documented through local media, oral knowledge or sources rarely encountered in international datasets. An AI answer can turn unequal visibility into an appearance of unequal quality. The familiar location wins partly because the machine has more readily available material with which to describe it.

It would nevertheless be careless to use this research as evidence that every factual comparison across countries is biased. A question about a published transport statistic is different from asking which population is the most trustworthy. The study’s value is to reveal the problem in the specific comparisons it tested and prompt scrutiny of the assumptions behind similar answers.

Why local stereotypes can travel faster through AI

An old-fashioned online search often presents several sources, allowing a reader to notice that claims about a neighbourhood come from different dates or viewpoints. A conversational answer can flatten that disagreement into one tidy paragraph. That convenience is useful when the underlying evidence is sound, but troubling when the subject involves reputation, prejudice or contested social claims.

There is a further feedback concern. If a generated description of a place is copied into blogs, marketing material or social posts, it can circulate without the original uncertainty. Later readers may encounter an apparently settled judgement without knowing that the phrase began as one model’s broad generalisation. This is a plausible risk identified by the pattern, not proof that a particular location has already suffered such a cycle.

Readers familiar with LiveAIWire’s examination of beauty scoring by AI will recognise a common difficulty: people are tempted to treat a numerical or comparative output as impartial simply because software produced it. Geography adds another layer, because reputation can influence tourism, housing and how strangers interpret communities they have not visited.

How to ask better questions about places

One useful response is to make the question narrower. Instead of asking which city is best, ask for a comparison of a clearly defined characteristic, on a stated date, using publicly accessible evidence. A person choosing somewhere to live might care about train connections, rents, walkability or flood risk. Those topics can be researched individually and checked against local government or other authoritative records.

Ask what the model does not know. Which data are missing? Is the answer using country-wide averages when you asked about a neighbourhood? Are observations from an old survey being treated as current? If a comparison concerns lived experience, whose experience is represented? These follow-up questions do not automatically make an AI answer correct, but they reveal where it may be overreaching.

It can also be better not to rank places at all. A region with a different language, economy or cultural history is not an inferior version of a better-documented one. The practical aim is to gather information relevant to an actual decision, not to translate complex communities into a flattering or insulting label.

LiveAIWire has previously explored public trust in AI decision-making. The Oxford results offer a concrete reason for caution: trust should be earned by the quality and relevance of evidence, not by the confidence or polish of a response.

A traveller might start by asking for three places with similar budgets and then check official visitor information and recent local reporting. Someone comparing areas to live could ask for documented measures such as access to healthcare, public transport and housing costs, specifying the year and geographical boundaries. A person interested in culture can look for local writers and institutions rather than accepting a league table that treats a subjective preference as a universal ranking.

None of these steps requires abandoning AI as a research assistant. They make the assistant responsible for explaining the evidence rather than granting it authority to declare whole communities better or worse. That is a useful distinction whenever a convenient summary risks becoming a stereotype.

The limits of the study and the question for AI developers

The research looked at a specific model and research design. ChatGPT changes over time, and the findings should not be described as a timeless property of every AI system or every possible prompt. Yet the broader concern does not disappear when a model version changes. Training data, evaluation methods and choices about what counts as a satisfactory answer all influence how places are described.

Independent audits matter precisely because different approaches can expose different failures. LiveAIWire has covered why AI bias audits may disagree, a warning against treating any single measurement as the final word. A useful evaluation would ask not only whether a model sounds fair, but whether its conclusions can be justified across different regions and groups.

The most constructive lesson is simple. When an AI describes an unfamiliar place, it is offering a generated interpretation of available information. It has not visited the streets, listened to residents or verified every implication. A human reader who remembers that distinction is far less likely to turn an automated stereotype into a personal judgement.

About the Author

Stuart Kerr is Technology Correspondent at LiveAIWire, covering artificial intelligence, cybersecurity and the social impact of emerging technology. LiveAIWire is an independent, human-led technology publication using AI-assisted research, editorial production and original AI-assisted editorial illustrations under his direction.