By
Stuart Kerr, Technology Correspondent, LiveAIWire
When a teenager in Lagos types a question into ChatGPT and
receives an answer, the exchange is mediated by a system trained
overwhelmingly on English-language text, weighted toward American usage, and
reflecting the cultural assumptions embedded in that corpus. The answer is
fluent. It may be accurate. It is also, in a meaningful sense, a product of a
particular cultural vantage point delivered as if it were a neutral response
to a universal question. This is the unexamined premise at the heart of AI
language’s global reach: that what large language models say, and how they
say it, is culturally neutral when it is not.
Language models are not merely translation tools. They are
generators of meaning, tone, and cultural frame. As they become the primary
interface through which millions of people access information, draft
communications, and receive advice, the question of whose language they speak
— and whose they marginalise — matters more than any technical
benchmark.
The Training Data Imbalance
The most widely used large language models are trained on datasets
in which English constitutes a dominant share of the text. Estimates vary,
but analyses of Common Crawl and similar web-scraped datasets consistently
find that English accounts for forty-five to sixty percent of available
training text, with other European languages making up most of the remainder.
The world’s most widely spoken languages by native speaker count — Mandarin,
Hindi, Spanish, Arabic, Bengali — are substantially underrepresented
relative to the populations who speak them.
The consequence is not only that models perform less well in
non-English languages, though they do. It is that models trained on
English-dominant corpora encode English-language conceptual structures,
cultural references, and rhetorical conventions that shape their outputs even
when generating text in other languages. A model producing Hindi output is
often, in a meaningful sense, translating English-language reasoning into
Hindi syntax rather than thinking in Hindi from the ground
up.
Research from the Meta
AI No Language Left Behind project has documented the performance
gap between high-resource and low-resource languages, and has invested
significantly in training multilingual models with better coverage of
underrepresented languages. The gap has narrowed but has not closed, and the
languages most severely underrepresented — those with fewer than a million
speakers, or those whose written form is less well standardised — remain at
the margins of what current AI can handle effectively.
Dialect, Slang, and the Varieties of Language
Even within English, AI language models perform unequally across
dialects and registers. Standard American English and Received Pronunciation
British English are well represented in training data; African American
Vernacular English, Scots English, Caribbean creoles, and the vast range of
regional varieties that constitute how most English speakers actually speak
and write are less so. Research has documented that language models exhibit
lower accuracy and coherence when processing non-standard dialects, and that
they are more likely to flag dialectal text as containing errors when it
reflects a linguistic convention rather than a mistake.
This matters for automated systems that use language models as
components. A customer service AI that struggles with a Scottish accent, a
content moderation system that flags African American Vernacular English at
higher rates than standard English, or a hiring tool that scores informal
registers lower than formal ones: these are practical manifestations of
linguistic bias that have real consequences for the speakers
affected.
What this means for you: if you communicate in a dialect or
register that differs from the standard varieties on which AI models are
primarily trained, your interactions with AI systems will be systematically
less effective than those of users whose language more closely matches the
training distribution. That is an inequality that most people experiencing it
do not have the technical context to identify as a design problem rather than
a personal failing.
AI and Linguistic Homogenisation
A subtler concern about AI language’s global reach is its
potential influence on language diversity itself. When millions of people
communicate through AI tools — editing their writing with AI assistance,
generating responses with AI help, learning languages through AI tutoring —
the output of those tools shapes what language use looks like at scale. If AI
tools consistently prefer certain constructions, vocabularies, and registers,
they exert a quiet pressure toward homogenisation that is difficult to
measure but potentially significant over time.
Historical analogies are instructive. The printing press, radio,
and television all exerted homogenising pressure on spoken and written
language, contributing to the decline of regional dialect variation in
several countries. AI language tools operate at a scale and intimacy that
exceeds any previous mass communication technology, interacting with
individual writers rather than broadcasting to passive audiences. The
potential for AI-mediated communication to accelerate linguistic convergence
— and with it the marginalisation of languages and dialects that already
struggle for digital presence — is a concern raised by linguists and
cultural heritage organisations.
The UNESCO
Atlas of the World’s Languages in Danger documents ongoing language
loss driven by economic and cultural globalisation. AI may either contribute
to that trend or potentially serve as a tool for language preservation,
depending on whether investment in low-resource language AI models follows
research interest or commercial incentive. So far, it has largely followed
the latter.
Translation, Mistranslation, and Consequential
Errors
AI translation systems have improved dramatically and are now used
in consequential contexts: medical consultations, legal proceedings,
immigration interviews, and emergency services. The improvement in average
translation quality is real and has expanded access to communication for
people who previously had no access to professional interpreting services.
The failure modes, however, are also real and can be consequential in
high-stakes contexts.
AI translation systems fail more frequently on low-resource
language pairs, on domain-specific vocabulary that is underrepresented in
training data, and on pragmatic nuances — politeness levels, implied meaning,
culturally specific references — that do not map cleanly between languages.
A medical AI translation that renders a patient’s description of pain
inaccurately, or that strips the pragmatic signals from a mental health
disclosure, can cause clinical harm that neither the patient nor the
clinician is aware has occurred.
The governance question is whether AI translation is being
deployed in high-stakes contexts with sufficient awareness of its failure
modes, and whether adequate human review is maintained for situations where
an error would be harmful. The evidence suggests inconsistency: some
healthcare systems maintain professional interpreter requirements for
clinical consultations, while others have substituted AI tools without
equivalent safeguards.
The broader challenge of AI
systems that generate meaning in ways their creators did not fully
anticipate is particularly acute in language, where meaning is
irreducibly cultural and where the AI’s confident fluency can conceal the
limitations of its cultural comprehension. A model that sounds authoritative
in a language it understands imperfectly is more dangerous than one that
acknowledges the limits of its competence, because neither the user nor the
system has a reliable mechanism for knowing when the limit has been
reached.
The path to genuinely multilingual AI — systems that think in
multiple languages rather than translating between them, that reflect diverse
cultural knowledge rather than a monoculture dressed in multilingual clothing
— requires both technical investment in training data diversity and a policy
environment that values linguistic equity alongside performance metrics. Both
are available; both are currently underprovided relative to the scale of the
challenge. The
digital divide that already disadvantages many communities has a
linguistic dimension that AI’s current trajectory is deepening rather than
closing.
The cultural encoding in AI language models connects to broader
concerns about AI
systems that reflect and amplify the biases present in their training
data — in language, the bias is not only demographic but
civilisational, shaping which ways of knowing and expressing the world are
treated as default.
About the
Author
Stuart Kerr is a technology correspondent at
LiveAIWire, covering artificial intelligence, emerging technologies, and
their impact on society and industry.