By Stuart Kerr, Technology Correspondent, LiveAIWire
A leading AI model explicitly built and marketed for multilingual coverage performs worse on the very languages it claims to support than models never designed for them at all, according to a 2026 benchmark testing 2,000 responses in Kazakh and Mongolian. Rather than simply giving a lower-quality answer, the model frequently produced a response in the wrong language entirely, a failure mode more basic than the accuracy gap itself. That finding captures something important about AI’s language problem in 2026: the technology has genuinely improved, and it remains structurally uneven in ways that marketing claims routinely paper over.
The core issue has not changed since large language models first emerged: these systems learn from whatever text already exists online in bulk, and the internet is not a fair sample of the world’s roughly 7,000 languages. English, Mandarin and a handful of other data-rich languages dominate training data, while thousands of languages, along with countless regional dialects of “high-resource” languages, remain thin or entirely absent. What has changed is how precisely researchers can now measure the resulting gap, and the picture it reveals is more specific, and more fixable, than a vague sense that AI works better in English.
Table of Contents
How Big the Gap Actually Is, Measured Precisely
The Kazakh and Mongolian benchmark found models scoring between 13.8 and 16.7 percentage points lower when prompted in those languages compared to English, a gap consistent enough across models to rule out random variation. Separate 2026 research from Oracle’s AI team, testing enterprise-grade retrieval-augmented systems, the technique many businesses rely on to ground AI answers in their own documents, found accuracy drops of up to 29 percent in non-English languages compared to English, even when the underlying retrieval system worked correctly. That finding matters because it shows the gap isn’t limited to raw model training. It persists even in systems specifically engineered to reduce hallucination and improve grounding.
What This Means for You
If your business serves customers in multiple languages, or if you personally rely on AI tools in a language other than English, the practical lesson from this research is that accuracy claims validated in English do not transfer automatically. A customer service chatbot that tests well in English can perform meaningfully worse for a French or Arabic-speaking customer using the exact same underlying system, not because the language itself is harder, but because less training and evaluation data exists for it. Businesses deploying multilingual AI should specifically test performance in each target language rather than assuming English-language benchmarks generalise, and individuals relying on AI in a non-English context should apply somewhat more scepticism to answers than they might in English.
Whose English, Exactly
The disparity does not stop at language boundaries. As Brookings Institution researchers have documented, English itself contains over 150 dialects beyond standardised US English, and non-standard varieties including African American Vernacular English and Chicano English are systematically underrepresented in the web-scraped text that trains most models. The consequences are measurable and specific: in one widely cited Stanford study, AI detectors built to flag AI-generated writing misclassified the majority of essays written by non-native English speakers as AI-generated, while performing with perfect accuracy on essays from native English speakers. A tool built to catch dishonesty ended up penalising people for how they were taught to write English, not for anything they had actually done wrong.
Brookings frames this as a digital language divide compounding an existing digital access divide: communities with less reliable broadband access produce a smaller digital footprint, which in turn means less representation in the datasets used to train the next generation of AI tools. The pattern is self-reinforcing unless it is deliberately corrected for, since under-representation in data produces worse tools, which in turn gives affected communities less reason to adopt those tools, further shrinking their digital footprint.
The Safety Dimension Nobody Markets
A less visible but more serious consequence of this gap involves AI safety rather than just accuracy. Researchers have documented that safety filters designed to block harmful requests are frequently trained and tested almost exclusively in English, creating what security researchers call cross-lingual jailbreaks, where simply translating a harmful prompt into an underrepresented language can bypass safety measures that would have caught the same request in English. This is not a hypothetical risk. It has been demonstrated repeatedly across multiple model families, and it means the language gap in AI is not merely an equity or quality issue but an active security vulnerability that scales with how many languages a safety team has resources to test in.
The Gap Is Real, but It Is Not Static
It would be inaccurate to present this purely as a worsening problem. Independent benchmarking from language technology firm RWS, published in April 2026, found that the performance gap between well-supported and underrepresented languages has narrowed significantly across recent model generations, with meaningful gains from major providers including OpenAI and Anthropic. That same research introduced an important caveat: model improvement is not linear or predictable release to release, with the study documenting cases where a newer model version actually underperformed its predecessor on specific non-English content generation tasks, a pattern the researchers termed benchmark drift. The practical implication is that closing the language gap requires continuous testing rather than a one-time fix, since progress in one model generation offers no guarantee the next will maintain it.
What Meaningful Progress Actually Looks Like
The research community’s proposed fixes have shifted from simply calling for more data toward more structural interventions. Grassroots projects like Masakhane, which coordinates linguistic data collection directly with African language speakers rather than scraping from the open web, represent one model for building training data with the affected community’s participation rather than around it. Academic researchers have separately built and released open corpora specifically for underrepresented dialects, including AAVE datasets exceeding 140,000 words, explicitly designed to close representation gaps that commercial data collection has left unaddressed.
This same dynamic, whose data gets collected, and by whom, echoes a pattern LiveAIWire’s reporting on AI-driven refugee forecasting has documented in a very different context: predictive systems built primarily from data-rich populations risk systematically underserving exactly the communities with the least existing digital footprint, whether the underlying task is predicting displacement or simply answering a question in someone’s first language. In both cases, the technology’s blind spots track closely with existing patterns of digital exclusion rather than introducing entirely new ones.
The Case for Getting This Right
None of this argues against multilingual AI deployment, and the researchers behind this work are not making that case either. Multilingual training, done well, has itself been shown to reduce certain forms of bias by exposing models to more culturally diverse data than English-only training would provide. The argument these studies collectively make is narrower and more actionable: claims of multilingual capability need to be tested per language rather than assumed from aggregate benchmarks, safety testing needs equivalent multilingual coverage to safety testing in English, and the communities whose languages are least represented in AI training data are, not coincidentally, often the same communities with the least resources to advocate for that gap being closed. Fixing it will require the industry to treat language coverage as a measured product requirement rather than a marketing claim.
About the Author
Stuart Kerr is Technology Correspondent at LiveAIWire, covering artificial intelligence, emerging technology, and their impact on business, society, and everyday life. LiveAIWire publishes original AI journalism every weekday at liveaiwire.com.