AI & Health

AI Preserving Lost Languages: Inside the Digital Museum Racing to Save Them

AI preserving lost languages illustration showing a digital museum concept
AI preserving lost languages is becoming possible with remarkably little data

By Stuart Kerr, Technology Correspondent, LiveAIWire

AI preserving lost languages sounds like science fiction. It is already happening in a Dartmouth computer science lab, built from just 35 matching sentences. Researchers there taught a large language model to translate Chinese into Nüshu, a nearly extinct script once used secretly among women in southern China, using almost no data at all. The result offers real hope for thousands of languages now sliding toward silence.

UNESCO estimates one Indigenous language disappears roughly every two weeks. That number is not abstract. Each loss erases stories, medicine, and ways of understanding the world that exist nowhere else. AI preserving lost languages has become one of the few tools moving fast enough to keep pace with that loss, though it comes with real limits worth understanding clearly.

How AI Preserving Lost Languages Actually Works

Nüshu had almost no digital footprint before this project began. Graduate student Ivory Yang, who learned a few words from her grandmother as a child, worked with Dartmouth researchers to build a framework called NüshuRescue. They started with 500 expert-translated sentence pairs. Then they trained GPT-4 Turbo on a much smaller subset, just 35 examples, and the model still learned to translate accurately.

That is the real breakthrough here. Most endangered languages lack the huge datasets modern AI usually needs. NüshuRescue proves a model can still learn useful patterns from a tiny amount of carefully validated data. The same framework, researchers say, could work for other under-documented languages, including Cherokee.

What This Means for You

If your own family speaks a language with few remaining fluent speakers, AI tools now exist that need far less data than most people assume. A handful of recordings, transcribed and translated carefully, may be enough to start building a usable digital resource. That does not replace a linguist. However, it can dramatically lower the cost of starting the work at all.

When the Technology Actively Gets It Wrong

Not every AI language tool performs equally well. Dartmouth researchers tested Google Translate’s language-identification system against Navajo, one of the most widely spoken Indigenous languages in North America. The system misidentified Navajo sentences as entirely different, unrelated languages. As a result, Navajo content often cannot even be reliably detected online, let alone translated.

Yang and her collaborators built a simple, far more accurate identification model specifically for Navajo and related Athabaskan languages. Their work highlights something important: mainstream AI tools are built around major world languages by default. Without deliberate correction, they can actively erase the languages that need the most help.

Turning Hundreds of Recorded Hours Into Usable Text

Speech recognition is solving a different, equally urgent problem. Linguist Sally Nicholas once told Dartmouth’s Rolando Coto Solano that she had recorded hundreds of hours of Cook Islands Māori and joked she would die before finishing the transcription. Coto Solano built an automatic speech-recognition model instead, one that identifies speech patterns from audio and converts them into text automatically.

He has since built similar tools for Bribri and Cabécar, two Indigenous languages spoken in Costa Rica. “Transcription is a very specialized and difficult task, especially in a language that very few people write,” he explains. AI preserving lost languages, in cases like these, is really about removing a bottleneck that has nothing to do with the language’s importance and everything to do with human time.

Why Community Control Has to Come First

Fast progress creates its own risk. Soroush Vosoughi, the Dartmouth professor who supervised the Nüshu project, is candid about this tension. Generative AI lowers barriers dramatically, he says, but these same models can introduce bias from dominant cultures, distorting or oversimplifying the nuance a language actually carries.

Vosoughi’s answer is direct: native speakers and linguists need to stay actively involved throughout, not just at the start. That view lines up closely with what researchers studying language-preservation ethics have concluded elsewhere. Effective AI language work depends on community leadership and genuine consent, not just clever engineering. Consequently, the strongest projects treat the technology as a tool a community controls, rather than a solution imposed on it from outside.

Where Big Tech Fits, and Where It Doesn’t

Large technology companies are involved too, though usually at a different scale. Meta’s No Language Left Behind initiative supports translation across roughly 200 languages, including many with few remaining speakers. Google has separately partnered with academic researchers to build speech-recognition pipelines for languages with very few documented speakers at all.

These corporate efforts matter, especially for scale. However, researchers repeatedly point to a gap between what big platforms can support and what individual communities actually need. As LiveAIWire’s reporting on AI language bias has shown, mainstream AI tools consistently perform worse on low-resource languages than on widely spoken ones, precisely the gap projects like NüshuRescue and Coto Solano’s Cook Islands Māori model are built to close from the ground up.

A Pattern Bigger Than Language Alone

This dynamic, useful global technology built primarily around data-rich populations, shows up well beyond linguistics. As LiveAIWire’s reporting on AI-driven humanitarian forecasting has documented, systems trained mostly on well-resourced regions can systematically underserve exactly the communities with the greatest need. AI preserving lost languages faces the identical challenge: the people with the most at stake often have the least existing digital footprint for a model to learn from.

What Comes Next

None of this technology guarantees a language survives. What it does offer is real, measurable progress against a problem that used to feel purely a matter of time and money. A framework built from 35 sentences. A tool correcting a major platform’s blind spot toward Navajo. Hundreds of recorded hours finally becoming searchable text. Small steps, individually. Together, they represent a genuine shift in what preservation can look like, provided the people whose languages are at stake stay firmly in control of how that technology gets used.

About the Author
Stuart Kerr is Technology Correspondent at LiveAIWire, covering artificial intelligence, emerging technology, and their impact on business, society, and everyday life. LiveAIWire publishes original AI journalism every weekday at liveaiwire.com.