Teaching AI to Speak Bashkir So the Language Can Live On
Russia's Republic of Bashkortostan is building AI-ready datasets of the Bashkir language and culture. It has become the country's first region to launch a systematic effort that uses artificial intelligence both to preserve cultural heritage and to power everyday digital services.

Artificial intelligence is rapidly becoming a mainstream tool for work and creativity. But what a neural network produces depends on the knowledge it has been trained on. The more accurate the underlying information, the more reliable the output. Researchers argue that training data deserves as much attention as the models themselves.
AI Needs an Ecosystem
The project's primary goal is to create open, legally compliant text, audio, and visual datasets of the Bashkir language and culture for training modern artificial intelligence systems. Bashkortostan is building an entire ecosystem designed to balance the interests of government, society, and the IT industry while supporting AI development.
"We see AI solutions spreading across every sector of the economy and everyday life, with many impressive examples already emerging. But isolated successes are not enough. We need to build a complete infrastructure, an ecosystem and, in effect, a market for IT companies and AI developers. Another priority is to build public trust in artificial intelligence and explain its practical value. To achieve that, we plan to launch a large-scale public awareness campaign. We will place special emphasis on schools, teacher training, and expanding continuing education programs related to digital technologies," said Bashkortostan Head Radiy Khabirov.
Work on open text, audio, and visual datasets of the Bashkir language and culture is already underway. An interagency working group has been formed, bringing together representatives from government agencies, universities, research institutions, libraries, museums, media organizations, publishers, and civil society groups. Particular attention is being paid to legal issues and securing the rights needed to use source materials.

Teaching a Neural Network to Speak Without Mistakes
The first outcome of the initiative is the publication of an open dataset of narrated audiobooks in the Bashkir language. It contains 476 audio excerpts from 14 literary works, including the Bashkir epics Akbuzat and Aldar and Zukhra, as well as works by Miftakhetdin Akmulla. Every audio excerpt is paired with its corresponding text in both Bashkir and Russian.
Over time, the accumulated materials will make it possible to build digital services capable of understanding and generating spoken Bashkir, reproducing knowledge of the people's history and culture, and creating images featuring traditional ornaments, clothing, and other cultural elements without errors or stereotypes. The datasets will also support Bashkir-language translators, voice assistants, and automatic subtitle generation.
Users will benefit from more accurate speech recognition, automatic translation, and faster text-to-speech generation in Bashkir. The datasets will also serve as the foundation for educational services designed for schools and cultural institutions.

The Time for Common Standards
Programmer Aygiz Kunafin believes Bashkortostan is becoming a pioneer by demonstrating how language data for artificial intelligence can be prepared systematically:
"In the past, the development of Bashkir-language resources for AI technologies was driven mostly by individual enthusiasts. We had very limited resources, so volunteers had to gather materials piece by piece, using whatever they could find on their own. Today, for the first time, this work is being carried out centrally. We can build on existing archives, organize them systematically, verify the quality of the materials, and maintain full control over what information goes into the datasets. That significantly accelerates the process while improving quality."

Bashkortostan is setting a model that other regions can follow in preparing cultural and historical materials for AI training. Other Russian regions could use this experience to preserve their native languages and cultural heritage. The growing emphasis on producing reliable training data for neural networks could eventually expand into a nationwide initiative. In that context, creating a unified catalog of language datasets and common standards for providing cultural materials to AI developers would be a logical next step.









































