bg
Science and new technologies
07:33, 17 September 2026
views
13

MSU Scientists Show That Information Estimates Depend on the Mathematical Approach Used

Researchers have proven that two objects that look identical under one way of measuring information can differ substantially when assessed with stricter methods. The work advances algorithmic information theory and refines the concept of informational dependence between objects.

A recent study by researchers at Moscow State University’s research and education school “Brain, Cognitive Systems, Artificial Intelligence” challenges simplified ideas about how information is measured. The work shows that two objects that are identical under one data-evaluation method can differ radically when an alternative mathematical approach is applied. The finding not only sharpens fundamental concepts in computer science but also has direct relevance to modern data-processing technologies.

The key part of the study is a comparison of two ways to measure information complexity: ordinary conditional complexity and total conditional complexity. The methodology is based on Kolmogorov complexity, which defines an object’s complexity by the minimum length of a program capable of reproducing it. For example, a structured sequence such as “101010...” can be described by a short algorithm (“repeat 10”), while a random string of symbols requires a program nearly as long as the object itself. MSU scientists showed that adding further constraints to the program, such as requiring unambiguous decoding, can radically change estimates of how closely related two objects are in informational terms. Thus, two data arrays previously considered equivalent can turn out to be incomparable under a stricter analysis.

Why Does This Matter for Science and Technology?

The findings are directly relevant to artificial intelligence and big-data analysis. Modern AI systems, including language models, work not simply with the amount of information but with its structure and hidden patterns. For example, the effectiveness of neural-network training depends on algorithms’ ability to identify minimal descriptions of data, a task where the accuracy of mathematical models is critical.

The practical significance of the study extends beyond pure mathematics. Refining the concept of “informational dependence” could:

  • improve the accuracy of data-compression algorithms, where an incorrect estimate of complexity can lead to unnecessary computations;
  • improve information-security methods, for example, when analyzing ciphers that are resistant to attacks based on statistical patterns;
  • lay the groundwork for “explainable” AI, making it important to understand exactly which patterns a model extracts from data.

The study is fundamental research, however. Its findings could become a building block for future technologies. As the authors point out, work of this kind is comparable to developing the mathematical tools of differential calculus, without which there would be neither skyscrapers nor space rockets.

From Theory to the Digital Economy

Over the past five years, Russia has made a rapid transition from applied solutions to deeper work on theoretical foundations. In 2022, Sber Cloud launched the ML Space platform, which allows developers to build machine-learning models as a turnkey process. Its success was driven not only by its technical capabilities but also by attention to data structures, including the optimization of intermediate-result storage. Even then, a question emerged: how can information redundancy be formally measured in real-world applications?

The year 2023 was a turning point for Russian AI: Yandex integrated YandexGPT into Alice, while Sber introduced GigaChat. These systems work with semantic relationships in text, requiring new approaches to measuring semantic complexity. For example, the phrase “the cat is sitting on the roof” and its machine translation into another language may have different algorithmic complexity despite having identical meanings. In 2024–2025, the development of multimodal models at VK and other companies raised the question of how to measure relationships between different types of data, including text, images and audio.

The 2026 MSU study is a straightforward continuation of this evolution. If the focus previously shifted toward scaling data, its qualitative analysis is now becoming more important. As the study shows, ignoring differences between methods of measuring information can lead to errors when designing algorithms, such as underestimating the complexity of a classification task.

Strategic Importance for Russia

For Russian science, the work has a dual significance. First, it strengthens the country’s position in mathematical computer science, an area in which Russian schools have traditionally been strong. Under sanctions and import substitution efforts, fundamental research becomes a tool for technological independence. Developing domestic cryptographic standards requires a deep understanding of information complexity, and the MSU findings could be relevant to the NII Kryptotsentr (Kryptotsentr Research Institute) or RAS laboratories.

Second, the work highlights the need to balance applied and theoretical projects. In 2025, the share of government investment in fundamental IT research fell to 12%, compared with 25% in 2020, creating risks for long-term development. The MSU example shows how even “abstract” mathematical models can become the basis for commercial solutions within five to 10 years. Kolmogorov theory, for instance, is already used in data-compression algorithms for the Gonets satellite system.

From Algorithms to Data Ethics

The most immediate prospects for applying the findings involve three areas. The first is the development of “adaptive” algorithms that automatically choose an information-measurement method based on the task. A speech-recognition system, for example, could switch between simplified and stricter complexity models to optimize speed and accuracy.

The second is the development of data-quality metrics. In the era of generative AI, it is critically important to distinguish “artificial” information from original data. The MSU study offers a mathematical tool for assessing the “naturalness” of data: if an object can be described by a short program, it is likely to be synthetic, as in GAN-generated data.

The third area involves ethical issues. As the volume of personal data grows, so does the question of how to measure the risk of a data leak. The MSU research could help create a “risk calculator” showing how much a combination of data collected about a person makes that person vulnerable. Such a tool would be useful not only for complying with laws governing the dissemination of personal data but also for ethical reasons, helping companies avoid collecting information that, even when anonymized, can easily be turned into a dossier on an individual.

A Scientific Foundation for Future Discoveries

The MSU study is part of a global trend. In 2025, a group of scientists from the Massachusetts Institute of Technology and colleagues in Austria and Italy showed that traditional information-entropy metrics are inadequate for quantum computing. In 2026, the European Commission launched the Q-Info program to adapt Kolmogorov theory to quantum algorithms. The Russian study fits into this broader context by proposing solutions for classical problems that are no less complex.

A resilient digital ecosystem cannot be built without advancing the fundamental foundations of computer science. Every new AI startup and every data-analysis platform relies on mathematical foundations that require constant refinement. With its strong theoretical tradition, Russia has an opportunity not only to follow global trends but also to help shape them.

Ordinary conditional complexity answers the question of how short a program can be used to obtain one object from another. Total conditional complexity adds an important constraint: the program must work correctly as a total function. It turns out that this constraint can substantially change the assessment of informational dependence even for objects that naturally arise in Kolmogorov complexity theory
quote
like
heart
fun
wow
sad
angry
Latest news
Important
Recommended
previous
next