bg
Science and new technologies
07:46, 19 September 2026
views
10

Russian TimeAdapter Is Transforming Generative Video and Bringing Physical AI Closer

The Kandinsky team at Sberbank has developed TimeAdapter, a module for generative models that enables more precise control over the speed of events in video. The neural network can use it to account not only for what is happening in a scene but also for its dynamics, making movement more natural.

Today’s neural networks can create stunning worlds, but they often struggle with basic physics. In generated footage, movement can look unnatural. The Kandinsky team at Sberbank has tackled this fundamental problem with TimeAdapter, a compact add-on that teaches AI to do more than generate pixels – it helps the system understand time and the dynamics of events.

How TimeAdapter Works

The main innovation is the separation of two concepts: smoothness and event speed. TimeAdapter accounts for less than 1% of the size of the main neural network, yet it handles two critical tasks. The first block controls the frame rate (fps – frames per second), allowing users to set anything from a cinematic 15 fps to an ultra-smooth 60 fps. The second block controls the actual pace of events within the frame.

To train the algorithm, researchers used about 40,000 videos showing a wide range of motion and dynamics, allowing the neural network to learn how time unfolds in the real world. TimeAdapter operates in two modes. In the first, it plugs into an existing model as a workaround that improves smoothness. In the second, the main neural network is fine-tuned along with the adapter, allowing it to start learning cause-and-effect relationships. The results, presented at the prestigious ICML conference, are striking: motion naturalness increased by 29%, visual quality by 8%, and prompt-following accuracy by 19%. The code is released under the MIT License, allowing it to be used freely, including for commercial purposes.

From Ads to Robots: The Era of Physical AI

Creative professionals stand to benefit directly from the technology. Marketers and bloggers will no longer have to spend hours editing footage by hand or making dozens of attempts to sync visuals with a music track. Sberbank itself already uses AI to create 60% of its advertising. But the larger goal goes much further.

We are approaching the era of physical AI – systems that understand the laws of the real world and can interact with it. This is where World Models come in. They are more than video generators: they are virtual simulators where robots and autonomous vehicles can train by encountering rare or dangerous situations without risking anything in reality. TimeAdapter provides an important bridge from “pretty pictures” to physics simulation. It is no coincidence that in August 2026, Sberbank established the Kandinsky World Model direction, releasing specialized models for autonomous transportation and robotics.

From Pixels to Simulation

To understand the scale of this technological shift, it is enough to look back at the past few years. The 2022–2023 period was defined by text-to-image systems and the first video generators, including Kandinsky 2.0 and Runway Gen-2. The main challenge was simply keeping frames from turning into visual mush. In 2024–2025, the race shifted to longer videos, higher quality and sound, with systems such as OpenAI Sora and Google Veo 3. Resolution improved and synchronized audio arrived, but complex physics and process logic remained a weak point.

The focus in 2026 is physical AI. The industry’s attention has shifted from visual appeal to the credibility of the underlying logic. TimeAdapter and Kandinsky WM 1.0 reinforce that trend, moving video generation into the realm of tools for modeling reality.

An Under-the-Radar Export Opportunity and Open Source

TimeAdapter is not yet a ready-to-export product, but it has a powerful advantage – open source. As countries in the Global South increasingly seek sovereign AI technologies from Russia, the open TimeAdapter codebase could become an entry point for Russian developers into the global market.

This is more than a patch for improving video. It is a bid for leadership in creating virtual sandboxes where future machines can be trained. Russian science has shown that it can set trends not only in language models but also in the most challenging segment of generative AI – where code meets the laws of physics.

Previously, the speed of events in a frame was controlled only through words in the prompt, while influencing the actual dynamics of the generated video was quite difficult. With TimeAdapter, the model understands how processes unfold over time and can naturally set the desired rhythm in any scene instead of memorizing ready-made patterns. This is a key condition for training physical AI and future World Models: a system must accurately understand the duration and dynamics of every action. The new method makes models better tools for following human intent – whether that intent comes from a director, an engineer, a robotics developer or any creative user
quote
like
heart
fun
wow
sad
angry
Latest news
Important
Recommended
previous
next