Russian TimeAdapter Is Transforming Generative Video and Bringing Physical AI Closer
The Kandinsky team at Sberbank has developed TimeAdapter, a module for generative models that enables more precise control over the speed of events in video. The neural network can use it to account not only for what is happening in a scene but also for its dynamics, making movement more natural.

Today’s neural networks can create stunning worlds, but they often struggle with basic physics. In generated footage, movement can look unnatural. The Kandinsky team at Sberbank has tackled this fundamental problem with TimeAdapter, a compact add-on that teaches AI to do more than generate pixels – it helps the system understand time and the dynamics of events.
How TimeAdapter Works
The main innovation is the separation of two concepts: smoothness and event speed. TimeAdapter accounts for less than 1% of the size of the main neural network, yet it handles two critical tasks. The first block controls the frame rate (fps – frames per second), allowing users to set anything from a cinematic 15 fps to an ultra-smooth 60 fps. The second block controls the actual pace of events within the frame.
To train the algorithm, researchers used about 40,000 videos showing a wide range of motion and dynamics, allowing the neural network to learn how time unfolds in the real world. TimeAdapter operates in two modes. In the first, it plugs into an existing model as a workaround that improves smoothness. In the second, the main neural network is fine-tuned along with the adapter, allowing it to start learning cause-and-effect relationships. The results, presented at the prestigious ICML conference, are striking: motion naturalness increased by 29%, visual quality by 8%, and prompt-following accuracy by 19%. The code is released under the MIT License, allowing it to be used freely, including for commercial purposes.

From Ads to Robots: The Era of Physical AI
Creative professionals stand to benefit directly from the technology. Marketers and bloggers will no longer have to spend hours editing footage by hand or making dozens of attempts to sync visuals with a music track. Sberbank itself already uses AI to create 60% of its advertising. But the larger goal goes much further.
We are approaching the era of physical AI – systems that understand the laws of the real world and can interact with it. This is where World Models come in. They are more than video generators: they are virtual simulators where robots and autonomous vehicles can train by encountering rare or dangerous situations without risking anything in reality. TimeAdapter provides an important bridge from “pretty pictures” to physics simulation. It is no coincidence that in August 2026, Sberbank established the Kandinsky World Model direction, releasing specialized models for autonomous transportation and robotics.

From Pixels to Simulation
To understand the scale of this technological shift, it is enough to look back at the past few years. The 2022–2023 period was defined by text-to-image systems and the first video generators, including Kandinsky 2.0 and Runway Gen-2. The main challenge was simply keeping frames from turning into visual mush. In 2024–2025, the race shifted to longer videos, higher quality and sound, with systems such as OpenAI Sora and Google Veo 3. Resolution improved and synchronized audio arrived, but complex physics and process logic remained a weak point.
The focus in 2026 is physical AI. The industry’s attention has shifted from visual appeal to the credibility of the underlying logic. TimeAdapter and Kandinsky WM 1.0 reinforce that trend, moving video generation into the realm of tools for modeling reality.

An Under-the-Radar Export Opportunity and Open Source
TimeAdapter is not yet a ready-to-export product, but it has a powerful advantage – open source. As countries in the Global South increasingly seek sovereign AI technologies from Russia, the open TimeAdapter codebase could become an entry point for Russian developers into the global market.
This is more than a patch for improving video. It is a bid for leadership in creating virtual sandboxes where future machines can be trained. Russian science has shown that it can set trends not only in language models but also in the most challenging segment of generative AI – where code meets the laws of physics.









































