Sber has released Kandinsky WM 1.0, a family of generative AI models designed to create video that closely follows the laws of physics. The company said the goal was to build a neural network capable of generating synthetic data for training Physical AI systems, including robots and autonomous devices.
The source code and model weights have been published on the Hugging Face platform under an MIT license, meaning developers and research companies can use them free of charge.
How the models were trained
Kandinsky WM 1.0 is built on the Kandinsky 5.0 Video Lite model, which was further trained on several million real video recordings from robot and self-driving car cameras. Sber added footage of manufacturing processes and complex mechanical interactions, along with specialized training data aimed at correcting violations of basic physical laws.
The models currently generate five-second videos, each depicting a single action or situation, such as handling objects, road maneuvers or manufacturing processes.
Where the models are already in use
Sber said Kandinsky WM is a first step toward full world models that can forecast how events unfold over time. The company said the models are already used at its own Robotics Center to train robotic systems, and by specialists at the AIRI institute to build synthetic road scenarios for self-driving vehicle simulators.
