- Zhang Yiming is personally overseeing the development of a new spatial-video model slated for launch as soon as October, according to people familiar with the matter.
- The new model aims to generate interactive virtual worlds for live streams, dramas, and games.
- This model would allow users to create virtual worlds that respond to Pico headset users' voices or movements.
- ByteDance has an existing AI model called Seedance for generating cinematic videos.
- The new spatial-video model is part of ByteDance's strategy to compete with Meta and Alphabet in the world models arena.
- Zhang Yiming has been coordinating efforts across business units and allocating AI resources and computing capacity for this project.
- The timing of the launch is uncertain, and plans may change.
ByteDance is preparing to launch a new AI model for real-time spatial video generation, with founder Zhang Yiming personally overseeing its development. Slated for an October release, this model aims to compete with tech giants Meta and Alphabet in the burgeoning field of virtual and mixed reality.1456
The model, built on ByteDance's existing Seedance technology, will allow users to create interactive virtual worlds for applications such as live streams, short-form dramas, and games. Zhang is coordinating efforts across various business units, leveraging AI resources and computing capacity to enhance the model's capabilities.2
Zhang hopes this new AI model will position ByteDance as a key player in the world models arena, joining experts like Fei-Fei Li and Yann LeCun in exploring visual AI approaches crucial for robotics and gaming. The spatial-video model aims to create immersive experiences, placing users in three-dimensional environments similar to Google's Genie.
If successful, this initiative could open a new front in ByteDance's rivalry with Meta and Apple, both of which have heavily invested in virtual and mixed-reality technologies. The model is designed to respond to user interactions, offering on-demand videos with minimal latency, thus enhancing user engagement in spatial-computing environments.
Additionally, the model seeks to reduce the cost of VR adoption by shifting the computational load to the cloud, making it more accessible for users.
“The model, built on ByteDance's Seedance, would generate interactive virtual worlds for live streams and games, responding to Pico headset users' voices or movements. It offers on-demand videos with about 0.05 seconds latency at 20 frames per second, aiming to lower VR adoption costs by moving spatial content generation to the cloud.”

