- xAI has unveiled Grok Imagine Video 1.5, which features faster generation and improved audio and physics.
- The new model nearly doubles generation speed, producing 6-second, 720p videos in about 25 seconds, down from over 40 seconds in the previous version.
- Grok Imagine Video 1.5 offers better sound effects, ambience, and dialogue, all generated in the same pass and synchronized with the action.
- According to xAI, this model is its best image-to-video model yet, addressing issues with fewer warps and more believable weight and momentum.
xAI has officially launched Grok Imagine Video 1.5, touted as its best image-to-video model to date. This version significantly enhances the user experience with faster generation speeds and improved audio quality.1
The model can now create 6-second videos at 720p resolution in approximately 25 seconds, a remarkable improvement from the previous model's time of over 40 seconds.
According to xAI, Grok Imagine Video 1.5 addresses critical issues in video generation, including better motion, better physics, and better audio. The new model generates sound effects, ambience, and dialogue in a single pass, ensuring that audio is synchronized with the action on screen.3

Users can expect clearer speech and improved synchronization, which enhances the overall viewing experience. The model also reduces visual artifacts, resulting in fewer warps and more believable weight and momentum in the generated clips.4
Grok Imagine Video 1.5 is now generally available through the xAI API as "grok-imagine-video-1.5", marking a significant step forward in the realm of AI-generated video content.
As xAI continues to innovate, the company aims to create visuals and sound that feel more natural and cohesive, setting a new standard in the industry.
“Grok Imagine Video 1.5 is now generally available, offering nearly double the generation speed for 6-second, 720p videos. The new model features improved sound effects, ambience, and dialogue, all generated in the same pass.”
