xAI’s rise in AI video generation has been staggering. On May 31, 2026, xAI officially released Grok Imagine Video 1.5 Preview, which has already climbed to the Artificial Analysis Video Arena. Number one on the Image-to-Video leaderboardWith an Elo rating of 1404 ±6, not only does it improve by 52 points over version 1.0, but it also successfully surpasses strong competitors like ByteDance’s Seedance 2.0. Looking back just 11 months ago, xAI had no video generation product whatsoever; now it has reached the pinnacle of global AI video generation.

From Zero to First: xAI’s 11-Month Miracle
xAI’s foray into AI video generation began with a quiet acquisition in March 2025, when it purchased video generation startup Hotshot (Natural Synthetics Inc.), which owned ready-made video foundation models like Hotshot-XL and Hotshot Act One. This move provided xAI with the technical foundation to enter the video generation space.

After that, xAI advanced product iteration at an extremely fast pace. On July 28, 2025, Grok Imagine Beta launched, with initial videos limited to just 6 seconds at 1280×720 resolution, and issues such as facial drift and physics simulation problems. That October 5th, version 0.9 was released, with frame rate increased to 24 FPS and native audio synchronization significantly improved. On January 28, 2026, the Grok Imagine API opened to developers, with the model ranking first on both the Text-to-Video and Image-to-Video leaderboards. On February 3rd, version 1.0 officially launched, supporting 720p and 10-second videos, with API pricing set at $4.20 per minute.
Technical Breakthroughs in Grok Imagine Video 1.5
Grok Imagine Video 1.5 Preview achieves significant improvements across multiple dimensions compared to version 1.0, with the most notable enhancement being the native synchronized audio feature. Unlike previous models that added audio after video generation, version 1.5 adopts a “single inference” architecture that generates video frames and audio simultaneously in one inference process, covering dialogue, lip-sync, ambient sounds, and background music. This design makes audio-video synchronization more natural, which is critical for film-level output. Elon Musk himself confirmed the official release of Grok Imagine 1.5 on June 4, 2026, and demonstrated an AI-generated trailer for The Iliad (Troy), further validating the model’s film-grade output capability, with indeed substantial quality improvements.
Iliad (Troy) trailer made by Grok Imagine 1.5, which was just released pic.twitter.com/o0zITVlvpn
— Elon Musk (@elonmusk) June 4, 2026
In terms of video length, version 1.5 increases the single video limit from 10 to 15 seconds, allowing users to specify any duration between 1 and 15 seconds, providing more flexibility for narrative creation. Regarding generation speed, Grok Imagine Video 1.5 generates a 5-second video at 720p quality in approximately 20-30 seconds, making it 2 to 3 times faster than ByteDance Seedance 2.0, significantly reducing inference bottlenecks in the production pipeline. Physical realism has also seen measurable improvements, with details such as cloth dynamics, water simulation, hair movement, and object interactions becoming more precise. Character deformation issues in high-action scenes are reduced, micro-expressions are clearer, and rendering quality for translucent and glass materials has improved.
Aurora Architecture and Colossus 2’s Computational Advantages
Behind xAI’s technical advantages lie two critical infrastructure components. The first is the Colossus 2 supercomputer cluster in Memphis, Tennessee, deploying approximately 555,000 NVIDIA GPUs, making it one of the world’s largest single AI computing clusters. The second is xAI’s self-developed Aurora engine: a self-regressive Mixture-of-Experts (MoE) architecture capable of simultaneous token prediction across text, images, video, and audio.
The core advantage of MoE architecture lies in its ability—unlike traditional dense models—to activate only dedicated subnetworks for specific inputs (for instance, realistic images and action scenes correspond to different experts). This enables the model to achieve higher parameter counts and quality ceilings without increasing inference costs. Aurora integrates text, images, and audio processing from the training stage, meaning video and audio are temporally aligned without the need for post-production stitching—this forms the technical foundation of native synchronized audio.
API Pricing and Developer Support
Grok Imagine Video 1.5 Preview is currently available via API, with the model alias being grok-imagine-video-1.5-2026-05-30Pricing structure is $0.08 per second for 480p and $0.14 per second for 720p, which translates to approximately $1.40 for a 10-second 720p video. Text input is free, while image input is $0.01 per request. API rate limit is 60 requests per minute, with service regions covering us-east-1 and eu-west-1.
Besides image-to-video conversion, the API also supports multiple workflows including text-to-video, video editing, multi-image editing, and reference videos, with support for video chaining to build longer multi-shot narratives.
Actual performance falls short of official claims
While the official promotional video for Grok Imagine Video 1.5 currently looks quite good in terms of quality, there aren’t enough examples available online yet. Based on the few comparison cases I’ve seen, I personally feel it hasn’t surpassed Seedance 2.0. After all, there’s still a significant gap between benchmark scores and real-world performance (for example, Happy Horse’s generation quality is quite far from Seedance 2.0, even though their scores are actually close). If the actual generation quality can maintain a certain standard, it should bring more users to xAI.
Arena AI has Grok Imagine Video 1.5 Preview ranked #1 right now, so I had to put it up against Seedance 2.0.
Seedance still feels stronger overall but Grok Imagine is getting close enough that this comparison actually matters.Full video with my thoughts below — including what… pic.twitter.com/V6Zfl8tzXd
— JSFILMZ (@JSFILMZ0412) May 31, 2026
Grok Imagine Video 1.5 Preview vs Kling 3.0.
Same prompts, wildly different results.
Which side are you on? 👇 https://t.co/IMJnDYV5Zl pic.twitter.com/dGLtm4OxNY— JSFILMZ (@JSFILMZ0412) June 1, 2026
Source: KOCPC Chinese