MAGI Preview Revolutionizes Video Generation with MoE Architecture, Slashes Costs, and Enhances Audio-Video Sync
September 3, 2026
MAGI Preview reveals three takeaways: rethinking video generation costs for high-frequency use, lowering entry barriers with open-source models, and advancing audio-video synchronization as a standard feature for cohesive content.
The model deploys an ultra-fine-grained MoE architecture with 12 heads per layer, 256-dimensional subspaces, and 72 active small experts per token, delivering high capacity through sparse activation.
Key figures include Cao Yue, founder of Sand.ai, along with investors Su Hua and Matrix Partners, with VidMuse as Sand.ai’s commercial product and the company raising over $100 million in 2026.
Sand.ai uses a single-stream audio-video unified architecture that processes text, video, and audio within one Transformer backbone, employing shared and modality-specific experts to synchronize lip movements, actions, and sounds.
MAGI Preview was released on August 5, 2026 as an open-source 100-billion-parameter Mixture-of-Experts video generation model, totaling about 114 billion parameters with roughly 6 billion activated per forward pass.
Historical context is drawn from DeepSeek, illustrating the shift from dense to MoE architectures for cost efficiency and scale, now applied to video generation.
The approach seeks to break the cost-quality-speed trade-off by combining MoE capacity with unified multi-modal generation for high-frequency, low-cost video production.
MAGI Preview ranks sixth on the Artificial Analysis text-to-video leaderboard and can produce a 10-second 1080p video at about 0.5 yuan, signaling a sharp drop in inference costs.
Beyond price, the model enables iterative, high-frequency experimentation and production workflows in AI-generated video content.
To tackle video-specific challenges, MAGI Preview introduces a Head Parallel mechanism for optimized cross-device communication, along with the MagiMoE operator library and MagiMuon optimizer to streamline routing, computation, and training.
Summary based on 1 source
