MiniMax Releases H3: 2K Video With Native Audio, Open Weights Promised

MiniMax released H3 on July 30, 2026 — a general-purpose multimodal video model that generates clips of up to 15 seconds at native 2K resolution with synchronized stereo audio, and takes text, images, video, and audio as reference inputs in a single request. The Shanghai company says it will publish H3’s weights “within days,” which would make it one of the first frontier-class video generators to ship as open weights in a field where the leading systems have stayed closed.
Intermediate
What H3 Does
H3 — marketed to consumers as Hailuo 3.0 — outputs native 2K video at 24fps in durations from 4 to 15 seconds, across aspect ratios from 21:9 through 9:16, with an Extend Video tool that pushes a sequence to roughly 30 seconds. Audio is not a separate pass: dialogue, sound effects, music, and room ambience are generated alongside the frames, which is the difference between a model that produces footage and one that produces a finished clip.
The feature MiniMax is leading with is omni-reference. A single generation request can carry up to 9 reference images, 3 reference video clips, and 3 reference audio clips — capped at 12 files total — with each reference video and audio clip running 2 to 15 seconds and the combined reference duration limited to 15 seconds. Audio references cannot be submitted on their own; they must accompany image or video content. In practice this means one prompt can pin down a character’s face, a product’s branding, a camera move borrowed from existing footage, and a speaker’s voice simultaneously.
The model also supports instruction-based editing — describing a change to characters, objects, scenes, sound, or pacing and having it applied without regenerating the clip from scratch — and video-to-video motion transfer, which maps the movement in one clip onto new subjects. Reference input formats are WAV and MP3 for audio, with AAC or MP3 audio tracks in the returned video; reference images and video must fall between 256 and 5760 pixels with aspect ratios between 2:5 and 5:2.
Cost, Access, and the Open-Weights Question
H3 is available now through MiniMax’s platform API under the model ID MiniMax-H3, and through third-party routers — OpenRouter lists it at minimax/hailuo-3 starting at $0.13 per second of generated video. MiniMax told Reuters that H3 generates 2K video at less than one-third the cost of mainstream competing products, and that the model is designed to run on Chinese-made chips. Access is still described as early, and several details remain unpublished: parameter count, architecture family, rate limits, license terms for generated content, and formal scores on standard video-generation benchmarks.
That last gap matters. Every capability claim above traces back to MiniMax’s own documentation and marketing; H3 has not yet been independently arena-scored against its rivals. The weights themselves are also still pending — as of publication they had not appeared on Hugging Face, so “open weights” remains a stated intention rather than a shipped artifact, and the license terms that will govern them are unknown.
What This Means
Video generation has been the most stubbornly closed corner of generative AI. Text models went open early and often; image models followed; but the systems at the top of the video leaderboards — ByteDance’s Seedance 2.0, Kuaishou’s Kling 3.0, Google’s Veo 3.1 — have remained API-only products. OpenAI is moving in the opposite direction entirely, with its Sora API scheduled for discontinuation at the end of September 2026. A downloadable H3 would be the first time researchers can inspect, fine-tune, and self-host a model in this tier.
The competitive framing is familiar from the text-model side: rather than beat closed rivals on raw fidelity, MiniMax is competing on cost, control, and openness. Kling 3.0 already offers native 4K, and Seedance 2.0 leads audio-inclusive rankings — H3’s pitch is the omni-reference control surface, one-pass audio, and a price roughly a third of the alternatives. For the advertising, e-commerce, product design, and games workflows MiniMax is targeting, reliable character and brand consistency across shots is often worth more than another resolution tier.
For students and researchers, the practical advice is to wait for the weights before drawing conclusions. If they arrive on the timeline MiniMax has stated, the interesting work begins immediately: independent benchmarking against Kling and Seedance, and the fine-tuning and quantization ecosystem that formed around MiniMax’s earlier open-weight text releases arriving in video for the first time.
Related Coverage
- MiniMax M3: Frontier Coding, 1M Context, and Sparse Attention — MiniMax’s open-weight text model from June 2026, and the sparse-attention work behind it
- MiniMax M2.7 Ships as Open Weights: Frontier Agentic Model on Hugging Face — the release pattern H3 would extend into video
- Wan2.2: Alibaba’s Open-Source Breakthrough in AI Video Generation — an earlier Chinese open-weight push in video generation
- OpenAI Launches Sora 2: A New Frontier in AI Video Generation — the closed-API approach H3 is positioned against
Sources
- China’s MiniMax releases H3 video model — Eduardo Baptista, Reuters
- MiniMax Platform — Video Generation guide (H3 model documentation)
- MiniMax official site — H3 announcement
- OpenRouter — MiniMax Hailuo 3 model page and pricing
- MiniMax H3 (Hailuo 3.0): 2K AI Video, Explained — OrcaRouter
- MiniMax H3 AI Video Generator: Native 2K Video With Audio — OpenArt






沪公网安备31011502017015号