8/3/2026, 1:02:19 PM · video-generation

MiniMax Launches H3 Multimodal Model With Video Generation at 2K Resolution

MiniMax's open-weight H3 model unifies text, image, video, and audio into a single generation pass, producing clips up to 15 seconds at 2K resolution with native stereo sound.

Shanghai-based artificial intelligence company MiniMax launched MiniMax H3 on July 31, 2026, <cite index="3-1">the successor to its Hailuo product line</cite> and the company's most capable video-generation system to date. <cite index="16-3,16-4">H3 is a general-purpose, omni-modal generative system that supports unified understanding of multimodal contexts composed of text, images, video, and audio, and can generate video with native stereo audio at resolutions up to 2K and durations of up to 15 seconds.</cite>

Architecture and Technical Specifications

<cite index="11-3">H3 is powered by a 33.1-billion-parameter dense single-stream omni transformer with a Qwen3-VL-32B text encoder.</cite> MiniMax describes four core enabling technologies on its official blog: <cite index="10-5">Contextual Omni Representation, H3-VAE (Variational AutoEncoder), H3-Omni Transformer, and In-Context Regeneration.</cite>

<cite index="14-13,14-14">H3-VAE is a full tokenizer overhaul whose high compression ratio delivers a stated 4× gain in effective sequence length, cutting training and inference cost and serving as the enabling technology for native 2K output.</cite> <cite index="4-6,4-7">In-Context Regeneration replaces a conventional super-resolution module: the base model regenerates its own lower-resolution output in-context by re-reading the original multimodal context, which recovers small text and fine detail relevant to brand and product rendering.</cite>

Output specifications documented in the official model card include <cite index="15-1">durations from 4 to 15 seconds, broad aspect-ratio support including 16:9, 1:1, 3:4, and 9:16, 24 frames per second (FPS), and 32 kHz stereo audio.</cite> <cite index="12-5">Every request returns a single MP4 containing H.264 video and native stereo audio, generated jointly with the video by the same diffusion transformer, not dubbed on afterward.</cite>

Unified Multimodal Design

A distinctive aspect of H3 is its rejection of the task-specialist architecture common in prior video systems. <cite index="4-13">Previous video stacks split into text-to-video, image-to-video, first-and-last-frame, subject reference, motion reference, and video editing — each often a separate expert model.</cite> <cite index="16-5">Thanks to its task-generalization-oriented system design, H3 already possesses broad multimodal context understanding and generation capabilities at the pre-training stage, enabling performance in following complex multimodal instructions.</cite>

<cite index="1-5,1-6">Early testing shows H3 is positioned for commercial content creation across a range of use cases, excelling at instruction following, accurate text and brand rendering, and video-to-video (V2V) motion transfer, with stated applications in advertising, branding, e-commerce, product design, UI/UX, and gaming.</cite>

Open-Weight Release

<cite index="11-4,11-5">MiniMax H3 was officially open-sourced on August 3, 2026 under the MiniMax H3 Community License Agreement, with the open release covering H3-Base as two task-specific checkpoints: FL2VA (text-to-video and first/last-frame conditioning) and Ref2VA (reference-based generation).</cite> <cite index="11-6">The H3-Context-IR preprocessing system and the H3-Regenerate-2K upscaling module remain hosted application programming interface (API) services.</cite>

<cite index="2-6">The move extends the open-weight strategy increasingly common among Chinese AI developers into the video-generation segment, where many leading models remain proprietary.</cite>

Competitive Context and Pricing

<cite index="2-9">Reuters reported in early July, citing a person with direct knowledge, that MiniMax planned to launch H3 later that month.</cite> <cite index="2-10">Competition among Chinese video-model developers has escalated since ByteDance released Seedance 2.0 earlier this year.</cite> <cite index="14-5">Citing evaluation firm Artificial Analysis, the South China Morning Post reports that H3 leads in video editing while trailing Google's Gemini Omni Flash in text-to-video and sitting behind both Seedance 2.0 and Gemini Omni Flash in image-to-video.</cite>

<cite index="2-12">MiniMax said generating 2K video with H3 would cost less than one-third of mainstream rival products.</cite> <cite index="9-2">MiniMax completed a Hong Kong initial public offering (IPO) in January 2026, reportedly raising about $619 million at a roughly $4 billion valuation, with backers including Alibaba and Tencent.</cite>

Sources

  1. [1]
    MiniMax H3: An Open Model Breaking the Boundaries Between Tasks and Modalities - MiniMax Research | MiniMax
  2. [2]
    MiniMax H3 Video Model: 2K Clips, Stereo Sound, Open Weights | 2026 - News and Statistics - IndexBox
  3. [3]
    MiniMax H3 Launches: 2K AI Video With Native Audio
  4. [4]
    MiniMax Releases MiniMax H3: An Omni-Modal Video Model That Generates 15-Second 2K Clips With Native Stereo Audio - MarkTechPost
  5. [5]
    MiniMax H3: Open Omni-Modal Video Model With Native ...
  6. [6]
    MiniMax H3 - Open-Weights General-Purpose Multimodal Video Model | fal
  7. [7]
    MiniMax Release Notes - July 2026 Latest Updates - Releasebot
  8. [8]
    MiniMax H3 Launches as a General-Purpose Multimodal ...
  9. [9]
    MiniMax H3 (Hailuo 3.0): 2K AI Video, Explained
  10. [10]
    MiniMax H3: Open Omni-Modal Video Model With Native Audio | ComfyUI Wiki
  11. [11]
    MiniMaxAI/MiniMax-H3 | vLLM Recipes
  12. [12]
    MiniMax H3 HuggingFace: Open Source Download, Setup, and Review
  13. [13]
    MiniMaxAI/MiniMax-H3 · Hugging Face
  14. [14]
    MiniMax H3: The Complete Guide to the Open-Weight 2K | PixMind
  15. [15]
    MiniMax H3 (Hailuo 3.0): full specs and input limits
  16. [16]
    Anikuku