MiniMax H3 ranks near the top of video arenas, but leaderboard results vary
MiniMax H3 is an open-weight multimodal video model that accepts text, images, video, and audio and can generate clips with native stereo sound. Third-party leaderboards place it near the top, but reported scores and open-weight rankings differ by snapshot and evaluation mode.
MiniMax H3 is presented as a general-purpose multimodal generation system that combines text, image, video, and audio understanding. Its model documentation describes clips of up to 15 seconds at resolutions reaching 2K, with native audio and support for multiple reference assets. Availability and maximum output specifications can differ across hosted and local implementations, and some auxiliary components are not included with the released weights.
Design Arena reports H3 in second place overall with an Elo of 1325. Artificial Analysis reports different values in separate leaderboard snapshots, including 1305 and 1240, while its audio-inclusive open-weight view places LTX-2.3 variants ahead. The claim that H3 is the leading open-weight video model therefore depends on the benchmark mode, audio settings, and collection date. These arena results should be treated as directional rather than a definitive industry benchmark.
Source evidence
Text to Video Leaderboard - Top AI Video Modelsartificialanalysis.ai · supportingLTX-2.3 Fast currently leads among open weights Text to Video models with audio in the Artificial Analysis Text to Video Arena with an Elo score of 980, followed by LTX-2.3 Pro (Elo 962) and LTX-2 Fast (Elo 945). Gemini Omni Flash currently leads the Artificial Analysis Text to Video Arena (without audio) with an Elo score of 1324. The top Text to Video models without audio by Elo rating are: 1. Gemini Omni Flash (Elo 1324), 2. MiniMax H3 (Elo 1305), 3. HappyHorse-1.0 (Elo 1284), 4. Dreamina Seedance 2.0 720p (Elo 1266), 5. HappyHorse-1.1 (Elo 1263). Rankings are based on blind user votes in the Artificial Analysis Video Arena. [...] | Range | Creator | Model | Elo | 95% CI | Samples | Released | API Pricing 1 | --- --- --- --- | | 1 | 1-2 | Google logoGoogle | Gemini Omni Flash | 1,244 | -8/8 | 10,010 | May 2026 | $6.00 /min | | 2 | 1-2 | MiniMax logoMiniMax | MiniMax H3 | 1,240 | -10/10 | 6,025 | Jul 2026 | $7.80 /min | | 3 | 3 | ByteDance Seed logoByteDance Seed | Dreamina See
MiniMax H3 - Open-Weights General-Purpose Multimodal ...fal.ai · supportingMiniMax H3 generates 5 to 15 seconds at 24 FPS. Output is 2K, which puts 1440 pixels on the short edge for ratios between 16:9 and 9:16 and reaches roughly 3.7 megapixels on wider formats, for example 2976x1248 at 21:9. Text-to-video and reference-to-video support 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16, plus an adaptive mode that lets MiniMax H3 pick the best ratio. First-and-last-frame follows the aspect ratio of the uploaded image. [...] Reference-to-video accepts up to 9 reference images, up to 3 reference video clips (2 to 15 seconds each, 15 seconds total), and up to 3 reference audio tracks (2 to 15 seconds each, 15 seconds total), with a maximum of 12 files in total. Audio must be paired with at least one image or video. First-and-last-frame takes a starting image plus an optional end image. Prompts can run up to 7,000 characters. Yes. Every generation includes native stereo audio, covering original score, dialogue, foley, and ambience timed to the picture. MiniMax H3 can also tra
What Is MiniMax H3 (Hailuo 3.0)? The Open-Weight ...huggingface.co · supportingTL;DR — MiniMax H3 (the official name of what most people call Hailuo 3.0) is the third generation of MiniMax's Hailuo video line, but it deliberately stops behaving like a video model. It reads text, images, video, and audio as one unified context, then generates a 4–15 s clip at 2K/24 fps with native stereo audio — no separate audio stage, no post-hoc upscaler. It supports up to 9 reference images, 3 reference video clips, and 3 reference audio clips per generation. On the Artificial Analysis leaderboards it ranks #1 in Video Editing, #2 in Text-to-Video, #3 in Image-to-Video. Weights are promised "in the coming days" under a planned MiniMax Community License (commercial use for organizations under $20M revenue, with attribution), but no Hugging Face model card exists yet as of August [...] Two caveats before you cite these numbers. First, this is a single third-party arena run, not a peer-reviewed suite, and MiniMax's methodology is unpublished — treat the rankings as a prior, not a
MiniMaxAI/MiniMax-H3huggingface.co · supporting# MiniMax H3 ## System Overview MiniMax H3 is a general-purpose, omni-modal generative system. It supports unified understanding of multimodal contexts composed of text, images, video, and audio, and can generate video with native stereo audio at resolutions up to 2K and durations of up to 15 seconds. Thanks to its task-generalization-oriented system design, H3 already possesses broad multimodal context understanding and generation capabilities at the pre-training stage, enabling outstanding performance in following complex multimodal instructions. H3 supports the following input and output specifications: [...] | Model Variant | Input Mode | Specifications | --- | H3-Base-FL2VA | First-and-last-frame mode | Supports zero, one, or two input images. - No image input: Text-to-video mode - One image input: First-frame-to-video or last-frame-to-video generation - Two image inputs: First-and-last-frame-to-video generation | | H3-Base-Ref2VA | Omni-reference mode | Supports multi-mo
China's MiniMax H3 is the first open model to top an AI video ...the-decoder.com · supportingAug 3, 2026 MiniMax releases H3 video model weights, putting an open model at the top of a video ranking for the first time. Artificial Analysis ranks H3 first in Video Editing, second in Text-to-Video, and third in Image-to-Video. The 33-billion-parameter model processes text, images, video, and audio together, generating four- to 15-second clips with stereo sound. According to the model card, a single prompt can include up to nine reference images, three video clips, and three audio clips. Video by MiniMax H3 [...] Video by MiniMax H3 Two pieces remain closed, though. The 2K resolution module and H3-Context-IR, which translates prompts and reference material into a structured intermediate format, aren't included. Running H3 locally in ComfyUI tops out at 768p, and users will need to handle context prep themselves using MiniMax's published prompting guides. The open weights do allow fine-tuning on custom footage, characters, or a specific visual style. One catch on the license side
Design Arena on X: "MiniMax H3 by @MiniMax_AI is 2nd overall on Video Arena with an Elo of 1325. This is a 209 Elo increase from @MiniMax_AI’s previous video model, MiniMax Hailuo-2.3 (Pro), putting them behind Gemini Omni Flash by @GoogleDeepMind and ahead of Seedance 2.0 Mini by @BytePlusGlobal. https://t.co/KjAEs3bQHo" / Xx.com · supportingKamryn Ohly Intelligence @KamrynOhly 23h A defining moment for open weights in video generation! user avatar Alice The Ai Expert @AliceInfoAi 9h MiniMax H3 just set the bar for open video. Huge leap. [...] Log inSign up ## Post user avatar Design Arena Intelligence @DesignArena MiniMax H3 by @MiniMax\_AI is 2nd overall on Video Arena with an Elo of 1325. This is a 209 Elo increase from @MiniMax\_AI’s previous video model, MiniMax Hailuo-2.3 (Pro), putting them behind Gemini Omni Flash by @GoogleDeepMind and ahead of Seedance 2.0 Mini by @BytePlusGlobal. With this performance, they establish themselves as the 2nd video lab overall. Among open weights, the model is 1st overall, ahead of LTX 2.3 by @Lightricks and Kandinsky 5.0 Pro by AI-Forever. By a substantial gap, MiniMax has set a new SOTA on open weight video generation. Congratulations to the @MiniMax\_AI team on the achievement! 10:06 PM · Aug 4, 20267.9KViews user avatar Kamryn O