MiniMax H3 获 SGLang Diffusion Day-0 支持,可本地生成带原生音频的 2K 视频
MiniMax H3 已接入 SGLang Diffusion,并获得 Day-0 服务支持。该开放权重视频模型支持文本、图像、视频和音频输入,最高可生成 2K、24fps、最长 15 秒且带立体声的片段。
MiniMax 表示,H3 现已在 LMSYS 维护的 SGLang Diffusion 中提供 Day-0 支持,开发者可以在本地进行部署和定制。相关信息显示,该模型面向视频生成、动效设计、电商创意、动画和视频编辑等场景,并可适配 NVIDIA Blackwell、Hopper 以及部分 AMD 硬件平台。
H3 支持在同一上下文中处理文本、图像、视频和音频,并生成原生立体声视频。ComfyUI 方面也称其已在模型发布当天提供支持,开放权重版本可用于本地推理。
不过,关于其性能相当于 Seedance 2.0、成本仅为三分之一,以及特定显卡即可免费运行等说法目前主要来自宣传材料,尚缺乏独立、可复现的对比测试。此外,部分 AMD 证据指向 MiniMax M3,而非 H3,因此不能直接作为 H3 AMD 兼容性的证明。
来源证据
MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Videoblog.comfy.org · supportingComfyUI Newsletter # ComfyUI Newsletter # MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video ### An open-weights omni-modal video model with real stereo sound and 2K output — this powerful model is greatly optimized in ComfyUI and can run locally on a 3060. Rob's avatar Alexis Rolland's avatar MiniMax H3 dropped today with open weights, and it’s natively supported in ComfyUI as of this morning. Day zero. This is a next-generation open-weights video model. Feed it text, images, video, or audio and it generates video with real stereo sound, up to 2K, up to 15 seconds a clip. It is MiniMax’s third-generation video model, following Hailuo 01 and Hailuo 02, and the first the company has released with open weights. ## Model Highlights [...] ## Optimized for local inference in ComfyUI Getting H3 to run well on consumer hardware took significant machine learning engineering. We found that the model's modulation weights (~40% of the total parameters) could be
MiniMax H3 Open Weights Land With Native ComfyUI Supportcomfyui-wiki.com · supporting| Component | Files | --- | | Diffusion model | `minimax_h3_fl2va_bf16` (61.7 GB), `minimax_h3_fl2va_int8_convrot` (31.7 GB), `minimax_h3_fl2va_pruned_int8_convrot` (19.5 GB), plus matching `ref2va_` variants | | Text encoder | `qwen3vl_32b_minimax_h3_bf16` (48.0 GB), `int8_convrot` (25.3 GB), `nvfp4_awq` (14.6 GB) | | VAEs | `minimax_h3_video_vae_fp16` (4.9 GB), `minimax_h3_audio_vae_fp32` (0.6 GB) | The pruned INT8 checkpoints are about 40% smaller than the standard INT8 files thanks to precomputed adaLN curve tables, and the NVFP4 AWQ text encoder runs on any GPU. H3-Base deploys through SGLang, vLLM, diffusers, and ComfyUI. ## Native ComfyUI support merged [...] Ref2VA accepts up to 9 reference images, 3 reference videos (2-15 s each, 15 s total max), and 3 reference audio clips (2-15 s each, 15 s total max, always alongside an image or video). Mixed input is capped at 12 files total. ## Open weights on Hugging Face Two repositories are now live: MiniMaxAI/MiniMax-H3 (Huggin
MiniMax H3 Gains Day-0 Support in SGLang Diffusion · Diggdigg.com · supporting5090 ordered, SGLang installed, H3 downloaded. 🤩 @MiniMax\_AI H3 is live in SGLang Diffusion, with day-0 serving support 🎬 This open model matches Seedance 2.0 at 1/3 the cost, or $0 if you run it locally on 2x 5090 or 1 RTX 6000. With SGLang Diffusion, you can build visual concepts, motion design, e-commerce creatives,… Show more H3 now has Day 0 support in @lmsysorg -- local, customizable, and ready to build on across @NVIDIAAI and @AMD hardware. #MiniMaxH3 #OpenWeights @MiniMax\_AI H3 is live in SGLang Diffusion, with day-0 serving support 🎬 This open model matches Seedance 2.0 at 1/3 the cost, or $0 if you run it locally on 2x 5090 or 1 RTX 6000. With SGLang Diffusion, you can build visual concepts, motion design, e-commerce creatives,… Show more [...] ## Sentiment ### Positive Read Many users praised MiniMax H3's day-0 SGLang support for enabling local customizable video generation with open weights and lower cost than Seedance, while a few raised concerns about matching p
SGLang is a high-performance serving framework for large ...github.com · supporting## News [2026/06] 🔥 The next generation of speculative decoding: DFlash and Spec V2 (blog). [2026/04] 🔥 DeepSeek-V4 on Day 0: From Fast Inference to Verified RL with SGLang and Miles (blog). [2026/06] SGLang provides day-0 support for latest open models (Nemotron 3 Ultra, Nemotron 3 Super, Higgs Audio v3 TTS). [2026/02] 🔥 Unlocking 25x Inference Performance with SGLang on NVIDIA GB300 NVL72 (blog). [2026/01] SGLang Diffusion accelerates video and image generation (blog). [2025/12] SGLang provides day-0 support for latest open models (MiMo-V2-Flash, Nemotron 3 Nano, Mistral Large 3, LLaDA 2.0 Diffusion LLM, MiniMax M2). [2025/10] SGLang now runs natively on TPU with the SGLang-Jax backend (blog). More [...] [2025/01] SGLang provides day one support for DeepSeek V3/R1 models on NVIDIA and AMD GPUs with DeepSeek-specific optimizations. (instructions, AMD blog, 10+ other companies) [2024/12] v0.4 Release: Zero-Overhead Batch Scheduler, Cache-Aware Load Balancer, Faster Struct
Day 0 Support for MiniMax M3 on AMD Instinct GPUsamd.com · supporting## Deploying with SGLang By leveraging SGLang with ROCm support, developers can unlock high-throughput serving in ROCm. Support is available in the lmsysorg build of SGLang via docker image using the SGLang MiniMax M3 recipe. `lmsysorg/sglang:-rocm720-mi35x (MI350X / MI355X) lmsysorg/sglang:-rocm700-mi30x (MI300X / MI325X)` The model is decode-bound for text and encoder/prefill-bound for vision, so there are two tuned recipes. (The same server handles both; the vision recipe is a safe superset for mixed workloads.) Notes: Text — MI350X / MI355X baseline (native MXFP8) [...] Notes: Text — MI350X / MI355X baseline (native MXFP8) `SGLANG_USE_AITER=1 sglang serve \ --model-path MiniMaxAI/MiniMax-M3 \ --tp 8 --mem-fraction-static 0.80 \ --quantization mxfp8 --dtype bfloat16 \ --chunked-prefill-size 8192 \ --reasoning-parser minimax-m3 --tool-call-parser minimax-m3-nom \ --trust-remote-code --host 0.0.0.0 --port 8080` Text —MI350X / MI355X baseline (native MXFP8) `SGLANG_USE_AITER=
MiniMax (official) on X: "H3 now has Day 0 support in @lmsysorg -- local, customizable, and ready to build on across @NVIDIAAI and @AMD hardware. #MiniMaxH3 #OpenWeights" / Xx.com · supportingLog inSign up ## Post user avatar MiniMax (official) @MiniMax\_AI H3 now has Day 0 support in @lmsysorg -- local, customizable, and ready to build on across @NVIDIAAI and @AMD hardware. #MiniMaxH3 #OpenWeights user avatar LMSYS Org @lmsysorg Aug 3 @MiniMax\_AI H3 is live in SGLang Diffusion, with day-0 serving support 🎬 This open model matches Seedance 2.0 at 1/3 the cost, or $0 if you run it locally on 2x 5090 or 1 RTX 6000. With SGLang Diffusion, you can build visual concepts, motion design, e-commerce creatives, 00:00 3:02 AM · Aug 3, 202671.4KViews user avatar SGLang @sgl\_project Aug 3 Go create something with MiniMax H3!!! 211 user avatar Chasen @chasen\_liao Aug 3 [...] 211 user avatar Chasen @chasen\_liao Aug 3 SGLang Diffusion能同时覆盖Blackwell/Hopper和AMD MI300X/MI355X,这种跨厂商Day-0支持很罕见 说明牢Mi 从设计阶段就考虑了硬件多样性,而不只是堆参数。对想做私有化部署的企业来说,这直接降低了锁定风险 点赞 96 user avatar Vector @PrasVector Aug 3 Hey this