MiniMax H3 在 ComfyUI 直播中展示开源权重视频生成能力
MiniMax 与 ComfyUI 的直播聚焦 H3 多模态视频模型,演示参考驱动视频、原生同步音频及多镜头创作等工作流。
MiniMax H3 在 ComfyUI 直播中展示了其面向视频生成的多模态能力。根据 ComfyUI 的介绍,该模型可处理文本、图像、视频和音频条件,并支持文生视频、图生视频、首尾帧生成以及参考驱动创作。
直播重点讨论了视频与立体声音频的联合生成,而非在视频完成后单独配音。公开权重检查点支持最高 768p、最长 15 秒的片段;MiniMax 托管服务则宣称可支持最高 2K 输出。
内容还涉及量化、卸载和本地部署工作流,为希望在 ComfyUI 中试用该模型的开发者提供了技术背景。相关材料表明,H3 的定位是将多种视频创作任务整合到同一模型架构中。
来源证据
MiniMax H3 on Comfy — Open-Weight Video Modelcomfy.org · supportingA lone rider crosses a glacial canyon in one continuous move, camera and native audio straight out of H3. A pint-size superhero calls out a towering city monster, character held consistent from a single reference image. One product shot becomes a full scene — the label stays crisp as the can pours out beside a waterfall. ## Choose a plan Access cloud-powered ComfyUI workflows with straightforward, usage-based pricing. Start free. Upgrade when you're ready. 5 free runs on real GPUs — no credit card required. Billed monthly Generates ~380 5s videos\ Billed monthly Generates ~670 5s videos\ Billed monthly Generates ~1,915 5s videos\ Built for teams collaborating on workflows together. Billed monthly Generates ~13,405 5s videos\ Everything in Pro, plus: Coming soon... [...] Now turn your agent into a creative technologist. Comfy Comfy Updated August 2026 # MiniMax H3 is here Full multi-modal I/O, native stereo clip. Up to 2K, 5 to 15s per generation. H3 actually condit
Native Audio + Motion Reference? Testing MiniMax H3 Limitsyoutube.com · supporting### Description 7024 views Posted: 5 Aug 2026 Is MiniMax H3 the king of open-source multimodal AI video? 🎬 Unlike traditional video models, MiniMax H3 generates synchronized high-fidelity audio and video simultaneously while demonstrating astonishing physical commonsense and narrative capability. In this deep-dive tutorial, we explore all three core generation modes in ComfyUI: Text-to-Video (T2V), Image-to-Video (I2V / First-Last Frame), and Reference-to-Video (Ref2V). From Chinese calligraphy ink flows and mechanical puzzle sound effects to liquid pouring physics and motion-driven stunt references, we put the INT8 Pruned model through its ultimate paces. 🔥 What you'll master in this tutorial: [...] 🔥 What you'll master in this tutorial: • 🧠 The H3 Multimodal Architecture: Understanding the FL2VA & Ref2VA base models, Qwen3-VL-32B text encoder, and dual VAEs. • 📐 The 17k+5 Frame Formula: Calculating 24FPS frame lengths correctly for smooth outputs. • ✍️ Structured Timeline Prompt
MiniMax H3 Deep Dive: Open Weights, Prompting Techniques ...youtube.com · supporting# MiniMax H3 Deep Dive: Open Weights, Prompting Techniques & Advanced ComfyUI Workflows ## Oxen 7340 subscribers 46 likes ### Description 704 views Posted: 8 Aug 2026 In this episode, we dive deep into MiniMax H3 - the best Open Weights video generation model as of August, 2026. This is a 33B parameter model that you can run at home. It is on par with Seedance 2.0 for some generations, and can be customized through fine-tuning. We will be showing you how to prompt the model, as well as some super interesting ways you can use the model in ComfyUI or Oxen.ai. [...] Cance which is pretty crazy. You can run this thing locally. I've got it up and running on my GPU under my desk right here. You can run it in Comfy UI. I've seen people run it on their Mac hardware. Uh, and the quality is pretty dang good. So, today we're going to be digging into this model, showing you some live generations, and showing you how to fine-tune it. We're going to try to pack a lot into this 45 minutes here. So
MiniMax H3: A New Open-Weight Video Model, Live in ComfyUIyoutube.com · supporting# MiniMax H3: A New Open-Weight Video Model, Live in ComfyUI ## ComfyUI 30800 subscribers 130 likes ### Description 2203 views Posted: 7 Aug 2026 MiniMax H3 is an open-weight, general-purpose multimodal video generation model that works across text, images, video, and audio. In ComfyUI, you can use H3 for text-to-video, image-to-video, first- and last-frame generation, and reference-driven creation. H3 jointly generates the visuals and synchronized stereo audio, including dialogue, sound effects, ambience, and music, rather than adding audio afterward. The open-weight H3 checkpoints support clips up to 15 seconds at 768p. MiniMax’s hosted H3 model also supports generation at up to 2K resolution. [...] with the Miniax community. Awesome. Awesome. So, I've just shared all of those links in our chat here, guys. Um, I'll also add them to the YouTube description after we hop off today. Um, you know, at that same link, you'll be able to revisit the stream, tune in, check it out if you'd l
MiniMax H3: 15-Second 2K AI Video With Native Audio - YouTubeyoutube.com · supportingSo, I said, "A woman walking along a marina is greeted by a talking sea lion who invites her to live under the sea with him." And she accepts and they dive into the water together. I choose that mixed live and animated visual style, and it gave us this. Join me, and I'll show you wonders beneath the surface. [music] Now, one of the other exciting things about this model is that they have made it available as open weight, which means that in theory, you could run this model on a local machine using ComfyUI or something else. However, they have a lot of restrictions around its use. If you are in the excluded territories, the US included, you are not licensed to use them at all. However, anyone can use the model as long as they go through a service that uses their API, like OpenArt or any [...] but maybe 20 15 10 All right, I'll do 10. I dropped in that video as the reference and said from 4 to 7 seconds his shirt turns red and for the entire video the background is an ice cream museum. A
ComfyUI on X: "Livestream Update: @MiniMax_AI H3: A New Open-Weight Video Model, Live in ComfyUI MiniMax H3 is an open-weight, general-purpose multimodal video generation model that works across text, images, video, and audio. In ComfyUI, you can use H3 for text-to-video, image-to-video, https://t.co/ziMNdGmodP" / Xx.com · supportingLivestream Update: @MiniMax\_AI H3: A New Open-Weight Video Model, Live in ComfyUI MiniMax H3 is an open-weight, general-purpose multimodal video generation model that works across text, images, video, and audio. In ComfyUI, you can use H3 for text-to-video, image-to-video, first- and last-frame generation, and reference-driven creation. H3 jointly generates the visuals and synchronized stereo audio, including dialogue, sound effects, ambience, and music, rather than adding audio afterward. The open-weight H3 checkpoints support clips up to 15 seconds at 768p. MiniMax’s hosted H3 model also supports generation at up to 2K resolution. During the stream, we’ll test the model live and discuss how H3 brings multiple generation tasks into one architecture, how its high-compression video [...] tasks into one architecture, how its high-compression video representation improves efficiency, and what developers should know when setting it up locally through ComfyUI. What we'll cover: →MiniMax H3