MiniMax H3 showcased in ComfyUI livestream
A ComfyUI livestream highlighted MiniMax H3’s reference-driven video, synchronized native audio, and multi-shot creation workflows.
MiniMax H3 was featured in a ComfyUI livestream covering the model’s multimodal video-generation capabilities. ComfyUI describes H3 as accepting text, image, video, and audio inputs, with workflows for text-to-video, image-to-video, first-and-last-frame generation, and reference-driven creation.
A central focus was joint generation of video and synchronized stereo audio, rather than adding audio after video rendering. The cited materials say open-weight checkpoints support clips of up to 15 seconds at 768p, while MiniMax’s hosted offering supports output up to 2K.
The session also addressed quantization, offloading, and local setup considerations for ComfyUI users. Taken together, the materials position H3 as a model intended to consolidate several video-creation tasks within one architecture.
Source evidence
MiniMax H3 on Comfy — Open-Weight Video Modelcomfy.org · supportingA lone rider crosses a glacial canyon in one continuous move, camera and native audio straight out of H3. A pint-size superhero calls out a towering city monster, character held consistent from a single reference image. One product shot becomes a full scene — the label stays crisp as the can pours out beside a waterfall. ## Choose a plan Access cloud-powered ComfyUI workflows with straightforward, usage-based pricing. Start free. Upgrade when you're ready. 5 free runs on real GPUs — no credit card required. Billed monthly Generates ~380 5s videos\ Billed monthly Generates ~670 5s videos\ Billed monthly Generates ~1,915 5s videos\ Built for teams collaborating on workflows together. Billed monthly Generates ~13,405 5s videos\ Everything in Pro, plus: Coming soon... [...] Now turn your agent into a creative technologist. Comfy Comfy Updated August 2026 # MiniMax H3 is here Full multi-modal I/O, native stereo clip. Up to 2K, 5 to 15s per generation. H3 actually condit
Native Audio + Motion Reference? Testing MiniMax H3 Limitsyoutube.com · supporting### Description 7024 views Posted: 5 Aug 2026 Is MiniMax H3 the king of open-source multimodal AI video? 🎬 Unlike traditional video models, MiniMax H3 generates synchronized high-fidelity audio and video simultaneously while demonstrating astonishing physical commonsense and narrative capability. In this deep-dive tutorial, we explore all three core generation modes in ComfyUI: Text-to-Video (T2V), Image-to-Video (I2V / First-Last Frame), and Reference-to-Video (Ref2V). From Chinese calligraphy ink flows and mechanical puzzle sound effects to liquid pouring physics and motion-driven stunt references, we put the INT8 Pruned model through its ultimate paces. 🔥 What you'll master in this tutorial: [...] 🔥 What you'll master in this tutorial: • 🧠 The H3 Multimodal Architecture: Understanding the FL2VA & Ref2VA base models, Qwen3-VL-32B text encoder, and dual VAEs. • 📐 The 17k+5 Frame Formula: Calculating 24FPS frame lengths correctly for smooth outputs. • ✍️ Structured Timeline Prompt
MiniMax H3 Deep Dive: Open Weights, Prompting Techniques ...youtube.com · supporting# MiniMax H3 Deep Dive: Open Weights, Prompting Techniques & Advanced ComfyUI Workflows ## Oxen 7340 subscribers 46 likes ### Description 704 views Posted: 8 Aug 2026 In this episode, we dive deep into MiniMax H3 - the best Open Weights video generation model as of August, 2026. This is a 33B parameter model that you can run at home. It is on par with Seedance 2.0 for some generations, and can be customized through fine-tuning. We will be showing you how to prompt the model, as well as some super interesting ways you can use the model in ComfyUI or Oxen.ai. [...] Cance which is pretty crazy. You can run this thing locally. I've got it up and running on my GPU under my desk right here. You can run it in Comfy UI. I've seen people run it on their Mac hardware. Uh, and the quality is pretty dang good. So, today we're going to be digging into this model, showing you some live generations, and showing you how to fine-tune it. We're going to try to pack a lot into this 45 minutes here. So
MiniMax H3: A New Open-Weight Video Model, Live in ComfyUIyoutube.com · supporting# MiniMax H3: A New Open-Weight Video Model, Live in ComfyUI ## ComfyUI 30800 subscribers 130 likes ### Description 2203 views Posted: 7 Aug 2026 MiniMax H3 is an open-weight, general-purpose multimodal video generation model that works across text, images, video, and audio. In ComfyUI, you can use H3 for text-to-video, image-to-video, first- and last-frame generation, and reference-driven creation. H3 jointly generates the visuals and synchronized stereo audio, including dialogue, sound effects, ambience, and music, rather than adding audio afterward. The open-weight H3 checkpoints support clips up to 15 seconds at 768p. MiniMax’s hosted H3 model also supports generation at up to 2K resolution. [...] with the Miniax community. Awesome. Awesome. So, I've just shared all of those links in our chat here, guys. Um, I'll also add them to the YouTube description after we hop off today. Um, you know, at that same link, you'll be able to revisit the stream, tune in, check it out if you'd l
MiniMax H3: 15-Second 2K AI Video With Native Audio - YouTubeyoutube.com · supportingSo, I said, "A woman walking along a marina is greeted by a talking sea lion who invites her to live under the sea with him." And she accepts and they dive into the water together. I choose that mixed live and animated visual style, and it gave us this. Join me, and I'll show you wonders beneath the surface. [music] Now, one of the other exciting things about this model is that they have made it available as open weight, which means that in theory, you could run this model on a local machine using ComfyUI or something else. However, they have a lot of restrictions around its use. If you are in the excluded territories, the US included, you are not licensed to use them at all. However, anyone can use the model as long as they go through a service that uses their API, like OpenArt or any [...] but maybe 20 15 10 All right, I'll do 10. I dropped in that video as the reference and said from 4 to 7 seconds his shirt turns red and for the entire video the background is an ice cream museum. A
ComfyUI on X: "Livestream Update: @MiniMax_AI H3: A New Open-Weight Video Model, Live in ComfyUI MiniMax H3 is an open-weight, general-purpose multimodal video generation model that works across text, images, video, and audio. In ComfyUI, you can use H3 for text-to-video, image-to-video, https://t.co/ziMNdGmodP" / Xx.com · supportingLivestream Update: @MiniMax\_AI H3: A New Open-Weight Video Model, Live in ComfyUI MiniMax H3 is an open-weight, general-purpose multimodal video generation model that works across text, images, video, and audio. In ComfyUI, you can use H3 for text-to-video, image-to-video, first- and last-frame generation, and reference-driven creation. H3 jointly generates the visuals and synchronized stereo audio, including dialogue, sound effects, ambience, and music, rather than adding audio afterward. The open-weight H3 checkpoints support clips up to 15 seconds at 768p. MiniMax’s hosted H3 model also supports generation at up to 2K resolution. During the stream, we’ll test the model live and discuss how H3 brings multiple generation tasks into one architecture, how its high-compression video [...] tasks into one architecture, how its high-compression video representation improves efficiency, and what developers should know when setting it up locally through ComfyUI. What we'll cover: →MiniMax H3