Notice
数据公告

QQ群和tg群已经启用,欢迎加入。公开信息来源均审核后发布;请结合来源、库存和更新时间判断。

Community & contactTelegram 群点击加入Telegram 频道点击订阅联系我们tgAIPricedb交流群979789483
Back to news
Products

MiniMax H3 x fal Livestream: Open-Weight Multimodal Video on Aug 13

MiniMax and fal host a live X Spaces session showcasing H3, the open-weight multimodal video model, with creators and developers covering omni-reference input, native audio, LoRA customization, and real workflows.

96% VERIFIED

MiniMax and fal are co-hosting a live X Spaces event on Aug 13, 11 AM PT, to break down how creators and developers are building with MiniMax H3 on fal. Guests include Odin Lovis, Creative Engineer at fal, and Ethan Wei and Victor SuOrtiz from MiniMax.

The session will cover multimodal references, native audio generation, prompting techniques, LoRAs and customization, open weights, and practical creative and technical workflows. MiniMax H3 is an open-weights general-purpose multimodal video model that accepts text, images, video, and audio in one unified context.

According to evidence, H3 supports up to nine reference images, three video clips, and three audio tracks per generation request, includes instruction-based editing, and outputs native stereo audio at up to 2K resolution and 15 seconds. fal is a Day 0 partner, offering a hosted API alongside direct weight access.

The livestream is positioned as a practical walkthrough rather than a polished case study, aiming to show real decisions, mistakes, and saves when generating a full sequence with H3.

Source evidence

MiniMax H3 - Open-Weights General-Purpose Multimodal Video ...fal.ai · supporting

## Common questions about MiniMax H3 MiniMax H3 is an open-weights, general-purpose multimodal video model. Instead of a separate model for each task, MiniMax H3 reads text, images, video, and audio in one unified context and generates coherent audiovisual results from any mix of them. It supports text-to-video, first-and-last-frame, reference-to-video, and precise video editing. MiniMax H3 is released with open weights, so it is an open foundation you can explore, customize, and build on rather than a closed endpoint. fal is a Day 0 ecosystem partner, which means you can call the hosted MiniMax H3 API on fal.ai from launch without provisioning GPUs, and still have the option to work with the weights directly for your own research and fine-tuning. [...] ### Sound Composed to Picture Every generation returns native stereo audio: original score, dialogue, foley, and room tone timed to the cut. Give MiniMax H3 a reference recording and it will transfer or clone that voice onto your cha

MiniMax H3 Brings Storytelling Control to AI Videotrilogyai.substack.com · supporting

## Where MiniMax H3 fits now H3 is ready for creators and developers who can review every output and discard convincing failures. It fits short narrative scenes, concept footage, visual transformations, stylized motion, dialogue, and frame-controlled transitions. Prompts work better when they describe a visible progression and a final state. Anatomy, science, machinery, exact geometry, logos, text, and multi-clip continuity need stricter supervision. The model’s surface quality makes review more important because obvious ugliness is no longer a reliable warning. [...] ## H3 weights are available MiniMax has released the H3 weights. Developers have started running the model locally on consumer hardware through early ComfyUI support. Published configurations differ in memory, quantization, resolution, and runtime, so they do not yet establish a hardware baseline or repeatable production setup. This article reports hosted API tests completed before the weight release. A follow-up will

Day 0 Support for MiniMax-H3 on AMD Instinct GPUsamd.com · supporting

MiniMax-H3 is MiniMax’s third-generation video model and a generational leap over its predecessors. Previous MiniMax video models (the Hailuo series) fragmented generation into separate expert pipelines for T2V, I2V, editing, and reference-driven tasks. H3 collapses all of these into a single unified model that jointly understands text, images, video, and audio, and generates video with native stereo audio at up to 2K resolution (1440p) and 15 seconds — up from 1080p and ~10 s in the prior generation. It also adds instruction-based video editing and omni-reference input, enabling subject-driven animation, voice cloning, and lip-sync in one pipeline. With open weights released under the MiniMax Community License, H3 is the strongest open-weight video generation model available today.

MiniMax H3 API Models | each::labseachlabs.ai · supporting

Text-to-Video and Image-to-Video (Hailuo V2.3, V2, V1 variants): Generate realistic videos with Pro, Standard, Fast, and Live modes. For instance, creators build dynamic marketing clips or social media reels; input an image of a cheetah and prompt "Cheetah turns toward the camera, sprinting across savanna at sunset" to produce a smooth 5-second clip with lifelike motion and high-resolution textures. [...] Minimax is a leading Chinese AI company specializing in multimodal generation, particularly AI video, music, and image creation through advanced models like Hailuo and Music series. Known for pushing boundaries in native multimodal processing, Minimax integrates text, visuals, audio, and video in a unified framework, rivaling top models like GPT-4o with features such as contextual fluidity, reduced latency, and high-fidelity outputs. Their Hailuo video models excel in cinematic-quality synthesis from text or images, while Music models deliver professional-grade tracks with precise str

MiniMax H3 Opens AI Video to Developers: Copyright Lawsuit ...techtimes.com · supporting

The model accepts any combination of text, images, video clips, and audio as input simultaneously. Specifically, it accepts up to nine reference images, three video clips, and three audio tracks in a single generation request — each contributing a different layer of creative control. A production team can supply character reference images, a video sample showing desired camera movement, and an audio reference defining the soundscape, and H3 produces a clip incorporating all of them in a single pass. MiniMax calls this the "omni-reference" system. The model also supports instruction-based editing: a creator can specify a change to part of an existing clip — swap a product, rewrite visible signage, change a background — and H3 applies the edit while leaving the rest of the frame intact. [...] MiniMax released H3 today — a multimodal video generation model that ranks first in video editing among all models tracked by independent benchmarking firm Artificial Analysis — and simultaneously

Building an AI Music Video Live with MiniMax H3: Full Multimodal Workflowyoutube.com · supporting

### Description 78 views Posted: 3 Aug 2026 Today, we’re finding out if MiniMax H3 can survive an entire music video. 🎬 On ImagineArt LIVE, we’re building the music video from scratch—live. Concept, visual language, shots, motion, edits. The whole slightly irresponsible workflow. No polished case study. You’ll see the decisions, mistakes, and saves as they happen. MiniMax H3 can work across text, images, video, and audio. We’re putting that multimodal control to work across a real sequence—not one lucky clip. The mission: make every shot feel like it belongs in the same world. We’ll shape the look, generate the scenes, and assemble the final cut. Then we’ll see what H3 nails, what needs fixing, and what breaks dramatically. Watch us build the music video today on ImagineArt LIVE. [...] live on the stream. I think it was last week or the week before. I can't remember. This is week four of the stream. I can't believe that. Um, we did it like an hour. Well, it was a little longer than th