MiniMax H3 × fal 直播:8月13日详解开放权重多模态视频模型
MiniMax 与 fal 将于太平洋时间8月13日上午11点在 X Spaces 联合直播,邀请创作者与开发者演示 H3 的构建方式,涵盖多模态参考、原生音频、提示词、LoRA 定制与真实工作流。
MiniMax 与 fal 将于太平洋时间8月13日上午11点在 X Spaces 举办联合直播,深入介绍创作者和开发者如何使用 fal 平台上的 MiniMax H3。参与嘉宾包括 fal 创意工程师 Odin Lovis,以及 MiniMax 的 AI 解决方案架构师 Ethan Wei 和 GTM 工程师 Victor SuOrtiz。
直播将涵盖多模态参考、原生音频生成、提示词技巧、LoRA 与定制化、开放权重,以及实际的创意与技术工作流。MiniMax H3 是一款开放权重、通用多模态视频模型,可在同一上下文中理解文本、图像、视频和音频。
资料显示,H3 单次生成请求最多可接受9张参考图像、3段视频片段和3条音轨,支持基于指令的视频编辑,并能输出原生立体声音频,最高支持2K(1440p)分辨率和15秒时长。fal 是 Day 0 合作伙伴,在提供托管 API 的同时也允许直接使用模型权重。
本次活动定位为实战演示,而非包装精美的案例研究,旨在展示使用 H3 生成完整视频时的真实决策、失误与补救过程。
来源证据
MiniMax H3 - Open-Weights General-Purpose Multimodal Video ...fal.ai · supporting## Common questions about MiniMax H3 MiniMax H3 is an open-weights, general-purpose multimodal video model. Instead of a separate model for each task, MiniMax H3 reads text, images, video, and audio in one unified context and generates coherent audiovisual results from any mix of them. It supports text-to-video, first-and-last-frame, reference-to-video, and precise video editing. MiniMax H3 is released with open weights, so it is an open foundation you can explore, customize, and build on rather than a closed endpoint. fal is a Day 0 ecosystem partner, which means you can call the hosted MiniMax H3 API on fal.ai from launch without provisioning GPUs, and still have the option to work with the weights directly for your own research and fine-tuning. [...] ### Sound Composed to Picture Every generation returns native stereo audio: original score, dialogue, foley, and room tone timed to the cut. Give MiniMax H3 a reference recording and it will transfer or clone that voice onto your cha
MiniMax H3 Brings Storytelling Control to AI Videotrilogyai.substack.com · supporting## Where MiniMax H3 fits now H3 is ready for creators and developers who can review every output and discard convincing failures. It fits short narrative scenes, concept footage, visual transformations, stylized motion, dialogue, and frame-controlled transitions. Prompts work better when they describe a visible progression and a final state. Anatomy, science, machinery, exact geometry, logos, text, and multi-clip continuity need stricter supervision. The model’s surface quality makes review more important because obvious ugliness is no longer a reliable warning. [...] ## H3 weights are available MiniMax has released the H3 weights. Developers have started running the model locally on consumer hardware through early ComfyUI support. Published configurations differ in memory, quantization, resolution, and runtime, so they do not yet establish a hardware baseline or repeatable production setup. This article reports hosted API tests completed before the weight release. A follow-up will
Day 0 Support for MiniMax-H3 on AMD Instinct GPUsamd.com · supportingMiniMax-H3 is MiniMax’s third-generation video model and a generational leap over its predecessors. Previous MiniMax video models (the Hailuo series) fragmented generation into separate expert pipelines for T2V, I2V, editing, and reference-driven tasks. H3 collapses all of these into a single unified model that jointly understands text, images, video, and audio, and generates video with native stereo audio at up to 2K resolution (1440p) and 15 seconds — up from 1080p and ~10 s in the prior generation. It also adds instruction-based video editing and omni-reference input, enabling subject-driven animation, voice cloning, and lip-sync in one pipeline. With open weights released under the MiniMax Community License, H3 is the strongest open-weight video generation model available today.
MiniMax H3 API Models | each::labseachlabs.ai · supportingText-to-Video and Image-to-Video (Hailuo V2.3, V2, V1 variants): Generate realistic videos with Pro, Standard, Fast, and Live modes. For instance, creators build dynamic marketing clips or social media reels; input an image of a cheetah and prompt "Cheetah turns toward the camera, sprinting across savanna at sunset" to produce a smooth 5-second clip with lifelike motion and high-resolution textures. [...] Minimax is a leading Chinese AI company specializing in multimodal generation, particularly AI video, music, and image creation through advanced models like Hailuo and Music series. Known for pushing boundaries in native multimodal processing, Minimax integrates text, visuals, audio, and video in a unified framework, rivaling top models like GPT-4o with features such as contextual fluidity, reduced latency, and high-fidelity outputs. Their Hailuo video models excel in cinematic-quality synthesis from text or images, while Music models deliver professional-grade tracks with precise str
MiniMax H3 Opens AI Video to Developers: Copyright Lawsuit ...techtimes.com · supportingThe model accepts any combination of text, images, video clips, and audio as input simultaneously. Specifically, it accepts up to nine reference images, three video clips, and three audio tracks in a single generation request — each contributing a different layer of creative control. A production team can supply character reference images, a video sample showing desired camera movement, and an audio reference defining the soundscape, and H3 produces a clip incorporating all of them in a single pass. MiniMax calls this the "omni-reference" system. The model also supports instruction-based editing: a creator can specify a change to part of an existing clip — swap a product, rewrite visible signage, change a background — and H3 applies the edit while leaving the rest of the frame intact. [...] MiniMax released H3 today — a multimodal video generation model that ranks first in video editing among all models tracked by independent benchmarking firm Artificial Analysis — and simultaneously
Building an AI Music Video Live with MiniMax H3: Full Multimodal Workflowyoutube.com · supporting### Description 78 views Posted: 3 Aug 2026 Today, we’re finding out if MiniMax H3 can survive an entire music video. 🎬 On ImagineArt LIVE, we’re building the music video from scratch—live. Concept, visual language, shots, motion, edits. The whole slightly irresponsible workflow. No polished case study. You’ll see the decisions, mistakes, and saves as they happen. MiniMax H3 can work across text, images, video, and audio. We’re putting that multimodal control to work across a real sequence—not one lucky clip. The mission: make every shot feel like it belongs in the same world. We’ll shape the look, generate the scenes, and assemble the final cut. Then we’ll see what H3 nails, what needs fixing, and what breaks dramatically. Watch us build the music video today on ImagineArt LIVE. [...] live on the stream. I think it was last week or the week before. I can't remember. This is week four of the stream. I can't believe that. Um, we did it like an hour. Well, it was a little longer than th