公告
数据公告

QQ群和tg群已经启用,欢迎加入。公开信息来源均审核后发布;请结合来源、库存和更新时间判断。

社群与联系Telegram 群点击加入Telegram 频道点击订阅联系我们tgAIPricedb交流群979789483
返回资讯列表
product

MiniMax H3 开放权重上线,fal 提供首日构建支持

MiniMax H3 的开放权重已发布,并在 fal 平台提供首日托管与部署支持。开发者可以研究、定制和自托管这一多模态音视频生成模型,但具体使用仍受地区和许可条件限制。

91% VERIFIED

MiniMax H3 的模型权重现已开放,fal 同步提供首日访问和生产级基础设施支持。H3 将文本、图像、视频和音频作为统一上下文,能够生成带原生立体声的短视频,并支持多种参考素材输入。

开放权重意味着开发者可以下载模型进行研究、定制和自行部署,也可以通过托管 API 使用。相关生态已出现 ComfyUI 和云 GPU 部署方案,降低了早期试用门槛。

不过,“开放权重”不等同于整个技术栈都采用开放源代码许可。部分组件和权重可能适用不同条款,部分国家或地区也可能无法进行本地部署,实际使用前应查看官方模型卡和许可说明。

来源证据

What Is MiniMax H3 (Hailuo 3.0)? The Open-Weight ...huggingface.co · supporting

`first_frame` `last_frame` `reference_` In content terms, one generation can inherit a face from an image, a motion/camera move from a video, and a voice from an audio clip at once (Morphic: "a shot can inherit a face, a motion, and a voice at once"), alongside instruction-based editing and voice transfer. Rule of thumb: think in reference sets, not single files — reuse the same 2–4 image + 1 video + 1 audio set across generations for character consistency; the first 5 reference images are free per generation. ## Open weights: the promised release, and the license to watch [...] TL;DR — MiniMax H3 (the official name of what most people call Hailuo 3.0) is the third generation of MiniMax's Hailuo video line, but it deliberately stops behaving like a video model. It reads text, images, video, and audio as one unified context, then generates a 4–15 s clip at 2K/24 fps with native stereo audio — no separate audio stage, no post-hoc upscaler. It supports up to 9 reference images, 3 refer

MiniMax H3: An Open Model Breaking the Boundaries ...minimax.io · supporting

Closed-source models have long dominated video generation, with slower iteration and a less open ecosystem than fields like large language models. To support the open-source community, accelerate compatibility with a broader range of AI hardware, and make it easier for users to build their own customized versions, we plan to open up the model weights in the coming days, subject to applicable laws and regulations. Hardware compatibility has been a key consideration since the earliest stages of H3's design. ### Multimodal context understanding [...] AIH3MultimodalVideo Generation Today, we're launching MiniMax H3, a general-purpose multimodal generation model. H3 understands unified context across text, images, video, and audio, generating video with native stereo sound, up to 15 seconds at 2K resolution. Early testing shows H3 is ready for commercial content creation across a wide range of use cases, excelling at instruction following, accurate text and brand rendering, and V2V motio

MiniMax H3: The Open-Weight Omni-Modal Video Model, ...runpod.io · supporting

## Getting started with MiniMax H3 on Runpod The easy part: ComfyUI has day 0 support for Minimax and templates already set up. So you can get started today with just a few quick downloads. Here are the specs for sizing a GPU: First, deploy a pod using the official Runpod ComfyUI template with ~600 GB of volume disk; we’re going to pull the entire repo which will allow you to test each quant and decide what’s best for your use case. [...] ``` cd ComfyUI/models/ cd ComfyUI/models/ ``` Download the entire repo with the following. This will automatically distribute the files into the appropriate subfolders. ``` hf download Comfy-Org/MiniMax-H3 --local-dir .hf download Comfy-Org/MiniMax-H3 --local-dir . ``` Lastly, on a fresh pod it’s generally good practice to update and restart ComfyUI when you’re using day-0 implementations like this. ‍ Once you're up and running, go to Templates on the left and select the Minimax H3 Text to Image template. At this point, you're ready to gener

MiniMax H3 Open Weights Exclude US, EU, UK, and Korea From Local Deploymenttechtimes.com · supporting

What shipped: H3-Base, a 33.1-billion-parameter dense, single-stream omni-modal transformer that generates video with native stereo audio at a native canvas of 768 pixels on the short edge. The release includes two task-specific checkpoints — FL2VA (text-to-video, first-frame, and last-frame conditioning) and Ref2VA (reference-to-video from images, video clips, and audio) — along with the H3 video VAE, the audio VAE, and the Qwen3-VL-32B text encoder, the last of which ships under Apache 2.0 and is the only genuinely OSI-licensed component in the stack. These details are documented in the H3 model card on Hugging Face. Native ComfyUI support landed the same day, with pull request Comfy-Org/ComfyUI #15224 merged on August 3, integrating joint audio-video generation via four new nodes and [...] MiniMax H3 is a three-module system: H3-Context-IR (input preprocessing), H3-Base (768p audio-video generation), and H3-Regenerate-2K (in-context upscaling to 2K). The open-source release covers o

Hands-On with the New Open-Weights AI Video Modelyoutube.com · supporting

because this is not copyright but still I'm just going to download but you can check it out. So, look. On the input side it's genuinely multimodal. You can feed it up to nine reference images, three reference videos and three audio clips to guide the output. And because the weights are open you can self-host it or hit it through a hosted API. I'm not sure how big that model is. Hopefully it will fit onto my one single H100 but if not we will see. So, this is animated poster as you can see. So, so far what I have seen uh what I really like about H3 is that it's built for real creative work. You can see that how cool it looks. Very impressive stuff by the way. And this is not a sponsored video so don't think that I'm hyping it up. So, it seems that if you go through their model card MiniMax [...] level. And this card tells you in one go what exactly this HiLo H3 is. It's a lightweight open weights video generation model from MiniMax, and the word HiLo literally means sea snail in Chinese

The Fal Episode: Building the Infrastructure Behind AI-Generated Videoyoutube.com · supporting

on file is a great distribution advantage for them. So you know we are also working with them to make sure that happens and on the on the user size side too we have security arrangements data protection agreements all these already set up. So when a new model comes out, there's no paperwork required from the enterprise customer to start using that model in day zero. Got it. So all it's a win-win for the user side and the and the researcher side. Super interesting. I think that we're seeing that flywheel bring in more and more interesting customers. I'm curious from your end, what have been your favorite foul use cases and particularly love to understand like which use cases do you feel like the optimized infrastructure have allowed that weren't possible before? I I think by far in the [...] own. Maybe at like the highest level Google and Facebook maybe some recommendation systems were running on GPUs but really these larger models like LLMs and at time it was most most popular stable d