Веса MiniMax H3 открыты, а fal предоставляет поддержку с первого дня
Веса MiniMax H3 опубликованы, а платформа fal предлагает доступ через API и инфраструктуру для запуска с первого дня. Разработчики могут изучать, адаптировать и размещать модель самостоятельно, с учетом региональных и лицензионных ограничений.
MiniMax открыла веса модели H3, а fal сразу предоставила доступ через хостируемую инфраструктуру. H3 объединяет текст, изображения, видео и аудио в едином мультимодальном контексте и создает видео с нативным стереозвуком, поддерживая несколько типов референсных материалов.
Разработчики могут исследовать модель, изменять ее под собственные задачи и запускать ее самостоятельно. Для более быстрого тестирования также доступны размещение через API, интеграции с ComfyUI и варианты развертывания на облачных GPU.
При этом открытые веса не означают, что все компоненты системы распространяются по лицензиям открытого исходного кода. Кроме того, в отдельных регионах локальное развертывание может быть ограничено, поэтому перед использованием необходимо проверить карточку модели, лицензии компонентов и местные требования.
Источники
What Is MiniMax H3 (Hailuo 3.0)? The Open-Weight ...huggingface.co · supporting`first_frame` `last_frame` `reference_` In content terms, one generation can inherit a face from an image, a motion/camera move from a video, and a voice from an audio clip at once (Morphic: "a shot can inherit a face, a motion, and a voice at once"), alongside instruction-based editing and voice transfer. Rule of thumb: think in reference sets, not single files — reuse the same 2–4 image + 1 video + 1 audio set across generations for character consistency; the first 5 reference images are free per generation. ## Open weights: the promised release, and the license to watch [...] TL;DR — MiniMax H3 (the official name of what most people call Hailuo 3.0) is the third generation of MiniMax's Hailuo video line, but it deliberately stops behaving like a video model. It reads text, images, video, and audio as one unified context, then generates a 4–15 s clip at 2K/24 fps with native stereo audio — no separate audio stage, no post-hoc upscaler. It supports up to 9 reference images, 3 refer
MiniMax H3: An Open Model Breaking the Boundaries ...minimax.io · supportingClosed-source models have long dominated video generation, with slower iteration and a less open ecosystem than fields like large language models. To support the open-source community, accelerate compatibility with a broader range of AI hardware, and make it easier for users to build their own customized versions, we plan to open up the model weights in the coming days, subject to applicable laws and regulations. Hardware compatibility has been a key consideration since the earliest stages of H3's design. ### Multimodal context understanding [...] AIH3MultimodalVideo Generation Today, we're launching MiniMax H3, a general-purpose multimodal generation model. H3 understands unified context across text, images, video, and audio, generating video with native stereo sound, up to 15 seconds at 2K resolution. Early testing shows H3 is ready for commercial content creation across a wide range of use cases, excelling at instruction following, accurate text and brand rendering, and V2V motio
MiniMax H3: The Open-Weight Omni-Modal Video Model, ...runpod.io · supporting## Getting started with MiniMax H3 on Runpod The easy part: ComfyUI has day 0 support for Minimax and templates already set up. So you can get started today with just a few quick downloads. Here are the specs for sizing a GPU: First, deploy a pod using the official Runpod ComfyUI template with ~600 GB of volume disk; we’re going to pull the entire repo which will allow you to test each quant and decide what’s best for your use case. [...] ``` cd ComfyUI/models/ cd ComfyUI/models/ ``` Download the entire repo with the following. This will automatically distribute the files into the appropriate subfolders. ``` hf download Comfy-Org/MiniMax-H3 --local-dir .hf download Comfy-Org/MiniMax-H3 --local-dir . ``` Lastly, on a fresh pod it’s generally good practice to update and restart ComfyUI when you’re using day-0 implementations like this. Once you're up and running, go to Templates on the left and select the Minimax H3 Text to Image template. At this point, you're ready to gener
MiniMax H3 Open Weights Exclude US, EU, UK, and Korea From Local Deploymenttechtimes.com · supportingWhat shipped: H3-Base, a 33.1-billion-parameter dense, single-stream omni-modal transformer that generates video with native stereo audio at a native canvas of 768 pixels on the short edge. The release includes two task-specific checkpoints — FL2VA (text-to-video, first-frame, and last-frame conditioning) and Ref2VA (reference-to-video from images, video clips, and audio) — along with the H3 video VAE, the audio VAE, and the Qwen3-VL-32B text encoder, the last of which ships under Apache 2.0 and is the only genuinely OSI-licensed component in the stack. These details are documented in the H3 model card on Hugging Face. Native ComfyUI support landed the same day, with pull request Comfy-Org/ComfyUI #15224 merged on August 3, integrating joint audio-video generation via four new nodes and [...] MiniMax H3 is a three-module system: H3-Context-IR (input preprocessing), H3-Base (768p audio-video generation), and H3-Regenerate-2K (in-context upscaling to 2K). The open-source release covers o
Hands-On with the New Open-Weights AI Video Modelyoutube.com · supportingbecause this is not copyright but still I'm just going to download but you can check it out. So, look. On the input side it's genuinely multimodal. You can feed it up to nine reference images, three reference videos and three audio clips to guide the output. And because the weights are open you can self-host it or hit it through a hosted API. I'm not sure how big that model is. Hopefully it will fit onto my one single H100 but if not we will see. So, this is animated poster as you can see. So, so far what I have seen uh what I really like about H3 is that it's built for real creative work. You can see that how cool it looks. Very impressive stuff by the way. And this is not a sponsored video so don't think that I'm hyping it up. So, it seems that if you go through their model card MiniMax [...] level. And this card tells you in one go what exactly this HiLo H3 is. It's a lightweight open weights video generation model from MiniMax, and the word HiLo literally means sea snail in Chinese
The Fal Episode: Building the Infrastructure Behind AI-Generated Videoyoutube.com · supportingon file is a great distribution advantage for them. So you know we are also working with them to make sure that happens and on the on the user size side too we have security arrangements data protection agreements all these already set up. So when a new model comes out, there's no paperwork required from the enterprise customer to start using that model in day zero. Got it. So all it's a win-win for the user side and the and the researcher side. Super interesting. I think that we're seeing that flywheel bring in more and more interesting customers. I'm curious from your end, what have been your favorite foul use cases and particularly love to understand like which use cases do you feel like the optimized infrastructure have allowed that weren't possible before? I I think by far in the [...] own. Maybe at like the highest level Google and Facebook maybe some recommendation systems were running on GPUs but really these larger models like LLMs and at time it was most most popular stable d