Maestro v1.5.5 发布,新增 MiniMax H3 支持
Maestro v1.5.5 已加入对 MiniMax H3 的支持。社区开发者在官方未测试的硬件上运行 H3,并在模型开放权重后不到 48 小时内推出了本地运行和相关工具。
MiniMax 表示,Maestro v1.5.5 现已支持 MiniMax H3。H3 是一款开放权重的通用多模态生成模型,可处理文本、图像、视频和音频,并生成最长 15 秒、最高 2K 分辨率且带原生立体声的视频。上述能力主要来自 MiniMax 的产品说明,独立评测仍然有限。
社区反馈显示,开发者已经在此前未经过 MiniMax 测试的设备上运行 H3,包括游戏显卡和 MacBook,并尝试完全离线执行。相关生态项目还陆续加入了 H3 支持,涵盖 ComfyUI、Diffusers、WanGP 以及面向 Mac 的 MLX 工具。
部分社区用户将本地 H3 称为“免费”或与商业视频模型相当,但这些说法属于个人评价。实际使用仍可能需要满足硬件、显存、软件配置和模型许可等条件。
来源证据
MiniMaxAI/MiniMax-H3huggingface.co · supporting# MiniMax H3 ## System Overview MiniMax H3 is a general-purpose, omni-modal generative system. It supports unified understanding of multimodal contexts composed of text, images, video, and audio, and can generate video with native stereo audio at resolutions up to 2K and durations of up to 15 seconds. Thanks to its task-generalization-oriented system design, H3 already possesses broad multimodal context understanding and generation capabilities at the pre-training stage, enabling outstanding performance in following complex multimodal instructions. H3 supports the following input and output specifications: [...] To reduce the computational cost of long multimodal sequences, H3 natively supports sparse-attention training and inference. The initial open-source release provides inference with full attention only. Our sparse-attention implementation will be released in a future update. [...] The complete H3 system consists of the following three modules:
MiniMax H3 draws local apps and training tools within 48 hours - RuntimeWireruntimewire.com · supportingYan Junjie, MiniMax's founder, chairman, CEO and CTO, saw outside developers run H3 on untested hardware and build new training, optimization and deployment tools within 48 hours of its open-weights release. On August 4, a day after MiniMax published H3's downloadable checkpoints, MiniMax said Maestro v1.5.5 had added H3 support. MiniMax also said developers had run the model on hardware it had never tested, spanning a gaming GPU and MacBooks operating fully offline. MiniMax on X [...] Maestro was one piece of a larger burst of work documented in MiniMax's ecosystem graphic. MiniMax listed native support in ComfyUI and Diffusers on the day the weights shipped. It said WanGP v12.41 followed with a path designed to run in 5 GB to 6 GB of VRAM, while Phosphene and a MiniMax-H3-MLX engine brought one-click local execution to Macs. DiffSynth-Studio added an NF4 build that MiniMax says lowers the requirement to 7 GB to 8 GB of VRAM. Those memory figures come from MiniMax's own summary rat
MiniMax H3: An Open Model Breaking the Boundaries ...minimax.io · supportingAIH3MultimodalVideo Generation Today, we're launching MiniMax H3, a general-purpose multimodal generation model. H3 understands unified context across text, images, video, and audio, generating video with native stereo sound, up to 15 seconds at 2K resolution. Early testing shows H3 is ready for commercial content creation across a wide range of use cases, excelling at instruction following, accurate text and brand rendering, and V2V motion transfer. With precise, controllable multimodal generation and editing, H3 is built for advertising, branding, e-commerce, product design, UI/UX, gaming, and more. [...] Closed-source models have long dominated video generation, with slower iteration and a less open ecosystem than fields like large language models. To support the open-source community, accelerate compatibility with a broader range of AI hardware, and make it easier for users to build their own customized versions, we plan to open up the model weights in the coming days, subject to
MiniMax-H3 now on huggingfacereddit.com · supportingMy Model is on the second page of Huggingface!Image 73 r/LocalLLM•5mo ago ### My Model is on the second page of Huggingface! 107 upvotes ·44 comments minimax just dropped m3 weights on huggingface. 428b total but only 23b active. anyone tried running it locally yetImage 74 r/ollama•1mo ago ### minimax just dropped m3 weights on huggingface. 428b total but only 23b active. anyone tried running it locally yet Image 75: r/ollama - minimax just dropped m3 weights on huggingface. 428b total but only 23b active. anyone tried running it locally yet 327 upvotes ·51 comments Image 76: Llama Image 77: Llama Public Anyone can view, post, and comment to this community 0 0 [...] Image 27: Claude + GPT + local. Free. Image 28: Your data stays. So does your money. Image 29: Switch models mid-sentence. Free. Image 30: Every model. One app. Free. Image 31: Mage Lab. Local. Free. Your data stays with you. magelab.ai Download is enough for th
MiniMax H3 Unifies AI Video—and Fooocus Has an RCE Warning | AI Signalyoutube.com · supporting[music] AI video models usually make you choose a lane, text to video, image to video, motion reference, editing, or audio. MiniMax says its new H3 model is designed to put those jobs into one system. The company launched H3 on July 31st as a general-purpose multi-model generator. It can take text, images, video, and audio as context, then generate video with native stereo sound for up to 15 seconds at 2K resolution. That could let a creator reference the motion from one clip, the character from an image, and the voice from an audio file in a single instruction. MiniMax also says H3 can handle native multi-shot video, accurate [music] text and brand rendering, motion transfer, and generalized editing. Those are company claims, and the full technical report has not been published yet. The [...] is a verified patched release, do not import metadata from untrusted images. Keep Fooocus off the public internet. Avoid exposing its web interface and run unfamiliar files in an isolated environ
Blaine Brown on X: "Running MiniMax H3 locally for FREE is blowing my mind! It's like having an unrestricted Seedance2.0/Sora2-level model you can run locally for FREE. https://t.co/dmqXrv8z6r" / Xx.com · supporting## Post ## Post user avatar user avatar user avatar user avatar user avatar ## Log in or sign up for X See what’s happening and join the conversation ## Relevant people Avatar ## Trending now