MiniMax 发布 Music 3 开放权重音乐生成模型,支持最长五分钟歌曲
MiniMax Music 3 已公开发布,可根据歌词和音乐描述生成包含人声与编曲的完整歌曲,最长约五分钟。模型权重已在 Hugging Face 提供,并支持多种本地推理工具。
MiniMax Music 3 是一款开放权重的文本生成音乐模型。用户提供歌词和详细的音乐描述后,模型能够生成结构完整、带有演唱与编曲的歌曲,最长时长约为五分钟,输出格式为 32 kHz、16-bit 立体声 WAV。
官方介绍称,该模型采用由 8B Global LLM 和 0.6B Local LLM 组成的混合架构:前者负责歌曲的长期结构与语义连贯性,后者负责帧级别的声学细节。官方还强调,模型通过连续隐藏状态、Flow Matching 和 Flow-VAE 等组件改善人声表现与长段落稳定性。
模型权重已发布至 Hugging Face,并提供 SGLang、diffusers 和 ComfyUI 等部署路径。官方资料将云端 API 描述为即将推出,因此目前更明确的使用方式是下载权重并进行本地推理。
来源证据
MiniMax open-sources Music 3, a Qwen3-based song generator | AI Weeklyaiweekly.co · supporting## Shared on Bluesky by 2 AI experts Sung Kim @sungkim.bsky.social: 🎵MiniMax-Music3 (open-weight) Next-Generation Open-Weights Production-Ready & Versatile Music Model Model: huggingface.co/MiniMaxAI/Mi... … → M MiniMax (official): huggingface: github: 魔搭: → Originally reported by github.com Read the original article → Original headline: GitHub - MiniMax-AI/MiniMax-Music3 Free AI alerts in your inbox Breaking AI news 3x/week. 50,000+ subscribers. We use essential cookies to keep the site working (login, form security). With your permission, we also use analytics cookies to understand how you use the site. Privacy policy [...] MiniMax has open-sourced MiniMax Music 3, a text-to-song model that generates complete tracks up to five minutes long from lyrics plus a music description. The output is 32 kHz, 16-bit stereo WAV, and lyrics can carry section tags like [Intro], [Verse], [Pre-Chorus], [Chorus], [Bridge], [Instrumental], [Solo] and [Outro] to shape structure.
MiniMaxAI/MiniMax-Music3huggingface.co · supporting# MiniMax Music 3 MiniMax Music 3 is a high-performance music generation model for creating complete songs up to five minutes long. Conditioned on lyrics and a detailed music description, it generates structurally coherent songs with expressive vocals, evolving arrangements, and stable long-form audio quality. MiniMax Music 3 combines an 8B Global LLM for long-range musical structure, a 0.6B Local LLM for frame-level acoustic detail, and a continuous hidden-state synthesis system based on Flow Matching and Flow-VAE. The model produces 32 kHz, 16-bit stereo WAV audio. ## Demo Explore music generation examples on the MiniMax Music 3 Demo. ## Complete Songs with Long-Range Coherence [...] ### Download the Model `hf download MiniMaxAI/MiniMax-Music3 --local-dir /path/to/minimax_ttm` We recommend the following inference frameworks to serve the model: SGLang - see cookbook diffusers - see diffusers docs ComfyUI see comfyUI tutorials ### Serve with SGLang-Omni `sgl-omni serve --mo
What Is MiniMax Music 3? Open-Weights Music Modelkie.ai · supportingMiniMax Music 3 is an open-weights AI music generation model from MiniMax that turns lyrics and a text music description into a complete, produced song of up to five minutes in a single generation. It outputs 32 kHz, 16-bit stereo audio and went live on August 13, 2026, with weights on Hugging Face and native support in ComfyUI. MiniMax's official announcement called it a "Next-Generation Open-Weights Production-Ready & Versatile Music Model," with a cloud API listed as coming soon. ## Key Takeaways [...] It targets the parts of music generation that short prompts struggle with: understanding a creator's expressive intent, holding that intent across an entire song, rendering instruments with physical realism, and producing vocals that sound performed rather than synthesized. The status is released, not leaked. MiniMax's official blog dated the launch August 13, 2026, and the weights are live on the `MiniMaxAI/MiniMax-Music3` Hugging Face repository, with a ComfyUI-optimized distributi
MiniMax Music 3: The Open-Weight AI Music Model ...mindstudio.ai · supportingMiniMax Music 3 is an open-weight AI music generator with downloadable weights on Hugging Face, built around a Qwen3-8B language model paired with a diffusion-based audio pipeline. It runs locally with a claimed minimum of 8GB of VRAM using layer streaming, though 20 to 24GB is recommended for smooth full-precision inference. The model already has day-one support in ComfyUI, and community fine-tunes started appearing on Hugging Face almost immediately after release. Its license permits commercial use with attribution to MiniMax Music 3, and only requires a separate agreement with MiniMax once a project earns more than $20 million. [...] ## What is MiniMax Music 3? MiniMax Music 3 is an open-weight text-to-music model released by MiniMax, the same company behind the open-weight video generator Hailuo (also referred to as H3 in some coverage) and the Seance 2.5 model. Music 3 takes a text prompt, and in some workflows a style or lyric reference, and generates a full song with vocals
MiniMax Music 3.0: Next-Generation Open-Weights, Production-Ready & Versatile Music Model - MiniMax Research | MiniMaxminimax.io · supportingAt the model level, Music 3.0 uses a global–local collaborative Hybrid-LM. The 8B Global LLM predicts core semantic and structural tokens frame by frame and maintains full-song context. The 0.6B Local LLM predicts acoustic tokens along the depth axis within each frame, supplying local sonic detail. This division of responsibilities maintains temporal stability across songs of up to five minutes while preserving rich variation within individual sections. [...] Music 3.0 introduces a new audio-rendering system designed to produce more natural, studio-quality vocal performances. The Structured Caption describes vocal timbre, delivery, techniques such as breathiness and falsetto, harmony arrangement, and effects such as delay and Auto-Tune in fine detail. Continuous hidden states fused from the global and local language models carry this performance information into the flow-matching and Flow-VAE generation process. Together, these mechanisms reduce the high-frequency digital artifacts com
🎵MiniMax-Music3 (open-weight) Next-Generation Open-Weights Production-Ready & Versatile Music Model Model: https://huggingface.co/MiniMaxAI/MiniMax-Music3 Repo: https://github.com/MiniMax-AI/MiniMax-Music3threads.com · supportingRelated threads digitalmatters.me's profile picture digitalmatters.me AI\_Music\_Generation MiniMax Music 3.0 Puts a Song Model on Your Own Hardware. Read the License First. MiniMax Music 3 is an open-weight music generation model from the Chinese lab MiniMax, published on August 13, 2026.... digitalmatters.me/artif… AI\_Music\_Generation #Generative\_AI #MiniMax #Open\_Weights MiniMax Music 3.0 Puts a Song Model on Your Own Hardware. Read the License First. digitalmatters.me MiniMax Music 3.0 Puts a Song Model on Your Own Hardware. Read the License First. ojoo.ai's profile picture ojoo.ai 💽 Chinese MinMax has Released Music 3, An Open Source Music Generator It Creates Songs of up to 5 Minutes and Runs Locally - even on consumer hardware. [...] The Full Model fits on a GPU with 24 GB of VRAM, but it can run with as little as 8 GB. ➡ You can Try it for FREE on Hugging Face ➡ huggingface.co/space… 1 1 Log in to see more replies. Log in Log in or sign up for ThreadsSe