公告
数据公告

QQ群和tg群已经启用,欢迎加入。公开信息来源均审核后发布;请结合来源、库存和更新时间判断。

社群与联系Telegram 群点击加入Telegram 频道点击订阅联系我们tgAIPricedb交流群979789483
返回资讯列表
product

MiniMax H3 团队公布开源路线图:计划推出 Regenerate-2K 与统一图像模型

MiniMax 在 Reddit AMA 总结中介绍了 H3 的后续计划,包括推进 Apache-2.0 许可、发布技术报告,以及开源 H3-Regenerate-2K 和统一图像生成与编辑模型。

86% VERIFIED

MiniMax 表示,H3 的 Apache-2.0 许可方案仍在讨论中,团队还计划发布更完整的技术报告。H3-Regenerate-2K 被描述为专用的潜空间 DiT 重生成模型,而不是简单的基础模型重跑或像素级放大器;目前团队仍在优化其效率和画质,目标是支持本地运行。

在架构和推理方面,团队介绍了类似 MoBA 的稀疏注意力机制,其参考实现将优先保证质量稳定。已经发布的检查点采用了 CFG 蒸馏,4 步和 8 步推理版本仍处于评估阶段,社区开发的 Turbo LoRA 可作为现阶段的加速方案。

团队还确认,正在对一个源自 H3 的统一文本生成图像和通用图像编辑模型进行后训练优化。此外,Ref2VA 支持通过前一段视频进行续接,社区已经展示了将工作流延长至约 60 秒的用法。上述功能和模型均应视为路线图或开发中项目,不能等同于已经全面发布。

来源证据

MiniMax-H3/README.md at main · MiniMax-AI/MiniMax-H3 · GitHubgithub.com · supporting

## Model Architecture ### H3-Context-IR H3-Context-IR is a hosted preprocessing and orchestration system designed for free-form multimodal inputs. It interprets the relationships among text, images, audio, and reference videos, as well as how these materials relate to the intended generation output. Its internal workflow includes instruction parsing, cross-modal association, temporal understanding, and complex logical reasoning. H3-Context-IR serializes its understanding of the context into a structured representation accepted by H3-Base. Without deviating from the user’s original intent, it may also supplement missing or underspecified semantic details where appropriate. [...] ## System Overview MiniMax H3 is a general-purpose, omni-modal generative system. It supports unified understanding of multimodal contexts composed of text, images, video, and audio, and can generate video with native stereo audio at resolutions up to 2K and durations of up to 15 seconds. Thanks to its task

MiniMaxAI/MiniMax-H3huggingface.co · supporting

# MiniMax H3 ## System Overview MiniMax H3 is a general-purpose, omni-modal generative system. It supports unified understanding of multimodal contexts composed of text, images, video, and audio, and can generate video with native stereo audio at resolutions up to 2K and durations of up to 15 seconds. Thanks to its task-generalization-oriented system design, H3 already possesses broad multimodal context understanding and generation capabilities at the pre-training stage, enabling outstanding performance in following complex multimodal instructions. H3 supports the following input and output specifications: [...] ## Online API Use MiniMax-H3 directly via API. Global: platform.minimax.io | CN: platform.minimaxi.com ## Online App Use MiniMax-H3 directly via App. WebApp Global: hailuoai.video | CN: hailuoai.com Desktop Global: hub.minimax.io | CN: hub.minimaxi.com ## Model Architecture ### H3-Context-IR H3-Context-IR is a hosted preprocessing and orchestration system designed

AMA: MiniMax H3 Team — Ask us anything about our open ...reddit.com · supporting

Skip to main contentAMA: MiniMax H3 Team — Ask us anything about our open video generation model, training, and future plans : r/StableDiffusion Open menu Open navigation u/Affectionate-War8374 -> Luigi (H3 Researcher) u/MM_Nero_H3 -> Nero (H3 Researcher) u/Kiro_Song -> Kiro (H3 Researcher) u/New_Estimate9277 -> Reynor (H3 system engineer) u/ryan85127704 - >Ryanlee (Head of Devrel) We are the MiniMax team behind MiniMax-H3. We’re here to answer your questions, including: Model architecture and training Video generation capabilities Image-to-video and reference-based generation Inference and optimization Future plans Ask us anything — we’d love to hear your feedback and discuss with the community! Coming Up Live AMA Finished·9 hr. ago Share [...] This is really terrifying - Minimax H3Image 32 r/StableDiffusion•5d ago ### This is really terrifying - Minimax H3 Image 33: r/StableDiffusion - This is really terrifying - Minimax H30:07 62

MiniMax H3 Open Weights Exclude US, EU, UK, and Korea From ...techtimes.com · supporting

What did not ship: H3-Context-IR and H3-Regenerate-2K. MiniMax's full system has three layers. H3-Context-IR is the hosted preprocessing and orchestration layer that parses free-form multimodal inputs — images, video clips, audio tracks, text — into a structured intermediate representation the generator can follow. H3-Regenerate-2K is the in-context regeneration pass that takes the 768p base output and reprocesses it at 2K resolution, using the original multimodal context to recover fine details that standard upscaling would lose. Both modules remain hosted API services. For the flagship advertised use case — native 2K generation — the local weights are the middle step of a workflow that still requires two MiniMax API calls. Those API calls route through infrastructure subject to China's

The Design Choices Behind Native 2K Multimodal Videoyoutube.com · supporting

NEXT VIDEO — ARCHITECTURE DEEP DIVE: MiniMax scheduled the H3 model release for August 3, 2026 at 00:00 China Standard Time. After the configuration, weight index, model code, and tensor files become publicly accessible and can be verified, the next video will reverse-engineer the real H3 architecture: layer structure, parameter distribution, attention design, VAE latent geometry, and inference requirements. Publication will follow verification of the released files. updated: [...] instead asks its conditional base model to regenerate the high-resolution output from both Y low and the original multimodal context. Formally, compare S phi of Y low with P theta of Y high given Y low in context. The second expression keeps prompt references and other conditioning evidence load-bearing while reusing generation capability already learned by the base model. Minimax says contextual regeneration can recover small text and fine detail that context-free super resolution must guess, but i

MiniMax Releases MiniMax H3: An Omni-Modal Video Model That ...marktechpost.com · supporting

JetBrains Open-Sources KotlinLLM ### JetBrains Open-Sources KotlinLLM: Smart Macros That Generate Kotlin Source Code at Runtime and Hot-Reload It Through JDI Nous Research Ships Three Integration Paths for Hermes Agent and Buzz ### Nous Research Ships Three Integration Paths for Hermes Agent and Buzz, Block’s Open Source Nostr Workspace for Humans and Agents PolyAI Releases Dialog-RSN-1 ### PolyAI Releases Dialog-RSN-1: An Audio-Native Dialog Model That Fuses Turn-Taking, Speech Recognition, Function Calling, And Response Building a Policy-Governed Multi-Agent Financial Research Workflow with Omnigent ### Building a Policy-Governed Multi-Agent Financial Research Workflow with Omnigent []( Discord Linkedin Reddit X [...] ## Key Takeaways H3 unifies text, image, video, and audio into one generation model — 2K, 4–15s, native stereo. Open weights are promised “in the coming days,” not shipped; the API is the only path today. H3-VAE’s 4× effective sequence-length ga