MiniMax представила генерацию видео H3 на мероприятии Magnific в Сан-Франциско
По словам MiniMax, на мероприятии для авторов в Сан-Франциско компания вместе с Magnific показала H3, провела обзор модели и организовала живую демонстрацию.
MiniMax сообщила, что мероприятие в Сан-Франциско собрало больше посетителей, чем вмещала площадка. Архитектор решений MiniMax Итан Вэй провёл обзор H3, а автор Энрике Лопес продемонстрировал работу модели непосредственно в Magnific.
В программе также была дискуссия о будущем генеративного видео и творческих рабочих процессов с участием Клэр Сюэ. Согласно официальным материалам MiniMax, H3 — это модель общего назначения с открытыми весами для мультимодальной генерации. Она принимает текст, изображения, видео и аудио и создаёт видео с нативным стереозвуком, разрешением до 2K и длительностью до 15 секунд.
MiniMax также объявила о скидке 50% на генерацию H3 в разрешении 2K через Magnific до 1 сентября. Поскольку предложение ограничено по сроку, его актуальность следует проверять непосредственно на сайте Magnific.
Источники
GitHub - MiniMax-AI/MiniMax-H3github.com · supporting## System Overview MiniMax H3 is a general-purpose, omni-modal generative system. It supports unified understanding of multimodal contexts composed of text, images, video, and audio, and can generate video with native stereo audio at resolutions up to 2K and durations of up to 15 seconds. Thanks to its task-generalization-oriented system design, H3 already possesses broad multimodal context understanding and generation capabilities at the pre-training stage, enabling outstanding performance in following complex multimodal instructions. H3 supports the following input and output specifications:
MiniMaxAI/MiniMax-H3 - vLLM Recipesrecipes.vllm.ai · supportingvLLMvLLM/Recipes MiniMax # MiniMaxAI/MiniMax-H3 Open-weight general-purpose multimodal generation model — jointly generates 24 FPS video with native stereo audio from text, image, video, and audio references, served via vLLM-Omni 8.7 s of 1248×768 video with synchronized stereo audio in ~87 s on 4×B300 View on HuggingFace dense64B0 ctxvLLM 0.26.0+vLLM-Omninightly path — nightly wheels")omni Guide ## Overview MiniMax H3 is an open-weight, general-purpose multimodal generation model. Rather than being confined to one specialized task — generate, edit, or reference — H3 reads a multimodal context that mixes text, images, video, and audio together, interprets the creative intent as a whole, and produces coherent audio-visual output end to end. [...] The modular root service loads both DiTs by default. Capacity profiles use `--task-type fl2va` or `--task-type ref2va` and expose only that task family. H3 executes one generation request per diffusion batch today. The first request
MiniMax's Postlinkedin.com · supportingJoin Lovis Odin, Creative Engineer at fal, Ethan Wei, AI Solutions Architect ... walkthrough of how creators and developers are building with H3.
MiniMax H3: An Open Model Breaking the Boundaries Between ...minimax.io · supportingAIH3MultimodalVideo Generation Today, we're launching MiniMax H3, a general-purpose multimodal generation model. H3 understands unified context across text, images, video, and audio, generating video with native stereo sound, up to 15 seconds at 2K resolution. For a quick hands-on experience, please visit . Early testing shows H3 is ready for commercial content creation across a wide range of use cases, excelling at instruction following, accurate text and brand rendering, and V2V motion transfer. With precise, controllable multimodal generation and editing, H3 is built for advertising, branding, e-commerce, product design, UI/UX, gaming, and more. [...] ### Multimodal context understanding Real creative work requires blending complex information across modalities, pulling in images, audio, video, and more as input sources. For example, the prompt for the shot below is: "Reference the Hitchcock camera movement from Video 1, have the character in Image 2 sing, with the vocals matchin
MINIMAX H3 JUST GOT EVEN BETTER!youtube.com · supportingthat it is much faster to first generate the video at a low resolution and then upscale it to a higher resolution at the end. So like for example, if I use the normal text to video workflow with a normal prompt and I use something like 1.2 megapixel resolution and then I click run, which gives us something like this. Yo, Mr. White. Yo, Jesse. What's up? Yeah, there you go. Now right now I'm renting a 4090. So this video took around 172 seconds to generate, which is already fairly fast. And now if I use the fast text to video workflow with the same exact parameters, but this time I first generate the video at 0.2 megapixel and then upscale it to 1.2 megapixel. If now I click run, okay, it gives us something like this at the end. Yo, Mr. White. Yo, Jesse. What's up? So as you can see, first [...] you to extend an already existing video. And I have tried a lot of different workflows, a lot of different nodes, and in my testing, this one is definitely the best. Now, the way it works is ver
We went a little over capacity at @magnific SF for MiniMax ...x.com · supportingEthan Wei, AI Solutions Architect at @MiniMax_AI, gave an H3 walkthrough 🛠️ Claire Xue joined us for a fireside on where AI video + creative