Notice
数据公告

QQ群和tg群已经启用,欢迎加入。公开信息来源均审核后发布;请结合来源、库存和更新时间判断。

Community & contactTelegram 群点击加入Telegram 频道点击订阅联系我们tgAIPricedb交流群979789483
Back to news
Products

Black Forest Labs Launches FLUX 3 Video With 20-Second Clips and Native Audio

Black Forest Labs has introduced FLUX 3 Video in early access, with text-to-video, image-to-video, keyframes, multilingual dialogue, and native audio. Individual generations can run for up to 20 seconds.

87% VERIFIED

Black Forest Labs has launched FLUX 3 Video as a multimodal generation model for video, images, and audio. It can create clips from text or image inputs and supports keyframe transitions, video continuation, multi-shot sequencing, and dialogue in multiple languages. The company says a single generation can last up to 20 seconds, with audio generated alongside the visuals.

The model also includes Draft Mode, which provides a faster, lower-cost preview for testing ideas before rendering a final version. Resolution and pricing details are not fully consistent across the available reports: official material references 720p and 1080p, while some coverage describes 720p as the initial launch resolution. Plans for 2K, 4K, and open weights are described as forthcoming, and the reported $0.06 draft price has not been independently confirmed.

Source evidence

FLUX 3 Video, Part 1: Generationbfl.ai · supporting

### FLUX 3 Video - Generation Capabilities We make FLUX 3 Video available to a general audience today. In its initial form, FLUX 3 Video can create video clips of up to 20 seconds length with native audio. We are releasing our model at HD (720p) and Full HD (1080p) resolutions, and we’re providing the following set of capabilities today: Text-to-Video: Describe a scene in simple language or using a detailed prompt. FLUX 3 follows complex instructions while generating natural movements, scene logic, and audio. Image-to-Video and Keyframes: Start with an image, specify an end frame, or set multiple keyframes in a clip. FLUX 3 Video connects these in sequence while following the intended visual language. [...] Draft Mode: Enables the exploration of creative directions easily. A draft generation returns a fast preview of your prompt at a fraction of the cost, so you can iterate on ideas instead of waiting for a full high-quality generation every time. When a draft is satisfactory, FLU

FLUX 3 多模态 AI 模型登场:支持单次生成 20 秒多样化视频_腾讯新闻news.qq.com · supporting

# FLUX 3 多模态 AI 模型登场:支持单次生成 20 秒多样化视频 头像 IT之家 2026-07-24 14:55发布于湖北IT之家官方账号 问AI · 与机器人合作将解锁哪些行为预测新场景? IT之家 7 月 24 日消息,Black Forest Labs 昨日(7 月 23 日)发布博文,宣布以 Early Access 方式,上线推出 Flux 3 多模态基础模型,采用统一架构联合学习图像、视频和音频。 在训练方面,IT之家援引博文介绍,该模型基于 Self-Flow(自监督流匹配)学习框架扩展,同步训练视频、图像和音频,在 FLUX.1 和 FLUX.2 系列基础上,扩大至多模态生成与理解任务。 FLUX 3 可在单次生成中输出最长 20 秒的视频,并附带原生音频。官方表示该模型支持文生视频、图生视频、视频生视频、输入视频与音频续写、关键帧转视频、多语言对话,以及多镜头串联。 性能方面,在带声音的 10 秒、720p 视频人工评测方面,结果显示,FLUX 3 对比 Grok Imagine Video 的胜率为 69%,对比 Seedance 2.0 的胜率为 52%,对比 Gemini Omni Flash 的胜率也为 52%。 在机器人方向,Black Forest Labs 正与 Mimic Robotics 合作,研究将 FLUX 3 用作机器人行为预测模型。 图像能力方面,官方称 FLUX 3 可完成图像生成与编辑,覆盖多种风格、宽高比和分辨率。

black-forest-labs/flux-3 | Run with an API on Replicatereplicate.com · supporting

## Settings `aspect_ratio` — `auto` (default), or one of `21:9`, `2:1`, `16:9`, `4:3`, `1:1`, `3:4`, `9:16`. `auto` picks a ratio from your prompt and inputs. `resolution` — `720p` (default) or `1080p`. `duration` — `auto` (default) or a whole number of seconds from 5 to 20. `generate_audio` — on by default. Turn it off for a silent clip. `draft` — generate a fast, low-cost 720p preview. Handy for iterating on a prompt before a full-quality run. `safety_tolerance` — moderation tolerance from 0 (strictest) to 4. Requests with image or video inputs are limited to 2. ## About FLUX 3 is built on Black Forest Labs’ Self-Flow architecture, a single multimodal flow-matching model trained across images, video, and audio. Model created Model updated

Flux 3 generates videos with native audio up to 20 seconds long, a first for Black Forest Labsthe-decoder.com · supporting

Flux 3 can now generate videos with native audio for the first time, with clips up to 20 seconds long. It supports text-to-video, image-to-video, video-to-video, keyframe-based transitions, multilingual dialogue, and agent-driven links between clips for longer multi-shot sequences. BFL says the model is especially good at human facial expressions and matching sounds to physical events. Ad DEC\_D\_Incontent-1 [...] ### AI News Without the Hype – Curated by Humans Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section. Subscribe now Source: BFL Blog wpDiscuz [...] BFL describes Flux 3 as a step toward "real-world visual intelligence," which it defines as models that can "perceive, predict, and act across physical and digital environments." The company is part of a broader push to build so-called world models. No single modality captures reality in full,

Black Forest Labs launches FLUX 3 capable of generating images and ...venturebeat.com · supporting

Every video output comes with native audio. For comparison, HappyHorse 1.0 tops out at 15 seconds of 1080p with synchronized audio — though BFL has not stated what resolution its 20-second clips run at, and its published evaluations were conducted at 720p. Still, a 20-second long clip from a single prompt is among the longest yet achieved, matching OpenAI's discontinued Sora model. The capability list BFL published covers: [...] | | | | | | | | --- --- --- | Model | Max single-generation duration | Max resolution | Key constraints | Price per 10-second clip (720p) | Price per 10-second clip (1080p) | Price per 10-second clip (4K) | | FLUX 3 Video | 20 seconds | Not stated; evaluations run at 720p | Early access; no published SLA or pricing | Not announced | Not announced | Not announced | | HappyHorse 1.1 | 15 seconds | 1080p | No 4K; closed weights | Not published (v1.0 reseller rate is ~$1.82) | Not published (v1.0 reseller rate is ~$3.12) | n/a | | Veo 3.1 | Per-second b

Flux 3 Is Here: What Black Forest Labs' New AI Video Model Can Domindstudio.ai · supporting

## What is Flux 3? Flux 3 is the video generation model from Black Forest Labs, released as the follow-up to the image-focused Flux 1 that launched back in August 2024. It’s a multimodal model, meaning it accepts text prompts, image references (up to 10 of them), audio, and video inputs for editing or remixing existing footage. It generates clips up to 20 seconds long, supports standard aspect ratios from 9:16 to 21:9, and produces native audio alongside video. It’s currently in early access via Discord, with resolution capped at 720p at launch and 1080p expected to roll out shortly after. ## TL;DR [...] Flux 3 launched in early access more than a year after Black Forest Labs first teased a video model alongside Flux 1, making this a long-awaited release for the company. Clip length hits 20 seconds, a meaningful jump from the 10-second ceiling that’s become the norm for most competing video models. The model is multimodal, accepting up to 10 image references plus audio and video in