Background
<cite index="20-2">Black Forest Labs was founded in August 2024 in Freiburg im Breisgau, Germany, by Robin Rombach, Andreas Blattmann, Patrick Esser, and Dominik Lorenz — all principal researchers behind the latent diffusion technology that powered Stable Diffusion.</cite> <cite index="25-4">The company has raised significant capital across multiple rounds, culminating in a $300 million Series B financing at a $3.25 billion valuation in February 2026, backed by investors including Andreessen Horowitz, General Catalyst, and Salesforce Ventures.</cite>
The Announcement
<cite index="7-2,7-3">On July 23, 2026, Black Forest Labs introduced FLUX 3, its new multimodal frontier model that jointly learns from images, video, and audio within a unified architecture, and can also be extended to predict actions.</cite> <cite index="4-6">This marks the company's first-ever video generation model — until now, Black Forest Labs had shipped image-only models, from the original FLUX.1 through the FLUX.2 line.</cite>
Architecture: Self-Flow
<cite index="12-12,12-13,12-14">FLUX 3 builds on Self-Flow, Black Forest Labs' approach for aligning multimodal generation and understanding within the same underlying architecture. Based on this approach, the company scaled up compute and data resources to train FLUX 3 across video, images, and audio simultaneously. Self-Flow is a training technique that unifies representation learning and generation in one pass, without relying on frozen external encoders such as CLIP or DINO.</cite> <cite index="7-8">Testing showed that generative video generation and action prediction do not require separate foundations, as the same underlying architecture could be extended to action prediction without sacrificing capabilities learned from videos.</cite>
Capabilities
<cite index="6-11">The model's most immediate feature is video generation: FLUX 3 can produce 20-second video clips with synchronized audio, where sound effects, dialogue, and ambient noise align with the visuals.</cite> <cite index="14-3">It supports text-to-video, image-to-video, video-to-video, and keyframe-to-video, with multi-shot sequences chained agentically.</cite> A companion product, FLUX-mimic, extends the same backbone to physical AI: <cite index="3-17">it is being tested with Mimic Robotics on real production tasks at Audi's manufacturing facilities.</cite>
Benchmark Claims
<cite index="10-6,10-7">Black Forest Labs has published several benchmark comparisons, qualified as preliminary, with full methodology to be published later during broader general availability. In early head-to-head preference testing on 10-second, 720p text-to-video clips with audio, the company says FLUX 3 was preferred over Luma Ray 3.2 in 93% of comparisons, Runway Gen-4.5 in 77%, Grok Imagine Video in 69%, Kling v3 Pro in 60%, Happy Horse v1 in 59%, Happy Horse 1.1 in 57%, and both Seedance 2.0 and Google's Gemini Omni Flash in 52%.</cite>
Those figures require context. <cite index="9-5,9-6">Black Forest Labs has not published the evaluation methodology, sample size, rater selection criteria, or the specific prompt set used; the clips used in comparisons were 10-second generations, not the full 20-second maximum.</cite> <cite index="9-7">Independent benchmarks from researchers or publication-grade review sites have not yet been conducted.</cite>
Availability and Pricing
<cite index="4-9">FLUX 3 Video and FLUX 3 Action entered a gated early-access program on launch day — anyone can apply, but Black Forest Labs must approve each applicant.</cite> <cite index="1-8">Image generation is to follow in the coming weeks, and an open-weight FLUX 3 Dev backbone is planned for later.</cite> <cite index="12-17">No pricing has been disclosed for any tier as of this writing.</cite>
Partnerships and Ecosystem
<cite index="7-4">The FLUX family already powers generative features inside leading platforms including Adobe Photoshop, Picsart, and Nous Research's Hermes Agent.</cite> <cite index="14-8">Canva, Burda, Magnific (formerly Freepik), Krea, and Picsart are testing FLUX 3, per Black Forest Labs.</cite> The lab's dual-track strategy — gated commercial access paired with a planned open-weight release — mirrors the approach that made earlier FLUX models among the most-downloaded image generation systems on Hugging Face.