READ. SCROLL. LISTEN.

Original briefings. Zero spin.

Every story is an original briefing written from 60+ sources across the spectrum — sources linked so you can verify it yourself.

← Back to headlines

Black Forest Labs Launches FLUX 3, a Video-Audio-Image Model, With Almost Nobody Able to Use It Yet

Black Forest Labs Launches FLUX 3, a Video-Audio-Image Model, With Almost Nobody Able to Use It Yet
The German AI lab announced FLUX 3 on July 23, 2026, a model that can generate 20-second video clips with matching audio from a single prompt. It's an early access waitlist right now, no pricing, no public API, no benchmarks anyone outside the company can check.

Black Forest Labs, based in Freiburg, Germany, announced FLUX 3 on Thursday, July 23, 2026. The pitch: one model trained jointly on images, video, audio and robotic actions, rather than separate models stitched together behind a shared interface.

CEO Robin Rombach framed it as a bet on multimodal training itself. "Joint training within one unified architecture is what will get us there, because each training modality strengthens the others," Rombach said in the company's announcement, distributed via GlobeNewswire. His argument: audio carries timing and physical cues that video alone misses, and language carries goals and instructions pixels can't express.

What FLUX 3 actually does

According to the-decoder, FLUX 3 generates video clips up to 20 seconds long with native audio, a first for the company. It supports text-to-video, image-to-video, video-to-video, keyframe transitions, multilingual dialogue and what BFL calls agent-driven links between clips for longer sequences.

The company says the model is particularly strong on human facial expressions and syncing sound to physical events on screen.

The numbers BFL is putting out there, and why you can't check them

BFL published comparison results against rival video models, all self-reported. Per the-decoder, in 10-second clip tests at 720p, BFL says FLUX 3 was preferred over Luma Ray 3.2 in 93 percent of comparisons, over Runway Gen-4.5 in 77 percent, and over Grok Imagine Video in 69 percent.

Against tougher competition the gap shrinks fast: 60 percent against Kling v3 Pro, 59 and 57 percent against two Happy Horse versions, and a coin-flip 52 percent against both Seedance 2.0 and Gemini Omni Flash.

BFL itself says these results are preliminary. No independent lab has verified them. VentureBeat noted the company hasn't disclosed rater counts, sample sizes, or full methodology, meaning there's no way for outside researchers or enterprise buyers to reproduce these comparisons right now. That's a real gap for anyone trying to evaluate whether FLUX 3 is actually competitive with Seedance, a model VentureBeat notes has already been used in Hollywood productions.

Four products, almost none of them available today

FLUX 3 is being split into four lines: FLUX 3 Video, FLUX 3 Image, FLUX 3 Action, and a promised open-weight FLUX 3 Dev. Only Video and Action are opening to a gated "Early Access" program starting today, and BFL has to approve every applicant. There's no public API access from BFL or any partner yet.

FLUX 3 Image is slated to roll out "in the coming weeks," according to both VentureBeat and the-decoder. General availability timing beyond that hasn't been specified.

VentureBeat points out this staggered, gated rollout mirrors what Anthropic and OpenAI have done with recent model launches, though those companies cited security concerns and government requests as justification. BFL hasn't offered a comparable explanation for withholding public access.

No pricing, no open weights, not yet

BFL has not announced pricing or any service-level commitments for enterprise customers, according to VentureBeat. That means no business can currently calculate what running FLUX 3 at scale would actually cost.

The bigger letdown for BFL's existing developer base: no downloadable weights. Previous FLUX releases built their reputation partly on open-weight versions arriving alongside or shortly after a major launch. This time, BFL says an open-weight FLUX 3 Dev is coming "later this year," described in the company's technical materials as covering video, audio, image and action prediction, a broader scope than any prior Dev release, which covered images only. But it's last in line, not first.

The robotics angle

Beyond content generation, BFL says the same architecture extends to predicting physical actions. The company built FLUX-mimic with Mimic Robotics, a video-action model reportedly being tested on production tasks at Audi, according to the-decoder. That's a concrete industrial application, not just a demo, though BFL hasn't disclosed what tasks or how extensively it's deployed on Audi's floor.

What's unresolved

The core question is whether FLUX 3's self-reported preference rates hold up once outside researchers or paying customers get real access. BFL says fuller benchmark results and methodology are coming later, per VentureBeat, but gave no firm date.

Until FLUX 3 Image ships in the coming weeks and the gated Early Access program starts letting people in, everything about this launch—capability claims, competitive standing, cost—remains something enterprises will have to take on Black Forest Labs' word.

Sources used for this briefing

This briefing was written by UBH's AI agent — these are the reporting inputs it draws on, linked so you can verify.

center
VentureBeatBlack Forest Labs launches FLUX 3 capable of generating images and 20-second video with audio — but in limited release to start
center-left
markets.businessinsiderBlack Forest Labs Unveils FLUX 3, A New Multimodal Frontier Model For Visual Intelligence
unknown
spiditsBlack Forest Labs Launches FLUX 3 Capable of Generating Images and 20-second Video with Audio - but in Limited Release to Start | AI Timeline | SPIDITS
unknown
the-decoderFlux 3 generates videos with native audio up to 20 seconds long, a first for Black Forest Labs