Black Forest Labs (BFL), a German AI company, has launched FLUX 3, a new model that can generate images, 20-second video clips with audio, and even predict robot actions—all from a single system. Unlike models that combine separate parts, FLUX 3 is trained jointly across these abilities, which BFL calls 'visual intelligence.'
The model is available in four versions: FLUX 3 Video, Image, Action, and an upcoming open-source Dev version. However, initial access is limited. Only FLUX 3 Video and Action are entering a gated early access program now, with Image rolling out in weeks. Developers hoping for downloadable weights will have to wait until later this year.
Pricing and detailed benchmarks have not been announced. Preliminary tests show FLUX 3 outperforming rivals like Luma Ray 3.2 and Runway Gen-4.5 in preference tests, but BFL cautions these results are from a pre-release version. The model can generate up to 20 seconds of video with native audio and supports text-to-video, image-to-video, and keyframe transitions.
BFL also highlights FLUX 3's potential for robotics. In partnership with Mimic Robotics, the model can be fine-tuned for specific tasks with as little as 30 minutes of robot data, far less than traditional methods. This is possible because FLUX 3 already understands physical dynamics from its video training.
The company, valued at $3.25 billion and backed by investors like a16z and NVIDIA, plans to release open-weight versions later. For now, enterprise buyers must evaluate based on limited information.