A company from Germany's Black Forest region is building one of the most ambitious AI models in the world right now. The name is fitting: Black Forest Labs, founded by former Stability AI people, is actually headquartered in Freiburg — and it just unveiled FLUX 3, a model that generates images, videos, sound, and even robot movements from a single set of weights.
Sounds like marketing talk. It's technically remarkable, though.
One model, many senses
Previous AI models have been specialists: one for images, one for video, one for language. FLUX 3 learns image, video, and audio together in the same training run — and extends the same idea to predicting robot actions. An AI agent tasked with grasping an object draws on the same "imagination" as the model generating a video on the side.
Concretely, FLUX 3 can currently:
- Generate and edit images, including readable text inside them — something many competitors still can't do reliably
- Produce videos up to 20 seconds long, with sound that actually matches what's happening on screen instead of being bolted on afterward
- Keep the same character consistent across multiple video clips
Who gets to play?
Here's the catch: FLUX 3 Video currently runs in gated early access only. Anyone who wants in has to apply on the Black Forest Labs website and hope to get approved. No price has been announced as of now. FLUX 3 Image is supposed to follow in the coming weeks, also with limited access at first.
The good news for anyone without a developer team behind them: Black Forest Labs already proved with the FLUX.1 series that open weights are part of the company philosophy. An open version of FLUX 3 is planned for later this year — at that point, the model can be self-hosted, no waitlist required.
What about robots?
The robotics angle sounds like a marketing add-on at first, but it's actually the more interesting part of the announcement. A model that has learned how objects plausibly move in a video can use that same knowledge to predict how a robot arm should grasp an object. Instead of collecting separate training data for image AI and robot AI, Black Forest Labs uses one shared foundation. Whether this works as well in practice as the announcement claims remains to be seen — no solid, independent tests exist as of now.
What does this mean for you?
You can't try FLUX 3 for free right now — the waitlist decides that. What you can use for free today is the predecessor, FLUX.1. Its open variants run on platforms like Hugging Face or directly in tools like ComfyUI, if you bring a bit of technical curiosity.
For developers or for everyone? Right now, clearly for developers and companies with API access. That changes only once the open version arrives — that's when it gets interesting for tinkerers with their own graphics card, too.
How does it stack up against the competition? OpenAI's Sora and Google's Veo also generate video with sound, but they run as pure cloud services with no open weights in sight. That's exactly the difference Black Forest Labs claims for itself: once you get FLUX 3 open, you can run it on your own hardware — no subscription, and your prompts don't end up anywhere else.
The more interesting part isn't the access model anyway, it's the consequence: a model that thinks image, sound, and motion together makes convincing fake videos easier to produce — and harder to spot. That's no reason to panic, but a good reason to keep deepfake detection tools like SynthID and C2PA in the back of your mind instead of relying on common sense alone.
Until the open version arrives, FLUX 3 remains mostly one thing: a promise. A very German one, too — announced soberly, without a big show, but with real technical substance behind it.
