Particle.news
Download on the App Store

Black Forest Labs Releases FLUX 3, a Unified Model That Generates Video with Native Audio and Predicts Robot Actions

The multimodal system trains on images, video and audio to learn physical motion so partners can test synchronized content generation and robotics before wider release.

Overview

  • Black Forest Labs opened FLUX 3 in Early Access on Thursday, offering text-to-video clips up to 20 seconds with native, synchronized audio plus early action-prediction tools available to partners and via APIs.
  • FLUX 3 is built on BFL’s Self-Flow approach that trains a single model on images, video and audio to capture motion, contact and sound so the same backbone can power image synthesis, video+audio generation, and action prediction.
  • A specialized variant, FLUX-mimic, was developed with mimic robotics and is being tested by Audi for high-dexterity tasks such as fitting flexible door seals on production lines, with the partners reporting much faster fine-tuning from limited demonstration data.
  • BFL published preliminary human-preference evaluations showing FLUX 3 favored over several competitors in early head-to-head tests, but the company describes these results as early and subject to further validation during the access phase.
  • The company closed a $300 million Series B that values BFL at $3.25 billion and outlined a phased rollout that will expose Video and Action to partners now, add Image generation in the coming weeks, and release an open-weight developer backbone later in 2026.