FLUX 3
FLUX 3 is Black Forest Labs' unified multimodal foundation model that jointly learns from images, video, and audio in one architecture, generating video up to 20 seconds with native synchronized audio. It supports text-to-video, image-to-video, video-to-video, keyframe-to-video, multilingual dialogue, and agentic multi-shot chaining, with early access available as of July 23, 2026, though image generation access was still pending as of mid-August 2026.