flux-3 — AI Video Model

black-forest-labs

black-forest-labs

flux-3

Commercial usevideo

black-forest-labs/flux-3-video is Black Forest Labs' advanced multimodal video generation model, designed to create high-quality video with native synchronized audio. It can generate videos from text prompts, animate images, and transform existing video content while maintaining strong visual consistency, realistic motion, and natural interactions between objects and characters. FLUX 3 Video is built around a unified multimodal architecture that jointly understands images, video, and audio, allowing it to generate video and sound together rather than treating audio as a separate post-production step. It can produce clips of up to 20 seconds in a single generation and is particularly strong at facial expressions, physical interactions, and connecting sounds with events in the scene.

Model Type
coin

Pricing: 17.5 joules / second

Input

aspect ratio

Aspect ratio of the generated video. 'auto' picks a ratio from your prompt and any inputs.

duration
Length of the generated clip in seconds. 'auto' lets the model pick to fit the content. Three or more images need an explicit duration.
520
OutputVIDEO

Output will be displayed here

Fill in the inputs on the left and hit RUN to preview the results.

Examples

Explore different use cases and parameter configurations

VIDEO
VIDEO
VIDEO