Black Forest Labs (BFL) began general availability of its video generation model "FLUX 3 Video" on August 4, 2026. The model can create videos up to 20 seconds long—complete with dialogue and sound effects—from text or still images, and can also generate continuations of existing videos, all through a single API. While the model was available in early access as of the July 23 announcement, it is now accessible to anyone through the BFL API and select partners. Beyond the performance benchmark numbers, what matters for production workflows is the combination of per-second pricing and draft-based workflow support now in place.
Generating 20-Second Videos with Audio Through a Single API
FLUX 3 Video accepts durations of 5 to 20 seconds in whole-second increments, with an automatic setting available to match the content. HD is the default resolution, with FHD also supported. Audio generation alongside video is enabled by default, producing dialogue, sound effects, and ambient sound from the same request.
The API offers three generation modes. Text-to-video creates footage from written descriptions, while image-to-video allows still images to be fixed as starting or ending frames. Video continuation takes existing video and audio as input and generates subsequent motion, camera movement, and dialogue that follows. According to BFL's announcement, the existing video used as a continuation source can be up to 4 seconds long.
Up to 10 still images can be specified as keyframes. Beyond using two images for start and end points, assigning timing to each image allows them to function as a storyboard that determines intermediate framing and poses. Multiple scenes or camera angles can also be included in a single generation, reducing the need to create short cuts separately and stitch them together afterward.
BFL states that the model supports lip-syncing across multiple languages. Examples cited include English, Chinese, and Japanese, along with several European and South Asian languages. However, per-language accuracy and synchronization rates for longer dialogue have not been disclosed.
From Draft to Final: $1.20 to $10.60 for 20 Seconds
The draft mode is what makes a real difference in production workflows. Generating a low-cost HD preview first returns an encrypted draft_cache from the API. This data locks in the original mode and prompt. Since the seed and reference media are also preserved, regenerating at standard quality later maintains the same composition and motion. Users can select their preferred draft and then specify HD or FHD, reducing the risk that content changes during final rendering.
Pricing is based on per-second rates tied to output duration, and includes generated audio. Calculating the cost of producing the maximum 20 seconds from published pricing yields the following:
| Mode | HD Draft | HD Standard | FHD Standard |
|---|---|---|---|
| Text/Image-to-Video | $0.06/sec ($1.20 for 20 sec) | $0.17/sec ($3.40) | $0.29/sec ($5.80) |
| Video Continuation | $0.12/sec ($2.40 for 20 sec) | $0.41/sec ($8.20) | $0.53/sec ($10.60) |
Video continuation costs more than text/image-to-video generation, with the price gap between the two modes widening to $4.80 for 20 seconds at FHD. For workflows involving multiple draft iterations, using the $0.06 or $0.12 draft tier makes costs more predictable. Meanwhile, FHD isn't generated directly at 1080p by the model itself—rather, HD-equivalent output is upscaled through a video upsampler. FHD standard costs $2.40 more per 20 seconds than HD standard. Whether this premium is justified by quality improvements for fine on-screen text or fast-motion scenes needs to be verified with actual footage.
1135 vs. 1069: How to Read the Internal Elo Ratings
In BFL's published human evaluation, FLUX 3 achieved an Elo score of 1135 for text-to-video. Gemini Omni Flash scored 1090, Minimax H3 scored 1082, and Seedance 2.0 scored 1069—putting FLUX 3 66 points ahead of Seedance 2.0. In July's early evaluation, FLUX 3 won 52% of votes in head-to-head comparisons using 720p, 10-second clips with audio. However, since the evaluation conditions and metrics differ between July's preference rate and August's Elo scores, one cannot conclude from these two figures that the gap has widened.
The image-to-video results are much closer. FLUX 3 scored 1051, while Seedance 2.0 scored 1049—a difference of just 2 points. While the chart shows FLUX 3 in first place, BFL's own description treats the two models as comparable. Extending the claim of "outperforming Seedance 2.0" to cover all modes would lose sight of this narrow margin.
Furthermore, this ranking comes from BFL's internal evaluation. The announcement page does not disclose the set of prompts used for evaluation, the number of generations, the composition of evaluators, or confidence intervals. Elo is a metric that converts head-to-head preference results into relative rankings, and the significance of a 2-point difference cannot be assessed without knowing the number of evaluations or their variance. The top score of 1135 is a promising early result, but it is not a number that guarantees reproducibility across specific production use cases.
Three Boundaries That Remain After General Availability
What has become generally available this time is the video generation component of FLUX 3. BFL has indicated plans to introduce features that combine image, video, and audio references, along with FLUX 3 Image for image generation and editing, and an open-weight version called FLUX 3 Dev. No release dates have been announced for any of these. The current API has not yet reached the stage where multiple input formats can be freely mixed together.
On the safety front, BFL states it worked with third-party company Cinder to evaluate risks—including non-consensual sexual imagery and child sexual abuse content—prior to general availability. The API's safety tolerance level is specified on a scale of 0 to 4, with 2 as the default. For requests that include reference images or video, the maximum is capped at 2, preventing users from relaxing the setting to its highest level.
What remains for practical decision-making includes quality as measured by external evaluations, the actual effectiveness of FHD upscaling, and real-world generation times. BFL has not disclosed generation speed or latency figures, nor has it revealed confidence intervals for its internal Elo ratings. By first testing project-specific prompts and languages in draft mode, then finalizing the same draft_cache in both HD and FHD for comparison, production teams can translate the changes brought by general availability into concrete cost and schedule terms.
