Black Forest Labs (BFL) has begun offering FLUX 3 Image, an image generation and editing model. Users can set rectangular boxes within an image to specify where people or text should be placed, or which area to edit. BFL says the model is better at preserving elements that were not targeted, even through repeated edits. It also supports combining up to 10 reference images and native 4K output.
However, the developer documentation states that shadows, reflections and ambient light can change even outside the specified box. For work that requires strict control over layout and detail, such as advertisements and print layouts, how reliably the model changes only the intended area will be a key factor in choosing it.
Specify placement by coordinates, then move the same element later
With FLUX 3 Image, you can specify where each element goes numerically, not just describe the image as a whole. BFL's approach treats the whole image as a coordinate space from 0 to 1000 both vertically and horizontally, and specifies a bounding box around each element in the order top, left, bottom, right. Even if the aspect ratio changes, positions can be expressed in the same 0–1000 coordinate system.
In the official usage guide, each element is given an ID, a description and a box for placement, and these are combined with a text description of the whole image.
For example, to place a headline at the top of a poster, a person in the center and a product at the bottom, you assign an ID and a box to each. This conveys which element goes where more clearly than a text instruction such as "move the product slightly to the right."
![FireShot Capture 080 - FLUX 3 Image - Black Forest Labs - [docs.bfl.ai].webp](https://media.xenospectrum.com/large_Fire_Shot_Capture_080_FLUX_3_Image_Black_Forest_Labs_docs_bfl_ai_8d338d6962.webp)
When editing, specifying different boxes before and after the change lets you move or resize the same element. To delete an element, remove its box on the output side and also instruct the deletion in text.
BFL's model page shows an example in which a surfer's wetsuit and board are both changed to red while surrounding elements such as the waves and sky are left intact. It shows how multiple changes can be requested at once while making clear which elements should be kept.
The model is designed for a workflow in which a generated image is not simply taken as finished, but is refined after the composition is settled, adjusting colors and subjects partially.
Not every box has to be specified by hand. BFL also describes a method in which an LLM is given the text and aspect ratio, works out the composition, and automatically produces descriptions and coordinates for each element. A person can then correct only the positions that look off.
Even when a short instruction is expanded into a detailed prompt, the element IDs and coordinates specified by the user are passed to the model unchanged. This lets AI propose the composition while a person fixes or adjusts the placements that matter.
The area outside the edit box is not guaranteed to stay fixed
Being able to specify the edit location with a rectangle does not mean pixels outside it will never change.
BFL's image editing guide says that pixels outside the box generally stay the same, but explicitly notes that shadows, reflections, ambient light and similar effects may change.
BFL also reported that in its testing, trying to add a new element in a very small box of roughly 40×25 pixels sometimes failed to make the element appear at all. This is one example BFL tried, and 40×25 pixels is not a defined minimum size that applies to every image.
The coordinate specification guide likewise explains that bounding boxes are meant to guide the placement and editing of elements, not to serve as strict masks for cutting out parts of an image.
In other words, when you change a subject's color or shape, its shadow and reflections in the surroundings may change naturally as well. That can be desirable for a photograph, but it can be a problem when the original image must be strictly preserved.
This difference matters especially for product photography and advertising. Even if a product's color can be changed exactly as intended, an ad asset may be unusable if a nearby logo or small text also changes.
BFL also publishes before-and-after images along with difference views showing which pixels changed. But a successful edit in an official example is not the same as the surroundings staying intact across every image and every round of repeated edits.
When testing in practice, it is worth comparing not only the changed area but also the text, logos and background you want to keep. Apply several rounds of edits to the same source image and see how far changes spread beyond the specified area.
A product photo where shadows and reflections can change naturally and a print layout where text and layout must be strictly fixed demand different levels of precision from the same editing feature.
"Up to 10 images" is not the only advance over FLUX.2
The previous-generation FLUX.2 also could combine multiple reference images and render text within images. According to BFL's official FLUX.2 specifications, [pro], [max] and [flex] support up to 8 reference images via the API and up to 10 in the Playground.
Treating the ability to use 10 reference images as something new in FLUX 3 Image would therefore misrepresent the difference between generations.
| Comparison | FLUX.2 official specs | FLUX 3 Image official specs |
|---|---|---|
| API reference images | [pro], [max], [flex]: up to 8 | Up to 10 |
| Playground reference images | [pro], [max], [flex]: up to 10 | Up to 10 reference images |
| Web search before generation | Offered in [max]; not supported in [pro] | Enabled by default; can be disabled |
| Placement and editing control | Supports structured instructions and control of color and composition | Explicit placement and editing method using element IDs and coordinates |
Comparison of BFL specifications as checked on October 3, 2026. Supported features differ by FLUX.2 variant, and this is not a table comparing image quality or editing success rates under identical conditions.
Being able to handle up to 10 images via the API as well makes it easier to combine multiple materials, such as people, products, clothing and backgrounds, into a single generation instruction.
For users who were already using 10 reference images in the Playground, however, the ability to specify position and changes for each element in detail, rather than the number of images itself, is likely the stronger reason to try FLUX 3 Image.
In the FLUX 3 Image API specifications, "grounding," which searches the web and images to supplement information before generation, is also enabled by default. If you disable it, images are generated from the prompt and reference images alone, without external search.
Search can help when depicting real-world subjects, but using it does not mean generated text or facts will always be accurate. When comparing models or settings, whether search is enabled should be kept consistent.
4K is for finishing; costs rise with resolution
On BFL's direct API, the price per image varies by output resolution. The prices listed on the official pricing page are as follows.
| Output tier | Regular price per image | Launch half-price per image |
|---|---|---|
| 768×768 | $0.041 | $0.0205 |
| 1K | $0.048 | $0.024 |
| 1.5K | $0.07 | $0.035 |
| 2K | $0.1 | $0.05 |
| 4K | $0.607 | $0.3035 |
BFL's FLUX 3 Image pricing, checked October 3, 2026. The half-price period runs from 15:00 UTC on October 1 to 15:00 UTC on October 8, which is 0:00 on October 2 to 0:00 on October 9 in Japan time.
At regular prices, generating 100 images costs $4.80 at 1K, $10 at 2K and $60.70 at 4K. 4K costs about 12.6 times as much as 1K.
The figures are each unit price multiplied by 100, and the ratio is 0.607 ÷ 0.048 rounded to one decimal place. During the half-price period, the same 100 images would cost $2.40 at 1K, $5 at 2K and $30.35 at 4K.
This compares only the cost of generating images on BFL's direct API. It does not include API prices via other providers or the number of generations needed to get a usable image.
The term "native 4K" also calls for care. BFL has published a soba restaurant example output directly by the model at 5456×3072 pixels. That is roughly 16.8 megapixels and is not the same size as the 3840×2160-pixel "4K" of typical TVs and displays.
With FLUX 3 Image, the actual output dimensions vary with the aspect ratio. For print or delivery files, therefore, you need to check the actual pixel count rather than relying on the name "4K."
In production, you can keep trial costs down by testing composition and editing approaches several times at low resolution and generating only the chosen candidates at high resolution.
However, an edit that works at low resolution will not necessarily give exactly the same result at 4K. In the end, text, fine details and unchanged areas need to be checked in the high-resolution output as well. Production cost should account for both the price of higher resolution and the number of times edits must be redone.
What to check before building it into a production workflow
FLUX 3 Image is available through BFL's Playground and API. BFL also says it offers commercial weights under contract to enterprises generating large volumes of images, supporting deployment on their own infrastructure and additional training.
The contracts and operating environment required differ between using the API and running the model on your own servers. The availability of commercial weights does not mean anyone can freely download the model weights and use them commercially without conditions.
For people who produce large numbers of color variants for ads, or who place multiple materials into a single layout, the workflow of layering changes onto the same image is worth testing. If elements can be specified and changed partially, it may reduce how often the whole image has to be regenerated from scratch once the composition is settled.
On the other hand, if parts you did not want to change move, the time spent checking and fixing grows. A low API unit price alone therefore does not show that overall production costs will fall.
Can logos and text you want to keep survive repeated changes to the same material? Can you tolerate changes in shadows and reflections? And how many generations does it take to get an image you can actually use?
If those conditions can be verified on your own work, FLUX 3 Image could reduce the effort of recreating the entire image each time a correction is needed after the composition is decided.
