Respuesta rápida: FLUX 3 is Black Forest Labs’ multimodal foundation model for images, video, audio, and action prediction. As of August 11, 2026, BFL lists FLUX 3 Video in Early Access with clips up to 20 seconds, optional native audio, and public per-second pricing. GlobalGPT also has a verified FLUX 3 Video route. In our three-run test, every task completed and returned a 5.04-second MP4 with an audio stream, but exact duration control, typography spacing, and object consistency were not perfect.
That status still needs careful wording. The publicación oficial de lanzamiento y official model page describe FLUX 3 Video as Early Access, while BFL’s pricing page now publishes video rates. GlobalGPT exposes a separate platform route powered by the model. Access through that route does not mean every BFL API, private-weight option, Image model, Action model, or FLUX 3 Dev release is generally available.
This update adds a small hands-on check: three generations on August 11 using the test environment’s flux-3-video model ID, including one exact prompt repeat. It is enough to document real behavior and visible failure points, but three outputs are not a broad benchmark.
Three first-party FLUX 3 Video runs—not a broad benchmark.
| Carreras | Returned duration | Resolución | Audio |
|---|---|---|---|
| 3/3 succeeded | 5.04 s per clip | 1280 × 704 | Decoded data in all files |

What Is FLUX 3?
FLUX 3 is a multimodal foundation model from Black Forest Labs. Unlike a system trained only to turn text into images, BFL says the model jointly learns from images, video, and audio in one architecture. The goal is to build representations that connect spatial structure, movement, sound, cause, and action rather than treating each medium as an isolated feature.

The name may sound like a routine successor to FLUX.1 and FLUX.2, but the scope is much wider. FLUX.2 remains a current image-generation and editing family; the FLUX 2 Pro guide covers that generation’s image workflow, pricing context, and prompting. FLUX 3 is being positioned as a shared backbone for content creation and action prediction.
BFL uses the phrase “world model” to describe this direction. Here, that does not mean the company has announced general intelligence. It means the model is intended to learn how visible objects, motion, audio, and actions relate across time. That distinction matters because the most interesting FLUX 3 claim is not simply “better images”; it is one underlying model serving several types of generation and control.
Why FLUX 3 Is More Than an Image Generator
Most creative AI products are still experienced as separate tools: one model generates images, another produces video, and a third adds speech or sound effects. FLUX 3 is designed around a different premise. If one architecture learns these modalities together, a prompt, reference image, motion sequence, and soundtrack can potentially share the same internal understanding.
- Images provide space: objects, composition, texture, typography, and visual relationships.
- Video adds time: motion, persistence, contact, and cause-and-effect across frames.
- Audio adds another synchronized signal: speech, ambience, and sounds linked to visible events.
- Action prediction connects perception to control: the model’s representation can be adapted to predict what a robot should do next.
For creators, the practical appeal is fewer handoffs between generation stages. Native audio is especially important because it can reduce the gap between making a clip and separately producing dialogue, effects, and synchronization. If sound is your main concern, the current Veo 3.1 sound guide shows how another leading video family approaches native audio today.
FLUX 3 Capabilities at a Glance
| Capacidad | What BFL announced | Status on August 11, 2026 |
|---|---|---|
| FLUX 3 Video | Video generation and editing with optional native audio, including speech, effects, styles, typography, and multi-shot workflows | BFL Early Access; official pricing published; verified GlobalGPT route |
| FLUX 3 Image | Image generation and editing through APIs and private-weight access | Early access announced; this article did not test it |
| FLUX 3 Action | Native action prediction plus specialized models built from the video backbone | Starting with selected partners |
| FLUX 3 Dev | Planned open-weight multimodal backbone for video, audio, image, and action prediction | Announced; date and license not published |
FLUX 3 Video is the clearest near-term product. BFL says it supports text-to-video and editing, synchronized audio, multilingual dialogue, varied aspect ratios, typography, and chained clips. Those features place it in an increasingly crowded market; the Comparativa de los mejores generadores de vídeo con IA provides useful context on what creators can already use while FLUX 3 access remains limited.
The launch post also describes 10-second, 720p text-to-video clips with audio in its preliminary evaluations. BFL explicitly says the model and evaluation harness are still in development. That makes the numbers useful as early company-reported evidence, but not a substitute for independent, repeatable testing after public access expands.
The current model page goes further: it says FLUX 3 Video can generate up to 20 seconds in one run, with audio optional and multilingual speech, effects, and ambience generated with the frames. Those are official capability limits; the hands-on test below records what the connected platform route actually returned from our prompts.
Is FLUX 3 Available Now?
Yes, but only through staged Early Access. BFL has opened an application route rather than a broad self-serve launch. The company can therefore work with selected users and partners while the models, APIs, private-weight options, and evaluation process are still changing.

| Pregunta | Known | Still limited or unknown |
|---|---|---|
| Can people apply to BFL? | Yes, through the official request form | Acceptance criteria and response time |
| Can users try FLUX 3 Video on GlobalGPT? | Yes; the verified product route is linked below | Route limits and quotas depend on the platform account |
| What does BFL price publicly? | Text/image-to-video, continuation, draft, HD, and FHD rates | Private-weight and enterprise terms |
| Will there be open weights? | FLUX 3 Dev is planned as open weight | Release date, license, and hardware requirements |

The safest reading is simple: FLUX 3 is real and the rollout is active, but “Early Access” should not be rewritten as “available to everyone.” If you need a production deadline, wait for written access confirmation and terms before designing a workflow around it.
We Tested FLUX 3 Video: Three Runs, Real Outputs
Test setup: We ran the tests on August 11, 2026 through the Broly task test environment connected to the platform, using the exact model ID flux-3-video. Each request contained only the model ID and prompt. We submitted three tasks: a product-ad prompt, a temporal-consistency prompt, and an exact repeat of the first prompt. The task API reported actual processing times of 88, 190, and 77 seconds respectively. Because the runs were concurrent and the route is still evolving, these timings should not be treated as a latency benchmark.
Test 1: typography, transformation, and native audio
Task ID: task_tl5xE4BIOv2ojEb0VOMFnVT56YtvzhZ8. The prompt asked for an eight-second, two-shot perfume ad with exact label text, a bottle-to-city transformation, a glass click, a whoosh, and one spoken line.

Lo que pasó: The output completed the requested visual arc and kept the blue perfume concept readable across both scenes. The label remained legible, but it visually collapsed the requested space and looked closer to “FLUX3” than the exact “FLUX 3.” The bottle also changed cap and body shape during the transformation. The MP4 contained decoded audio data, but our automated inspection did not transcribe the spoken line, so we are not scoring the dialogue word for word. Most importantly, the returned clip was 5.04 seconds rather than the requested eight seconds.
Test 2: one-shot motion and object consistency
Task ID: task_izFhE8RhTeA5Tb1EaESJqbb7sP6wCwoE. This prompt removed scene changes and asked one wind-up robot to cross three markers in a fixed order while retaining its defining features.

Lo que pasó: This was the cleaner result. The robot moved in the requested direction and crossed the yellow, blue, and green markers in order. Its main identifying details remained visible through the shot, with no extra character or obvious camera cut. The motion looked simple rather than cinematic, but that simplicity made consistency easier to judge. The output again lasted 5.04 seconds and contained decoded audio data.
Test 3: exact repeat of the first prompt
Task ID: task_JsnPFkB0Yqf55EwaoPU8BShBRkGLJT44. We sent the first prompt again without changing a word.

Lo que pasó: The repeat preserved the concept rather than the exact result. It again rendered a blue bottle, readable “FLUX3” lettering, a ribbon of light, and a night skyline, but the bottle geometry, cap, framing, and transition differed. That is normal for an unseeded generation, yet it matters when a brand team expects repeatable product details.
| Test dimension | Resultado observado | Conclusión práctica |
|---|---|---|
| Task completion | 3 of 3 tasks succeeded | The connected route was operational during this test |
| Duration control | All three prompts requested 8 seconds; all returned 5.04 seconds | Verify route-level duration controls before production |
| Typography | Readable label, but exact “FLUX 3” spacing was not preserved | Plan for retries or post-production on brand text |
| Coherencia temporal | The simple robot kept its defining features and marker order | Constrained single-shot prompts can hold up well |
| Repeatability | Same story, different product geometry and camera choices | An exact prompt repeat is not an exact visual repeat |
| Audio | Every MP4 contained decoded audio data; dialogue was not automatically transcribed | Listen and review sync manually before approval |
Nuestro veredicto: FLUX 3 Video followed the high-level scene logic better than the exact production constraints. The smooth product transformation and the robot’s ordered movement were encouraging. The fixed 5.04-second return, imperfect text spacing, and bottle-shape drift are the reasons to test your own workflow before committing. This is a useful hands-on signal, not a model-wide score.
FLUX 3 vs FLUX.2: What Actually Changed?
| Área | FLUX.2 | FLUX 3 |
|---|---|---|
| Primary public identity | Image generation and editing family | Multimodal foundation model |
| Modalidades | Images and image editing workflows | Images, video, audio, and action prediction |
| Acceso actual | Published API models and pricing | BFL Video Early Access, official video pricing, and a verified GlobalGPT route |
| Open model direction | Several FLUX.2 variants and documented endpoints | FLUX 3 Dev open weights announced, details pending |
| Fair performance verdict | Can be tested as an image model today | Video can be tested; a broad family-wide verdict is still premature |
The biggest change is scope, not a verified percentage improvement. FLUX 3 is not merely FLUX.2 with a new image-quality score; it is intended to unify creative generation and physical action in one family. For a current image-only decision, the FLUX 2 vs Nano Banana Pro comparison is more actionable because both sides can be evaluated as image tools today.
A full FLUX 3 vs FLUX.2 comparison still needs matched image prompts, reference edits, typography, prompt adherence, latency, and cost. The three-run video check above narrows the evidence gap, and BFL now publishes FLUX 3 Video rates, but it does not compare the Image model or open-weight variants.
What the Official Demos and Robotics Results Show
BFL’s launch materials show generated video, synchronized audio, visual styles, animated text, and early editing behavior. They demonstrate the product direction and the kinds of outputs the company is targeting. Because these are selected official examples, they do not tell us how often an ordinary user will reproduce the same quality, how long generation takes, or what each result will cost.

The robotics work is more specific. In the FLUX 3 x mimic report, BFL and mimic robotics describe adapting the FLUX 3 video backbone for robot action prediction. The displayed soft-body kitting chart reports a 95% median success rate for FLUX-mimic, with each dashed median based on 20 autonomous trials.
That is a robotics result, not an image- or video-quality benchmark. Its relevance is conceptual: BFL argues that a backbone trained to model motion, contact, weight, and cause can expose representations that are useful for physical tasks. Whether the same approach generalizes across robots, factories, and tasks will require broader independent evidence.
Early Community Reaction on X and Reddit
FLUX 3 attracted immediate attention. On the July 27 verification snapshot, the official BFL announcement on X showed 6,088 likes, 862 reposts, 365 quotes, and 314 replies. Those figures measure launch interest, not model reliability, and they will continue to change.

A Reddit post in r/StableDiffusion titled “Flux 3 looks insane. This was 1 prompt” showed 370 upvotes and 63 comments in the same verification window. The post’s split-screen sample uses a demanding prompt, but it is still one shared example rather than a controlled benchmark.

The right conclusion is that creators are curious about prompt complexity, motion, native sound, and visual consistency. Once access expands, prompt reproducibility should become a core test. In the meantime, the Kling 3.0 prompt formula is a useful reference for structuring scenes, camera behavior, motion, and constraints in currently accessible video tools.
FLUX 3 API, Pricing, and Open Weights
As of August 11, BFL’s published pricing documentation includes FLUX 3 Video. One credit equals $0.01. Text-to-video and image-to-video cost $0.17 per second at HD or $0.29 per second at FHD, while HD drafts cost $0.06 per second. Video continuation costs $0.43 per second at HD or $0.54 per second at FHD, with drafts at $0.12 per second.
| Artículo | Published or verified | Still limited or unknown |
|---|---|---|
| FLUX 3 Video pricing | T2V/I2V: $0.17/s HD, $0.29/s FHD; draft: $0.06/s. Continuation: $0.43/s HD, $0.54/s FHD; draft: $0.12/s | Private-weight and enterprise terms |
| BFL access | Video Early Access, model overview, and pricing are public | General API acceptance, quotas, and account-specific limits |
| Acceso GlobalGPT | Verified FLUX 3 Video product route and successful three-run task test | Account quotas and route-level controls may differ from BFL’s direct offering |
| FLUX 3 Dev | Planned as an open-weight multimodal backbone | Release date, license, parameter size, and hardware needs |
Do not treat “open weight” as identical to “open source.” The weights, code, training data, and license can each have different access conditions. A meaningful open-weight verdict must wait until BFL publishes the actual FLUX 3 Dev release and license.
How to Request FLUX 3 Early Access
- Abra el official FLUX 3 model page and review the current capability labels.
- Follow the request button to the official FLUX 3 Early Access form.
- Describe the capability you need—Video, Image, Action, private weights, or developer access—and the workflow you want to build.
- Include a realistic use case, expected volume, organization details, and why early access matters to the project.
- Wait for written confirmation before assuming access, pricing, limits, or a delivery date.
A precise application is stronger than a general “I want to try it.” For example, a studio could explain that it needs multilingual dialogue and synchronized effects for short branded clips, while a robotics team could describe its manipulation task, robot platform, data, and deployment environment.
If dialogue is the goal, plan the full pipeline rather than only the generation prompt. The Veo 3.1 dialogue and lip-sync workflow highlights the practical issues—speaker timing, voice consistency, shot design, and regeneration—that will also matter when evaluating FLUX 3 Video.
Who Should Watch FLUX 3?
- AI video creators should watch native audio, multilingual dialogue, editing, multi-shot chaining, consistency, speed, and final pricing.
- Image creators should wait for repeatable comparisons covering prompt adherence, reference editing, typography, and cost. Until then, the best AI image generators tested in 2026 are the more practical shortlist.
- Desarrolladores should monitor public model IDs, rate limits, SDK support, private-weight terms, and FLUX 3 Dev’s eventual license.
- Robotics and physical-AI teams should examine whether the FLUX-mimic approach transfers beyond the published task and how much task-specific data adaptation requires.
Before committing a production workflow, use a four-gate FLUX 3 readiness check. A launch demo can clear the capability gate without clearing access, economics, or evidence.
| Gate | What must be true | Evidence to collect |
|---|---|---|
| Acceda a | Your team has written approval for the required capability | Account invitation, terms, region, quotas, and support contact |
| Capacidad | The model completes your real task, not only a showcase prompt | Matched prompts, reference inputs, failure cases, and repeat runs |
| Economics | Cost and turnaround fit the production plan | Price, generation time, retry rate, volume limits, and infrastructure |
| Pruebas | Results remain credible outside selected launch examples | Independent tests, reproducible settings, and versioned outputs |
Want to try the model used in our test? Abrir FLUX 3 Video on GlobalGPT. The product route was live and the connected task environment completed all three runs on August 11, 2026. Check the controls and quota shown in your own account before starting a production batch.
Frequently Asked Questions About FLUX 3
When was FLUX 3 announced?
Black Forest Labs announced FLUX 3 on July 23, 2026. The launch post introduced a multimodal foundation model spanning images, video, audio, and action prediction, with capabilities rolling out through staged Early Access rather than one broad public release.
Is FLUX 3 publicly available?
BFL offers FLUX 3 Video through Early Access and now publishes its video pricing. GlobalGPT also has a verified FLUX 3 Video product route. That does not make FLUX 3 Image, Action, private weights, or FLUX 3 Dev broadly self-serve; each capability still has its own rollout and terms.
Can FLUX 3 generate video with native audio?
Yes. BFL says FLUX 3 Video can generate multilingual speech, effects, and ambience with the frames. All three MP4s in our August 11 test contained decoded audio data. We did not automatically transcribe the tracks, so the test confirms audio presence rather than word-perfect dialogue.
Is FLUX 3 also an image model?
Yes. Image generation and editing are part of the FLUX 3 family, but BFL said FLUX 3 Image Early Access would open in the weeks after the initial announcement. Public endpoint names, final availability, and pricing were not published at the time of this update.
Is FLUX 3 open source or open weight?
BFL has announced future open-weight access through FLUX 3 Dev. It has not yet published the release date, license, parameter size, code scope, or hardware requirements. Until those details arrive, it is more accurate to say “planned open weight” than “open source.”
How is FLUX 3 different from FLUX.2?
FLUX.2 is primarily a current image-generation and editing family with published endpoints and pricing. FLUX 3 is a broader multimodal foundation intended to connect images, video, audio, and action prediction. A fair quality or cost comparison must wait for wider FLUX 3 access.
What did the FLUX 3 hands-on test show?
Three of three tasks succeeded with the flux-3-video model ID. The outputs followed the requested high-level scenes, and the robot test kept its movement order and main features. All three clips were 5.04 seconds despite eight-second prompts; exact label spacing and repeated product geometry were imperfect.
Can I use FLUX 3 on GlobalGPT?
Yes. The verified route is FLUX 3 Video on GlobalGPT. We also completed three generations through the connected task environment on August 11, 2026. Platform access is separate from direct BFL API or private-weight access.
Conclusión final
FLUX 3 is no longer only a launch-page promise. BFL publishes Video capabilities and pricing, GlobalGPT has a verified route, and our three runs produced usable MP4s with audio streams. The model handled scene logic and constrained motion well, but it missed the requested duration and softened exact text and product consistency. Try your real prompt before scaling, and keep watching FLUX 3 Image, Action, private weights, and the FLUX 3 Dev license as those parts of the family mature.



