Recensione di FLUX 3: video, audio, prezzi e accesso

recensione-di-flux-3-video

Risposta rapida: FLUX 3 is Black Forest Labs’ multimodal foundation model for images, video, audio, and action prediction. As of August 11, 2026, BFL lists FLUX 3 Video in Early Access with clips up to 20 seconds, optional native audio, and public per-second pricing. GlobalGPT also has a verified FLUX 3 Video route. In our three-run test, every task completed and returned a 5.04-second MP4 with an audio stream, but exact duration control, typography spacing, and object consistency were not perfect.

That status still needs careful wording. The post di lancio ufficiale e official model page describe FLUX 3 Video as Early Access, while BFL’s pricing page now publishes video rates. GlobalGPT exposes a separate platform route powered by the model. Access through that route does not mean every BFL API, private-weight option, Image model, Action model, or FLUX 3 Dev release is generally available.

This update adds a small hands-on check: three generations on August 11 using the test environment’s flux-3-video model ID, including one exact prompt repeat. It is enough to document real behavior and visible failure points, but three outputs are not a broad benchmark.

REVIEW SNAPSHOT

Three first-party FLUX 3 Video runs—not a broad benchmark.

CorseReturned durationRisoluzioneAudio
3/3 succeeded5.04 s per clip1280 × 704Decoded data in all files

What Is FLUX 3?

FLUX 3 is a multimodal foundation model from Black Forest Labs. Unlike a system trained only to turn text into images, BFL says the model jointly learns from images, video, and audio in one architecture. The goal is to build representations that connect spatial structure, movement, sound, cause, and action rather than treating each medium as an isolated feature.

FLUX 3 official page showing image, video, audio, and action prediction
Black Forest Labs presents FLUX 3 as one multimodal model spanning image, video, audio, and action prediction.

The name may sound like a routine successor to FLUX.1 and FLUX.2, but the scope is much wider. FLUX.2 remains a current image-generation and editing family; the FLUX 2 Pro guide covers that generation’s image workflow, pricing context, and prompting. FLUX 3 is being positioned as a shared backbone for content creation and action prediction.

BFL uses the phrase “world model” to describe this direction. Here, that does not mean the company has announced general intelligence. It means the model is intended to learn how visible objects, motion, audio, and actions relate across time. That distinction matters because the most interesting FLUX 3 claim is not simply “better images”; it is one underlying model serving several types of generation and control.

Why FLUX 3 Is More Than an Image Generator

Most creative AI products are still experienced as separate tools: one model generates images, another produces video, and a third adds speech or sound effects. FLUX 3 is designed around a different premise. If one architecture learns these modalities together, a prompt, reference image, motion sequence, and soundtrack can potentially share the same internal understanding.

  • Images provide space: objects, composition, texture, typography, and visual relationships.
  • Video adds time: motion, persistence, contact, and cause-and-effect across frames.
  • Audio adds another synchronized signal: speech, ambience, and sounds linked to visible events.
  • Action prediction connects perception to control: the model’s representation can be adapted to predict what a robot should do next.

For creators, the practical appeal is fewer handoffs between generation stages. Native audio is especially important because it can reduce the gap between making a clip and separately producing dialogue, effects, and synchronization. If sound is your main concern, the current Veo 3.1 sound guide shows how another leading video family approaches native audio today.

FLUX 3 Capabilities at a Glance

CapacitàWhat BFL announcedStatus on August 11, 2026
FLUX 3 VideoVideo generation and editing with optional native audio, including speech, effects, styles, typography, and multi-shot workflowsBFL Early Access; official pricing published; verified GlobalGPT route
FLUX 3 ImageImage generation and editing through APIs and private-weight accessEarly access announced; this article did not test it
FLUX 3 ActionNative action prediction plus specialized models built from the video backboneStarting with selected partners
FLUX 3 DevPlanned open-weight multimodal backbone for video, audio, image, and action predictionAnnounced; date and license not published

FLUX 3 Video is the clearest near-term product. BFL says it supports text-to-video and editing, synchronized audio, multilingual dialogue, varied aspect ratios, typography, and chained clips. Those features place it in an increasingly crowded market; the Confronto tra i migliori generatori di video basati sull'intelligenza artificiale provides useful context on what creators can already use while FLUX 3 access remains limited.

The launch post also describes 10-second, 720p text-to-video clips with audio in its preliminary evaluations. BFL explicitly says the model and evaluation harness are still in development. That makes the numbers useful as early company-reported evidence, but not a substitute for independent, repeatable testing after public access expands.

The current model page goes further: it says FLUX 3 Video can generate up to 20 seconds in one run, with audio optional and multilingual speech, effects, and ambience generated with the frames. Those are official capability limits; the hands-on test below records what the connected platform route actually returned from our prompts.

Is FLUX 3 Available Now?

Yes, but only through staged Early Access. BFL has opened an application route rather than a broad self-serve launch. The company can therefore work with selected users and partners while the models, APIs, private-weight options, and evaluation process are still changing.

FLUX 3 official launch plan for Video, Image, Action, and Dev
BFL’s launch plan separates FLUX 3 Video, Image, Action, and Dev into staged early-access releases.
DomandaKnownStill limited or unknown
Can people apply to BFL?Yes, through the official request formAcceptance criteria and response time
Can users try FLUX 3 Video on GlobalGPT?Yes; the verified product route is linked belowRoute limits and quotas depend on the platform account
What does BFL price publicly?Text/image-to-video, continuation, draft, HD, and FHD ratesPrivate-weight and enterprise terms
Will there be open weights?FLUX 3 Dev is planned as open weightRelease date, license, and hardware requirements
GlobalGPT FLUX 3 Video product route showing text-to-video and image-to-video access
The verified GlobalGPT product route lists FLUX 3 Video with text-to-video, image-to-video, and native-audio positioning.

The safest reading is simple: FLUX 3 is real and the rollout is active, but “Early Access” should not be rewritten as “available to everyone.” If you need a production deadline, wait for written access confirmation and terms before designing a workflow around it.

We Tested FLUX 3 Video: Three Runs, Real Outputs

3/3tasks completed
5.04 seach returned clip
1280 × 704each MP4
Audio presentedecoded in every file

Test setup: We ran the tests on August 11, 2026 through the Broly task test environment connected to the platform, using the exact model ID flux-3-video. Each request contained only the model ID and prompt. We submitted three tasks: a product-ad prompt, a temporal-consistency prompt, and an exact repeat of the first prompt. The task API reported actual processing times of 88, 190, and 77 seconds respectively. Because the runs were concurrent and the route is still evolving, these timings should not be treated as a latency benchmark.

Test 1: typography, transformation, and native audio

Task ID: task_tl5xE4BIOv2ojEb0VOMFnVT56YtvzhZ8. The prompt asked for an eight-second, two-shot perfume ad with exact label text, a bottle-to-city transformation, a glass click, a whoosh, and one spoken line.

PROMPTTest Prompt 1
Create an 8-second 16:9 cinematic product commercial in two connected shots. Shot 1: on a dark studio table, a translucent cobalt-blue perfume bottle rotates slowly under a narrow white spotlight; the front label must read exactly 'FLUX 3' in clean white sans-serif letters, with no other text. Shot 2: the bottle opens into a ribbon of blue light that becomes a night city skyline while the same bottle remains centered and recognizable. Use smooth premium camera motion, realistic glass reflections, and synchronized audio: a soft glass click when the cap lifts, a rising electronic whoosh during the transformation, and a calm voice saying exactly 'Make the moment move.' Do not add subtitles, logos, watermarks, or extra text.

Select the prompt manually if your WordPress setup disables the copy control.

FLUX 3 hands-on test 1: the returned product-ad video.
Five sampled frames from the FLUX 3 typography and audio test
Five frames show the label, bottle rotation, blue-light transition, and city reveal.

Cosa è successo: The output completed the requested visual arc and kept the blue perfume concept readable across both scenes. The label remained legible, but it visually collapsed the requested space and looked closer to “FLUX3” than the exact “FLUX 3.” The bottle also changed cap and body shape during the transformation. The MP4 contained decoded audio data, but our automated inspection did not transcribe the spoken line, so we are not scoring the dialogue word for word. Most importantly, the returned clip was 5.04 seconds rather than the requested eight seconds.

Test 2: one-shot motion and object consistency

Task ID: task_izFhE8RhTeA5Tb1EaESJqbb7sP6wCwoE. This prompt removed scene changes and asked one wind-up robot to cross three markers in a fixed order while retaining its defining features.

PROMPTTest Prompt 2
Create an 8-second 16:9 single continuous tracking shot with no cuts. A small red wind-up robot with a square head, two round silver eyes, one brass key on its back, and white boots walks from left to right across three floor markers in this exact order: yellow, blue, then green. The camera tracks at waist height while the robot keeps the same body shape, colors, eye count, back key, and walking direction for the entire shot. Add synchronized metallic footsteps and a quiet room tone. No speech, no text, no extra characters, no camera cut, and no change of costume.

Select the prompt manually if your WordPress setup disables the copy control.

FLUX 3 hands-on test 2: the returned one-shot robot video.
Five sampled frames from the FLUX 3 temporal consistency test
The robot moves left to right across yellow, blue, and green while keeping its red body, two eyes, white boots, and back key.

Cosa è successo: This was the cleaner result. The robot moved in the requested direction and crossed the yellow, blue, and green markers in order. Its main identifying details remained visible through the shot, with no extra character or obvious camera cut. The motion looked simple rather than cinematic, but that simplicity made consistency easier to judge. The output again lasted 5.04 seconds and contained decoded audio data.

Test 3: exact repeat of the first prompt

Task ID: task_JsnPFkB0Yqf55EwaoPU8BShBRkGLJT44. We sent the first prompt again without changing a word.

PROMPTRepeat Prompt
Create an 8-second 16:9 cinematic product commercial in two connected shots. Shot 1: on a dark studio table, a translucent cobalt-blue perfume bottle rotates slowly under a narrow white spotlight; the front label must read exactly 'FLUX 3' in clean white sans-serif letters, with no other text. Shot 2: the bottle opens into a ribbon of blue light that becomes a night city skyline while the same bottle remains centered and recognizable. Use smooth premium camera motion, realistic glass reflections, and synchronized audio: a soft glass click when the cap lifts, a rising electronic whoosh during the transformation, and a calm voice saying exactly 'Make the moment move.' Do not add subtitles, logos, watermarks, or extra text.

Select the prompt manually if your WordPress setup disables the copy control.

Five sampled frames from the repeated FLUX 3 product-ad prompt
The repeated prompt produced the same high-level story but a different bottle design, cap, camera path, and city reveal.

Cosa è successo: The repeat preserved the concept rather than the exact result. It again rendered a blue bottle, readable “FLUX3” lettering, a ribbon of light, and a night skyline, but the bottle geometry, cap, framing, and transition differed. That is normal for an unseeded generation, yet it matters when a brand team expects repeatable product details.

Test dimensionRisultato osservatoConclusione pratica
Task completion3 of 3 tasks succeededThe connected route was operational during this test
Duration controlAll three prompts requested 8 seconds; all returned 5.04 secondsVerify route-level duration controls before production
TypographyReadable label, but exact “FLUX 3” spacing was not preservedPlan for retries or post-production on brand text
Coerenza temporaleThe simple robot kept its defining features and marker orderConstrained single-shot prompts can hold up well
RepeatabilitySame story, different product geometry and camera choicesAn exact prompt repeat is not an exact visual repeat
AudioEvery MP4 contained decoded audio data; dialogue was not automatically transcribedListen and review sync manually before approval

Il nostro verdetto: FLUX 3 Video followed the high-level scene logic better than the exact production constraints. The smooth product transformation and the robot’s ordered movement were encouraging. The fixed 5.04-second return, imperfect text spacing, and bottle-shape drift are the reasons to test your own workflow before committing. This is a useful hands-on signal, not a model-wide score.

FLUX 3 vs FLUX.2: What Actually Changed?

AreaFLUX.2FLUX 3
Primary public identityImage generation and editing familyMultimodal foundation model
ModalitàImages and image editing workflowsImages, video, audio, and action prediction
Accesso attualePublished API models and pricingBFL Video Early Access, official video pricing, and a verified GlobalGPT route
Open model directionSeveral FLUX.2 variants and documented endpointsFLUX 3 Dev open weights announced, details pending
Fair performance verdictCan be tested as an image model todayVideo can be tested; a broad family-wide verdict is still premature

The biggest change is scope, not a verified percentage improvement. FLUX 3 is not merely FLUX.2 with a new image-quality score; it is intended to unify creative generation and physical action in one family. For a current image-only decision, the FLUX 2 vs Nano Banana Pro comparison is more actionable because both sides can be evaluated as image tools today.

A full FLUX 3 vs FLUX.2 comparison still needs matched image prompts, reference edits, typography, prompt adherence, latency, and cost. The three-run video check above narrows the evidence gap, and BFL now publishes FLUX 3 Video rates, but it does not compare the Image model or open-weight variants.

What the Official Demos and Robotics Results Show

BFL’s launch materials show generated video, synchronized audio, visual styles, animated text, and early editing behavior. They demonstrate the product direction and the kinds of outputs the company is targeting. Because these are selected official examples, they do not tell us how often an ordinary user will reproduce the same quality, how long generation takes, or what each result will cost.

FLUX-mimic soft-body kitting benchmark chart
In BFL’s soft-body kitting chart, FLUX-mimic shows a 95% median success rate across 20 autonomous trials.

The robotics work is more specific. In the FLUX 3 x mimic report, BFL and mimic robotics describe adapting the FLUX 3 video backbone for robot action prediction. The displayed soft-body kitting chart reports a 95% median success rate for FLUX-mimic, with each dashed median based on 20 autonomous trials.

That is a robotics result, not an image- or video-quality benchmark. Its relevance is conceptual: BFL argues that a backbone trained to model motion, contact, weight, and cause can expose representations that are useful for physical tasks. Whether the same approach generalizes across robots, factories, and tasks will require broader independent evidence.

Early Community Reaction on X and Reddit

FLUX 3 attracted immediate attention. On the July 27 verification snapshot, the official BFL announcement on X showed 6,088 likes, 862 reposts, 365 quotes, and 314 replies. Those figures measure launch interest, not model reliability, and they will continue to change.

Black Forest Labs X post announcing FLUX 3 Early Access
BFL’s launch post describes FLUX 3 as a unified model for image, video, audio, and action prediction.

A Reddit post in r/StableDiffusion titled “Flux 3 looks insane. This was 1 prompt” showed 370 upvotes and 63 comments in the same verification window. The post’s split-screen sample uses a demanding prompt, but it is still one shared example rather than a controlled benchmark.

Reddit FLUX 3 split-screen video sample post
A Reddit user shared a demanding split-screen FLUX 3 prompt; the post drew 63 comments and an enthusiastic title.

The right conclusion is that creators are curious about prompt complexity, motion, native sound, and visual consistency. Once access expands, prompt reproducibility should become a core test. In the meantime, the Kling 3.0 prompt formula is a useful reference for structuring scenes, camera behavior, motion, and constraints in currently accessible video tools.

FLUX 3 API, Pricing, and Open Weights

As of August 11, BFL’s published pricing documentation includes FLUX 3 Video. One credit equals $0.01. Text-to-video and image-to-video cost $0.17 per second at HD or $0.29 per second at FHD, while HD drafts cost $0.06 per second. Video continuation costs $0.43 per second at HD or $0.54 per second at FHD, with drafts at $0.12 per second.

ArticoloPublished or verifiedStill limited or unknown
FLUX 3 Video pricingT2V/I2V: $0.17/s HD, $0.29/s FHD; draft: $0.06/s. Continuation: $0.43/s HD, $0.54/s FHD; draft: $0.12/sPrivate-weight and enterprise terms
BFL accessVideo Early Access, model overview, and pricing are publicGeneral API acceptance, quotas, and account-specific limits
Accesso al GlobalGPTVerified FLUX 3 Video product route and successful three-run task testAccount quotas and route-level controls may differ from BFL’s direct offering
FLUX 3 DevPlanned as an open-weight multimodal backboneRelease date, license, parameter size, and hardware needs

Do not treat “open weight” as identical to “open source.” The weights, code, training data, and license can each have different access conditions. A meaningful open-weight verdict must wait until BFL publishes the actual FLUX 3 Dev release and license.

How to Request FLUX 3 Early Access

  1. Aprire la sezione official FLUX 3 model page and review the current capability labels.
  2. Follow the request button to the official FLUX 3 Early Access form.
  3. Describe the capability you need—Video, Image, Action, private weights, or developer access—and the workflow you want to build.
  4. Include a realistic use case, expected volume, organization details, and why early access matters to the project.
  5. Wait for written confirmation before assuming access, pricing, limits, or a delivery date.

A precise application is stronger than a general “I want to try it.” For example, a studio could explain that it needs multilingual dialogue and synchronized effects for short branded clips, while a robotics team could describe its manipulation task, robot platform, data, and deployment environment.

If dialogue is the goal, plan the full pipeline rather than only the generation prompt. The Veo 3.1 dialogue and lip-sync workflow highlights the practical issues—speaker timing, voice consistency, shot design, and regeneration—that will also matter when evaluating FLUX 3 Video.

Who Should Watch FLUX 3?

  • AI video creators should watch native audio, multilingual dialogue, editing, multi-shot chaining, consistency, speed, and final pricing.
  • Image creators should wait for repeatable comparisons covering prompt adherence, reference editing, typography, and cost. Until then, the best AI image generators tested in 2026 are the more practical shortlist.
  • Sviluppatori should monitor public model IDs, rate limits, SDK support, private-weight terms, and FLUX 3 Dev’s eventual license.
  • Robotics and physical-AI teams should examine whether the FLUX-mimic approach transfers beyond the published task and how much task-specific data adaptation requires.

Before committing a production workflow, use a four-gate FLUX 3 readiness check. A launch demo can clear the capability gate without clearing access, economics, or evidence.

GateWhat must be trueEvidence to collect
AccessoYour team has written approval for the required capabilityAccount invitation, terms, region, quotas, and support contact
CapacitàThe model completes your real task, not only a showcase promptMatched prompts, reference inputs, failure cases, and repeat runs
EconomicsCost and turnaround fit the production planPrice, generation time, retry rate, volume limits, and infrastructure
ProveResults remain credible outside selected launch examplesIndependent tests, reproducible settings, and versioned outputs

Want to try the model used in our test? Aperto FLUX 3 Video on GlobalGPT. The product route was live and the connected task environment completed all three runs on August 11, 2026. Check the controls and quota shown in your own account before starting a production batch.

Frequently Asked Questions About FLUX 3

When was FLUX 3 announced?

Black Forest Labs announced FLUX 3 on July 23, 2026. The launch post introduced a multimodal foundation model spanning images, video, audio, and action prediction, with capabilities rolling out through staged Early Access rather than one broad public release.

Is FLUX 3 publicly available?

BFL offers FLUX 3 Video through Early Access and now publishes its video pricing. GlobalGPT also has a verified FLUX 3 Video product route. That does not make FLUX 3 Image, Action, private weights, or FLUX 3 Dev broadly self-serve; each capability still has its own rollout and terms.

Can FLUX 3 generate video with native audio?

Yes. BFL says FLUX 3 Video can generate multilingual speech, effects, and ambience with the frames. All three MP4s in our August 11 test contained decoded audio data. We did not automatically transcribe the tracks, so the test confirms audio presence rather than word-perfect dialogue.

Is FLUX 3 also an image model?

Yes. Image generation and editing are part of the FLUX 3 family, but BFL said FLUX 3 Image Early Access would open in the weeks after the initial announcement. Public endpoint names, final availability, and pricing were not published at the time of this update.

Is FLUX 3 open source or open weight?

BFL has announced future open-weight access through FLUX 3 Dev. It has not yet published the release date, license, parameter size, code scope, or hardware requirements. Until those details arrive, it is more accurate to say “planned open weight” than “open source.”

How is FLUX 3 different from FLUX.2?

FLUX.2 is primarily a current image-generation and editing family with published endpoints and pricing. FLUX 3 is a broader multimodal foundation intended to connect images, video, audio, and action prediction. A fair quality or cost comparison must wait for wider FLUX 3 access.

What did the FLUX 3 hands-on test show?

Three of three tasks succeeded with the flux-3-video model ID. The outputs followed the requested high-level scenes, and the robot test kept its movement order and main features. All three clips were 5.04 seconds despite eight-second prompts; exact label spacing and repeated product geometry were imperfect.

Can I use FLUX 3 on GlobalGPT?

Yes. The verified route is FLUX 3 Video on GlobalGPT. We also completed three generations through the connected task environment on August 11, 2026. Platform access is separate from direct BFL API or private-weight access.

Conclusione finale

FLUX 3 is no longer only a launch-page promise. BFL publishes Video capabilities and pricing, GlobalGPT has a verified route, and our three runs produced usable MP4s with audio streams. The model handled scene logic and constrained motion well, but it missed the requested duration and softened exact text and product consistency. Try your real prompt before scaling, and keep watching FLUX 3 Image, Action, private weights, and the FLUX 3 Dev license as those parts of the family mature.

Condividi il post:

Messaggi correlati