One model can build the whole sound scene.
Seed Audio 1.0 can combine voice, Foley, ambience, timing, and instrumental music. Our tests found an unusually high creative ceiling—and a repeatability problem you should not ignore.
Seed Audio 1.0 is one of the first audio models I have tested that feels designed around a complete scene rather than a single output type. It can speak, act, add ambience, create Foley, place dialogue on a timeline, and produce instrumental music from one prompt.
The short verdict: Seed Audio 1.0 has an unusually high ceiling for unified audio creation, especially for timed dialogue and cinematic sound scenes. Its weakness is repeatability. A strong run can sound remarkably complete, while another run from the same prompt may add garbled speech, miss one event, or change the requested voice.
I generated 23 scored samples through the Anywhere/BrolyAI test environment. Together, they covered narration, two-character dialogue, timestamped speech, SFX-only sequences, voice-plus-music scenes, standalone music, multilingual performance, two-minute monologues, and reference-audio continuation.

Seed Audio 1.0 review: the quick verdict

| Categoria | Result from our tests |
|---|---|
| Expressive voice | High quality on the best run, but inconsistent across repeats |
| Multi-speaker dialogue | Accurate scripts and strong speaker separation |
| Dialogue timing | The most reliable feature in our test set |
| Efeitos sonoros | Realistic space and event order; weak at exact event counts |
| Voice + SFX + music | Convincing complete scenes when every requested event lands |
| Standalone music | More capable than expected; produced structured 30-second cues |
| Multilingual voice | Excellent best case, but one of two runs hallucinated extra speech |
| Two-minute audio | Strong voice identity and narrative coherence across both runs |
| Reference audio | Unreliable through our test venue; voice similarity was poor |
Melhor para: creators making radio drama, game scenes, podcast concepts, cinematic prototypes, and audio-first storyboards.
Less suitable for: jobs that require deterministic wording, exact repeated event counts, or dependable voice cloning from one unattended generation.
What is Seed Audio 1.0?

Seed Audio 1.0 is ByteDance Seed’s unified text-to-audio creation model. Instead of treating speech, ambience, effects, and music as separate jobs, it can generate them together as one sound scene. The official product page says it supports dialogue timing at 100-millisecond intervals and can generate up to two minutes in one pass.
That unified design is the real attraction. A prompt can describe who speaks, how they sound, what happens in the room, which effects enter later, and what kind of music sits under the scene. In a good run, the result sounds composed rather than stitched together.
There is an important boundary. ByteDance’s own roadmap says fine-grained timing currently focuses mainly on character dialogue. More precise control over sound effects, ambience, and music remains an area for further development. Our tests matched that distinction: dialogue timing was excellent, while exact Foley counts were less reliable.
Official sources checked August 6, 2026:
- Seed Audio 1.0 product page
- ByteDance Seed introduction
- BytePlus Audio 1.0 API documentation
- BytePlus Audio 1.0 pricing
How we tested Seed Audio 1.0
We designed nine tasks and ran 23 scored generations. Most capability tests used two or three repeats, because one polished sample tells you very little about production reliability.
| Teste | Tarefa | Runs |
|---|---|---|
| T01 | Expressive English narration | 3 |
| T02 | Two-character radio dialogue | 3 |
| T03 | Dialogue at 0, 4, and 9 seconds | 3 |
| T04 | SFX-only hallway sequence | 2 |
| T05 | Voice, ambience, SFX, and music | 3 |
| T06 | Standalone instrumental music | 3 |
| T07 | One character across English, Japanese, and German | 2 |
| T08 | 120-second podcast monologue | 2 |
| T09 | Continuation from a reference voice | 2 |
The 23 planned outputs contained about 12.5 minutes of generated audio and reported $1.7504 in upstream generation cost through the test environment. Seed Audio did not return text-token usage. Billing was based on generated seconds, so generation tokens are recorded as not applicable.
All scored outputs were 40 kHz, 16-bit, stereo PCM WAV files. When direct automated playback was not available, Gemini 3.6 Flash received the WAV files as audio input and returned structured transcripts, event checks, timing estimates, and artifact notes. Those judgments are automated listening evaluations, not human MOS scores.
Voice generation: impressive highs, uneven repeatability
T01 · Expressive narration
High varianceOne run hallucinated extra speech, one sounded stilted, and one was natural and complete.
Prompt utilizado
One calm female narrator delivers two exact sentences, moving from concern to confidence. No music or SFX.
The narration prompt asked one calm female speaker to deliver two exact sentences, moving from restrained concern to growing confidence, with no background sound.
The three runs behaved very differently:
- Run 1: delivered the requested sentence, then continued with several seconds of garbled, unscripted speech.
- Run 2: spoke the exact text but sounded slow and staccato.
- Run 3: delivered the complete script naturally, with the best pacing and voice consistency.
This is the central pattern of the review. Seed Audio 1.0 can produce a strong voice performance, but a single run is not enough for wording-critical work. Generate alternatives and listen through the entire tail before publishing.
The model did respect the negative instructions in all three narration runs: no music, no sound effects, and no unwanted background were detected.
Multi-character dialogue and speaker separation
T02 · Two-character radio drama
3/3 exact scriptsAll three scripts were exact. Room tone and lead-in length varied between takes.
Prompt utilizado
Maya whispers, Eli answers calmly, and a refrigerator hum sits under three fixed dialogue turns.
Our convenience-store scene used two adult characters and three fixed lines. Maya had to whisper tensely, while Eli tried to sound calm. The only background requested was a subtle refrigerator hum.
All three runs delivered the dialogue in the correct order with two distinct voices. The best run combined exact lines, clean speaker separation, convincing emotion, and a continuous refrigerator hum.
The other runs showed smaller scene-level problems. One spent roughly the first half of its 30-second duration on ambience before beginning the dialogue. Another omitted the requested refrigerator hum. Neither mistake ruined the spoken scene, but both reduce editing efficiency.
For audio drama, the practical result is encouraging: Seed Audio understands character turns and vocal contrast. You may still need to trim long lead-ins or regenerate for the right room tone.
Dialogue timing was the strongest result
T03 · Prompt-level timing
Mais confiávelAll three runs placed the lines near the requested points within the evaluator’s timing uncertainty.
Prompt utilizado
Place three radio lines at 0.0, 4.0, and 9.0 seconds with male/female/male speakers.

The timing test requested three radio lines:
- “Unit seven, report.” at 0.0 seconds.
- “North entrance secure.” at 4.0 seconds.
- “Hold position.” at 9.0 seconds.
Across all three runs, the lines, order, and male-female-male speaker pattern were correct. Gemini’s audio estimates placed the repeated-run onsets around 0.1, 4.0, and 9.0 seconds. The first run was estimated at approximately 0.1, 3.9, and 8.9 seconds.
Those measurements should not be presented as laboratory-grade proof of 100-millisecond precision—the evaluator itself stated roughly 0.1-second uncertainty. Even with that caveat, this was the most repeatable capability we tested. If you need dialogue to enter near planned timeline positions, Seed Audio 1.0 is unusually promising.
SFX generation: realistic scenes, imperfect counting
T04 · SFX-only hallway
Counting failedThe space and order were convincing, but the two runs produced four footsteps and two footsteps.
Prompt utilizado
Keys, a heavy door, exactly three footsteps, thunder, and a final slam. No voice or music.
The SFX-only prompt described a small tiled hallway: keys near the microphone, a heavy wooden door opening, three slow footsteps moving away, distant thunder, and a final door slam. Speech and music were forbidden.
Both outputs created a believable acoustic space, preserved the event order, and avoided voice or music leakage. The effects were judged realistic, with good distance and room consistency.
Neither run followed the exact count. One produced four footsteps; the other produced two. That makes Seed Audio useful for generating a convincing Foley concept, but less dependable when an editor needs exactly three impacts, knocks, footsteps, or gunshots.
The best workflow is to use one-pass SFX generation for atmosphere and rapid prototyping, then replace or edit count-critical events in a DAW.
Voice, ambience, SFX, and music in one scene
T05 · Complete cinematic mix
2/3 completeTwo of three runs completed every cue. One missed the final metallic slam.
Prompt utilizado
Rain alley, traffic, siren, whispered detective, string pulse, fast footsteps, and metallic slam.

The complete-scene test was deliberately crowded. It asked for a wet night alley, water drips, traffic, a moving police siren, a whispered detective line, restrained strings, fast footsteps, and a metallic door slam.
Two of three runs included the complete event sequence. Dialogue remained clear over ambience and music, and the scene felt spatially coherent rather than like unrelated clips layered together. One run missed the final metallic slam, despite otherwise producing the right atmosphere and exact spoken line.
This is where the unified model makes the strongest case for itself. For storyboarding and creative ideation, generating an entire mixed scene from one prompt is faster and more expressive than sourcing every component separately. For final production, important end cues still need verification.
Can Seed Audio 1.0 generate music?
T06 · Standalone music
Music worksAll three created structured music. One made the muted piano difficult to distinguish.
Prompt utilizado
A 72 BPM suspense cue with cello ostinato, muted piano, analog drone, build, and unresolved ending.

Yes. In our tests, Seed Audio 1.0 generated recognizable standalone instrumental music rather than only a vague ambient texture.
The prompt asked for a 30-second suspense cue at 72 BPM with low cello ostinato, muted piano pulses, a soft analog drone, a clear build, and an unresolved final chord. All three runs formed a coherent musical cue with the requested tempo feel and ending. Two included every requested element; in one run, the muted piano was difficult to distinguish.
This does not automatically make Seed Audio a replacement for a dedicated song generator. Our task was a short instrumental cinematic cue, not a verse-chorus song with lyrics. Still, the music capability is substantial enough to support game scenes, trailers, podcasts, and previsualization—especially when music must interact with speech and sound design in the same generation.
Multilingual performance: excellent best case, weak worst case
T07 · One voice, three languages
1/2 cleanOne run was flawless; the other repeated and hallucinated speech around the language switch.
Prompt utilizado
One female character speaks English, Japanese, and German with the same identity and calm emotion.

The multilingual task asked one female character to speak the same recording-space role in English, Japanese, and German. Two runs produced opposite impressions.
The stronger run delivered all three sentences accurately, preserved the same voice, kept calm emotion, and used clean pauses. The weaker run started correctly in English, then repeated and hallucinated speech around the Japanese section and damaged the German opening. Its perceived cross-language voice consistency fell sharply.
This means Seed Audio 1.0 can achieve convincing multilingual character performance, but the feature needs multiple candidates and language-aware review. Do not approve a multilingual output by checking only the first language.
Two-minute generation and long-form voice stability
T08 · 120-second monologue
2/2 stableBoth files reached 120 seconds with one stable voice. One had a brief pronunciation stumble.
Prompt utilizado
One male podcast host tells a structured weather-station story with three discoveries and no background.

Both long-form tasks produced exact 120-second WAV files. Each used one male podcast host with dry studio sound, no music, and no sound effects.
The first run maintained one voice and a coherent investigative structure throughout: a clear opening, three chronological discoveries, a move from curiosity to unease, and a final unresolved question. No obvious voice drift or garbled segment was detected.
The second run was similarly coherent and stable, with one brief pronunciation stumble around 1:31. That is a strong result for a two-minute one-pass monologue. It suggests Seed Audio’s long-form claim is not merely a duration limit; the model can sustain identity and narrative structure across the full window.
Reference audio continuity needs caution
T09 · Reference continuation
Route uncertainText was exact, but voice similarity was poor; one output changed to a male voice and added Foley.
Prompt utilizado
Continue from an authorized female reference while preserving voice identity and intimate recording style.
The reference test used the strongest narration sample as an authorized input and asked for a continuation in the same general female identity and intimate recording style.
Both outputs spoke the requested continuation exactly, but neither matched the reference well. The first shifted into a breathy full whisper and received a conservative voice-similarity score of 2/5. The second was judged to use a male voice, scored 1/5 for similarity, and added roughly 17 seconds of unwanted Foley before speaking.
There is an important venue limitation here. The Anywhere task endpoint accepted the reference field, but it did not return confirmation showing how the upstream model consumed that reference. These results prove that reference continuity was unreliable through our test route; they do not prove that BytePlus’s native API always behaves the same way.
Seed Audio 1.0 pricing and limits

Billed by the second, with 60 free trial minutes when the service is activated.
maximum output
prompt characters
audio references
reference image
BytePlus lists the official pay-as-you-go price at $0.15 per generated minute, billed by the second, with 60 free trial minutes when the service is activated. The API documentation sets a 120-second maximum output, supports prompts up to 3,000 characters, allows up to three reference audio clips, and accepts one reference image.
Each reference audio clip can be up to 30 seconds and 10 MB. The reference image limit is 10 MB. These are BytePlus-native limits; third-party platforms may expose a smaller input form or use different pricing.
Our Anywhere/BrolyAI test route reported upstream cost separately and averaged about $0.14 per generated minute. That test-environment field should not be confused with BytePlus retail pricing or the final price of another platform.
Seed Audio 1.0 pros and cons
Where it wins
- Unified voice, Foley, ambience, and music
- Excellent dialogue timing
- Strong long-form identity
- Useful cinematic music
Where it breaks
- Large take-to-take variance
- Occasional garbled speech
- Weak exact SFX counting
- Unreliable reference continuity in our route
Prós
- Generates voice, ambience, Foley, and music in one pass.
- Excellent dialogue timing in our repeated tests.
- Strong two-speaker separation and exact script handling.
- Produces useful standalone cinematic music.
- Maintains voice identity and narrative coherence across two minutes.
- Outputs clean 40 kHz stereo WAV through our test route.
Contras
- Repeat runs can vary dramatically in naturalness and reliability.
- Occasional hallucinated or garbled speech.
- Exact SFX counts are unreliable.
- Some complete scenes miss a final requested event.
- Multilingual performance needs careful review.
- Reference-voice continuity was poor through our test environment.
- High parallel submission triggered upstream concurrency errors, although failed tasks were not billed.
Who should use Seed Audio 1.0?
Seed Audio 1.0 is a strong fit for audio creators who think in scenes rather than isolated assets. It is particularly useful for radio drama, game dialogue, cinematic prototypes, audio storyboards, podcast concepts, and short suspense cues.
It is less convincing as a one-click final-delivery system. If exact wording, voice identity, or event counts are legally or creatively critical, plan for repeated generations and human review. The model works best when you treat it as a fast audio director with several takes—not a deterministic renderer.
Perguntas frequentes
Is Seed Audio 1.0 a text-to-speech model?
Can Seed Audio 1.0 generate sound effects without speech?
Can Seed Audio 1.0 generate music?
How long can Seed Audio 1.0 generate?
Does Seed Audio 1.0 support reference audio?
How much does Seed Audio 1.0 cost?
Veredicto final
Seed Audio 1.0 is a genuinely interesting unified audio model. The ability to produce convincing speech, Foley, ambience, and music from one prompt is not just a launch claim: our mixed-scene and music tests showed that the pieces can work together.
Its best feature was dialogue timing. Its biggest weakness was consistency. Across 23 samples, the model repeatedly showed an excellent best case and a noticeably weaker alternate take.
For creative prototyping, that tradeoff is easy to accept. Generate several versions, keep the strongest take, and inspect every ending. For deterministic production, exact voice matching, or event-count-sensitive work, keep a human review and editing stage in the workflow.
Overall verdict: Seed Audio 1.0 is one of the most capable scene-level audio generators we have tested, but it is best used as a multi-take creative tool rather than a one-shot final renderer.



