Seed Audio 1.0 리뷰: 음성, 효과음, 음악을 한 대의 기기로 모두 해결

seed audio 1.0 리뷰
SEED AUDIO / FIELD REVIEW
23 hands-on generations

One model can build the whole sound scene.

Seed Audio 1.0 can combine voice, Foley, ambience, timing, and instrumental music. Our tests found an unusually high creative ceiling—and a repeatability problem you should not ignore.

Best resultPrecise dialogue timing and coherent two-minute narration.
주요 위험Alternate takes can add garbled speech, miss cues, or change voice identity.

Seed Audio 1.0 is one of the first audio models I have tested that feels designed around a complete scene rather than a single output type. It can speak, act, add ambience, create Foley, place dialogue on a timeline, and produce instrumental music from one prompt.

The short verdict: Seed Audio 1.0 has an unusually high ceiling for unified audio creation, especially for timed dialogue and cinematic sound scenes. Its weakness is repeatability. A strong run can sound remarkably complete, while another run from the same prompt may add garbled speech, miss one event, or change the requested voice.

I generated 23 scored samples through the Anywhere/BrolyAI test environment. Together, they covered narration, two-character dialogue, timestamped speech, SFX-only sequences, voice-plus-music scenes, standalone music, multilingual performance, two-minute monologues, and reference-audio continuation.

Seed Audio 1.0 review: the quick verdict

Complete BytePlus post describing one-pass voice, sound effect, and music generation
BytePlus presented Seed Audio 1.0 as a one-pass system for voice, sound effects, and music.
카테고리Result from our tests
Expressive voiceHigh quality on the best run, but inconsistent across repeats
Multi-speaker dialogueAccurate scripts and strong speaker separation
Dialogue timingThe most reliable feature in our test set
음향 효과Realistic space and event order; weak at exact event counts
Voice + SFX + musicConvincing complete scenes when every requested event lands
Standalone musicMore capable than expected; produced structured 30-second cues
Multilingual voiceExcellent best case, but one of two runs hallucinated extra speech
Two-minute audioStrong voice identity and narrative coherence across both runs
Reference audioUnreliable through our test venue; voice similarity was poor

최적 대상: creators making radio drama, game scenes, podcast concepts, cinematic prototypes, and audio-first storyboards.

Less suitable for: jobs that require deterministic wording, exact repeated event counts, or dependable voice cloning from one unattended generation.

What is Seed Audio 1.0?

Official Seed Audio text-to-timbre benchmark comparison
Official Seed Audio material comparing text-to-timbre performance.

Seed Audio 1.0 is ByteDance Seed’s unified text-to-audio creation model. Instead of treating speech, ambience, effects, and music as separate jobs, it can generate them together as one sound scene. The official product page says it supports dialogue timing at 100-millisecond intervals and can generate up to two minutes in one pass.

That unified design is the real attraction. A prompt can describe who speaks, how they sound, what happens in the room, which effects enter later, and what kind of music sits under the scene. In a good run, the result sounds composed rather than stitched together.

There is an important boundary. ByteDance’s own roadmap says fine-grained timing currently focuses mainly on character dialogue. More precise control over sound effects, ambience, and music remains an area for further development. Our tests matched that distinction: dialogue timing was excellent, while exact Foley counts were less reliable.

Official sources checked August 6, 2026:

How we tested Seed Audio 1.0

23scored audio outputs
12.5 mingenerated audio reviewed
$1.7504planned upstream cost
120초longest single-pass test

We designed nine tasks and ran 23 scored generations. Most capability tests used two or three repeats, because one polished sample tells you very little about production reliability.

테스트작업Runs
T01Expressive English narration3
T02Two-character radio dialogue3
T03Dialogue at 0, 4, and 9 seconds3
T04SFX-only hallway sequence2
T05Voice, ambience, SFX, and music3
T06Standalone instrumental music3
T07One character across English, Japanese, and German2
T08120-second podcast monologue2
T09Continuation from a reference voice2

The 23 planned outputs contained about 12.5 minutes of generated audio and reported $1.7504 in upstream generation cost through the test environment. Seed Audio did not return text-token usage. Billing was based on generated seconds, so generation tokens are recorded as not applicable.

All scored outputs were 40 kHz, 16-bit, stereo PCM WAV files. When direct automated playback was not available, Gemini 3.6 Flash received the WAV files as audio input and returned structured transcripts, event checks, timing estimates, and artifact notes. Those judgments are automated listening evaluations, not human MOS scores.

Voice generation: impressive highs, uneven repeatability

T01 · Expressive narration

High variance
Repeated-run finding

One run hallucinated extra speech, one sounded stilted, and one was natural and complete.

Play sample · T01

사용된 프롬프트

One calm female narrator delivers two exact sentences, moving from concern to confidence. No music or SFX.

The narration prompt asked one calm female speaker to deliver two exact sentences, moving from restrained concern to growing confidence, with no background sound.

The three runs behaved very differently:

  • Run 1: delivered the requested sentence, then continued with several seconds of garbled, unscripted speech.
  • Run 2: spoke the exact text but sounded slow and staccato.
  • Run 3: delivered the complete script naturally, with the best pacing and voice consistency.

This is the central pattern of the review. Seed Audio 1.0 can produce a strong voice performance, but a single run is not enough for wording-critical work. Generate alternatives and listen through the entire tail before publishing.

The model did respect the negative instructions in all three narration runs: no music, no sound effects, and no unwanted background were detected.

Multi-character dialogue and speaker separation

T02 · Two-character radio drama

3/3 exact scripts
Repeated-run finding

All three scripts were exact. Room tone and lead-in length varied between takes.

Play sample · T02

사용된 프롬프트

Maya whispers, Eli answers calmly, and a refrigerator hum sits under three fixed dialogue turns.

Our convenience-store scene used two adult characters and three fixed lines. Maya had to whisper tensely, while Eli tried to sound calm. The only background requested was a subtle refrigerator hum.

All three runs delivered the dialogue in the correct order with two distinct voices. The best run combined exact lines, clean speaker separation, convincing emotion, and a continuous refrigerator hum.

The other runs showed smaller scene-level problems. One spent roughly the first half of its 30-second duration on ambience before beginning the dialogue. Another omitted the requested refrigerator hum. Neither mistake ruined the spoken scene, but both reduce editing efficiency.

For audio drama, the practical result is encouraging: Seed Audio understands character turns and vocal contrast. You may still need to trim long lead-ins or regenerate for the right room tone.

Dialogue timing was the strongest result

T03 · Prompt-level timing

가장 신뢰할 수 있는
Repeated-run finding

All three runs placed the lines near the requested points within the evaluator’s timing uncertainty.

Play sample · T03

사용된 프롬프트

Place three radio lines at 0.0, 4.0, and 9.0 seconds with male/female/male speakers.
Official Seed Audio dialogue timing controls with two fully loaded examples
Official documentation describes unified orchestration and dialogue timing control at 100-millisecond intervals.

The timing test requested three radio lines:

  1. “Unit seven, report.” at 0.0 seconds.
  2. “North entrance secure.” at 4.0 seconds.
  3. “Hold position.” at 9.0 seconds.

Across all three runs, the lines, order, and male-female-male speaker pattern were correct. Gemini’s audio estimates placed the repeated-run onsets around 0.1, 4.0, and 9.0 seconds. The first run was estimated at approximately 0.1, 3.9, and 8.9 seconds.

Those measurements should not be presented as laboratory-grade proof of 100-millisecond precision—the evaluator itself stated roughly 0.1-second uncertainty. Even with that caveat, this was the most repeatable capability we tested. If you need dialogue to enter near planned timeline positions, Seed Audio 1.0 is unusually promising.

SFX generation: realistic scenes, imperfect counting

T04 · SFX-only hallway

Counting failed
Repeated-run finding

The space and order were convincing, but the two runs produced four footsteps and two footsteps.

Play sample · T04

사용된 프롬프트

Keys, a heavy door, exactly three footsteps, thunder, and a final slam. No voice or music.

The SFX-only prompt described a small tiled hallway: keys near the microphone, a heavy wooden door opening, three slow footsteps moving away, distant thunder, and a final door slam. Speech and music were forbidden.

Both outputs created a believable acoustic space, preserved the event order, and avoided voice or music leakage. The effects were judged realistic, with good distance and room consistency.

Neither run followed the exact count. One produced four footsteps; the other produced two. That makes Seed Audio useful for generating a convincing Foley concept, but less dependable when an editor needs exactly three impacts, knocks, footsteps, or gunshots.

The best workflow is to use one-pass SFX generation for atmosphere and rapid prototyping, then replace or edit count-critical events in a DAW.

Voice, ambience, SFX, and music in one scene

T05 · Complete cinematic mix

2/3 complete
Repeated-run finding

Two of three runs completed every cue. One missed the final metallic slam.

Play sample · T05

사용된 프롬프트

Rain alley, traffic, siren, whispered detective, string pulse, fast footsteps, and metallic slam.
Community example discussing combined Seed Audio voice and Foley generation
A community example highlighting combined voice and Foley generation.

The complete-scene test was deliberately crowded. It asked for a wet night alley, water drips, traffic, a moving police siren, a whispered detective line, restrained strings, fast footsteps, and a metallic door slam.

Two of three runs included the complete event sequence. Dialogue remained clear over ambience and music, and the scene felt spatially coherent rather than like unrelated clips layered together. One run missed the final metallic slam, despite otherwise producing the right atmosphere and exact spoken line.

This is where the unified model makes the strongest case for itself. For storyboarding and creative ideation, generating an entire mixed scene from one prompt is faster and more expressive than sourcing every component separately. For final production, important end cues still need verification.

Can Seed Audio 1.0 generate music?

T06 · Standalone music

Music works
Repeated-run finding

All three created structured music. One made the muted piano difficult to distinguish.

Play sample · T06

사용된 프롬프트

A 72 BPM suspense cue with cello ostinato, muted piano, analog drone, build, and unresolved ending.
Official Seed Audio notes about music and timing limitations
Official notes frame music generation as capable but still subject to timing limits.

Yes. In our tests, Seed Audio 1.0 generated recognizable standalone instrumental music rather than only a vague ambient texture.

The prompt asked for a 30-second suspense cue at 72 BPM with low cello ostinato, muted piano pulses, a soft analog drone, a clear build, and an unresolved final chord. All three runs formed a coherent musical cue with the requested tempo feel and ending. Two included every requested element; in one run, the muted piano was difficult to distinguish.

This does not automatically make Seed Audio a replacement for a dedicated song generator. Our task was a short instrumental cinematic cue, not a verse-chorus song with lyrics. Still, the music capability is substantial enough to support game scenes, trailers, podcasts, and previsualization—especially when music must interact with speech and sound design in the same generation.

Multilingual performance: excellent best case, weak worst case

T07 · One voice, three languages

1/2 clean
Repeated-run finding

One run was flawless; the other repeated and hallucinated speech around the language switch.

Play sample · T07

사용된 프롬프트

One female character speaks English, Japanese, and German with the same identity and calm emotion.
Community report discussing Japanese output limitations in Seed Audio
A community report describing limitations in Japanese output.

The multilingual task asked one female character to speak the same recording-space role in English, Japanese, and German. Two runs produced opposite impressions.

The stronger run delivered all three sentences accurately, preserved the same voice, kept calm emotion, and used clean pauses. The weaker run started correctly in English, then repeated and hallucinated speech around the Japanese section and damaged the German opening. Its perceived cross-language voice consistency fell sharply.

This means Seed Audio 1.0 can achieve convincing multilingual character performance, but the feature needs multiple candidates and language-aware review. Do not approve a multilingual output by checking only the first language.

Two-minute generation and long-form voice stability

T08 · 120-second monologue

2/2 stable
Repeated-run finding

Both files reached 120 seconds with one stable voice. One had a brief pronunciation stumble.

Play sample · T08

사용된 프롬프트

One male podcast host tells a structured weather-station story with three discoveries and no background.
Official Seed Audio long-form stability and two-minute generation statement
Official documentation states that Seed Audio 1.0 can generate up to two minutes in one pass and supports continuation.

Both long-form tasks produced exact 120-second WAV files. Each used one male podcast host with dry studio sound, no music, and no sound effects.

The first run maintained one voice and a coherent investigative structure throughout: a clear opening, three chronological discoveries, a move from curiosity to unease, and a final unresolved question. No obvious voice drift or garbled segment was detected.

The second run was similarly coherent and stable, with one brief pronunciation stumble around 1:31. That is a strong result for a two-minute one-pass monologue. It suggests Seed Audio’s long-form claim is not merely a duration limit; the model can sustain identity and narrative structure across the full window.

Reference audio continuity needs caution

T09 · Reference continuation

Route uncertain
Repeated-run finding

Text was exact, but voice similarity was poor; one output changed to a male voice and added Foley.

Play sample · T09

사용된 프롬프트

Continue from an authorized female reference while preserving voice identity and intimate recording style.

The reference test used the strongest narration sample as an authorized input and asked for a continuation in the same general female identity and intimate recording style.

Both outputs spoke the requested continuation exactly, but neither matched the reference well. The first shifted into a breathy full whisper and received a conservative voice-similarity score of 2/5. The second was judged to use a male voice, scored 1/5 for similarity, and added roughly 17 seconds of unwanted Foley before speaking.

There is an important venue limitation here. The Anywhere task endpoint accepted the reference field, but it did not return confirmation showing how the upstream model consumed that reference. These results prove that reference continuity was unreliable through our test route; they do not prove that BytePlus’s native API always behaves the same way.

Seed Audio 1.0 pricing and limits

Official BytePlus Seed Audio pricing table
BytePlus lists pay-as-you-go Seed Audio pricing by generated duration.
Official BytePlus pricing
$0.15 / generated minute

Billed by the second, with 60 free trial minutes when the service is activated.

120초
maximum output
3,000
prompt characters
3
audio references
1
reference image

BytePlus lists the official pay-as-you-go price at $0.15 per generated minute, billed by the second, with 60 free trial minutes when the service is activated. The API documentation sets a 120-second maximum output, supports prompts up to 3,000 characters, allows up to three reference audio clips, and accepts one reference image.

Each reference audio clip can be up to 30 seconds and 10 MB. The reference image limit is 10 MB. These are BytePlus-native limits; third-party platforms may expose a smaller input form or use different pricing.

Our Anywhere/BrolyAI test route reported upstream cost separately and averaged about $0.14 per generated minute. That test-environment field should not be confused with BytePlus retail pricing or the final price of another platform.

Seed Audio 1.0 pros and cons

Where it wins

  • Unified voice, Foley, ambience, and music
  • Excellent dialogue timing
  • Strong long-form identity
  • Useful cinematic music

Where it breaks

  • Large take-to-take variance
  • Occasional garbled speech
  • Weak exact SFX counting
  • Unreliable reference continuity in our route

장점

  • Generates voice, ambience, Foley, and music in one pass.
  • Excellent dialogue timing in our repeated tests.
  • Strong two-speaker separation and exact script handling.
  • Produces useful standalone cinematic music.
  • Maintains voice identity and narrative coherence across two minutes.
  • Outputs clean 40 kHz stereo WAV through our test route.

단점

  • Repeat runs can vary dramatically in naturalness and reliability.
  • Occasional hallucinated or garbled speech.
  • Exact SFX counts are unreliable.
  • Some complete scenes miss a final requested event.
  • Multilingual performance needs careful review.
  • Reference-voice continuity was poor through our test environment.
  • High parallel submission triggered upstream concurrency errors, although failed tasks were not billed.

Who should use Seed Audio 1.0?

Seed Audio 1.0 is a strong fit for audio creators who think in scenes rather than isolated assets. It is particularly useful for radio drama, game dialogue, cinematic prototypes, audio storyboards, podcast concepts, and short suspense cues.

It is less convincing as a one-click final-delivery system. If exact wording, voice identity, or event counts are legally or creatively critical, plan for repeated generations and human review. The model works best when you treat it as a fast audio director with several takes—not a deterministic renderer.

자주 묻는 질문

Is Seed Audio 1.0 a text-to-speech model?
It includes text-to-speech, but it is broader than a conventional TTS model. One prompt can combine character dialogue, emotion, ambience, sound effects, timing instructions, and music. That makes it closer to a unified scene-generation model than a voice-only service.
Can Seed Audio 1.0 generate sound effects without speech?
Yes. Both of our SFX-only runs avoided speech and music while producing the requested hallway space and event sequence. However, neither followed the exact requested number of footsteps, so count-critical Foley may need editing or regeneration.
Can Seed Audio 1.0 generate music?
Yes. Three tests produced structured 30-second instrumental suspense cues with a stable tempo feel, identifiable instrumentation, a dynamic build, and an unresolved ending. It appears most useful for cinematic cues and mixed sound scenes rather than as a proven replacement for dedicated full-song models.
How long can Seed Audio 1.0 generate?
The official limit is 120 seconds in one pass. Both of our long-form tests returned exact two-minute WAV files with one stable narrator and coherent story structure. One contained a brief pronunciation stumble, but neither showed a speaker change.
Does Seed Audio 1.0 support reference audio?
The BytePlus API supports up to three reference clips. Our third-party test route accepted a reference input, but the two outputs matched the source voice poorly. Native API behavior may differ, so reference continuity should be verified in the exact platform used for production.
How much does Seed Audio 1.0 cost?
BytePlus lists $0.15 per generated minute and 60 free trial minutes at activation, checked August 6, 2026. Third-party services may charge different rates. Always separate official BytePlus pricing from platform-specific credits or margins.

최종 판결

Seed Audio 1.0 is a genuinely interesting unified audio model. The ability to produce convincing speech, Foley, ambience, and music from one prompt is not just a launch claim: our mixed-scene and music tests showed that the pieces can work together.

Its best feature was dialogue timing. Its biggest weakness was consistency. Across 23 samples, the model repeatedly showed an excellent best case and a noticeably weaker alternate take.

For creative prototyping, that tradeoff is easy to accept. Generate several versions, keep the strongest take, and inspect every ending. For deterministic production, exact voice matching, or event-count-sensitive work, keep a human review and editing stage in the workflow.

Overall verdict: Seed Audio 1.0 is one of the most capable scene-level audio generators we have tested, but it is best used as a multi-take creative tool rather than a one-shot final renderer.

게시물을 공유하세요:

관련 게시물