What Seed Audio 1.0 is best at
Twenty preset voices
Multilingual presets spanning English, Chinese, Japanese, Spanish, Indonesian and Portuguese, each with a preview. The voice is optional — omit it and the model narrates in a voice it chooses.
Whole audio scenes
Not just a narrator reading a line. Describe a conversation, an ambience, or a scripted radio piece and Seed Audio renders it as one continuous take.
Reference conditioning
Attach up to three reference audio clips and cite them in the prompt as @Audio1 to @Audio3, or attach a single reference image for the model to voice — one or the other, not both.
Delivery controls
Speed from 0.5x to 2x, pitch across two octaves in semitones, volume, and four output formats from lossy MP3 to raw PCM at up to 48 kHz.
Seed Audio 1.0 at a glance
Speech and scene prompts
Built to Seed Audio 1.0’s strengths — steal these to get started.
How to use Seed Audio 1.0 on Genfire
- 1
Describe it
Write a prompt, and pick Seed Audio 1.0 in the model selector.
- 2
Generate & compare
Render the track, iterate on the prompt, or run the same brief on a sibling model to compare — all from one credit balance.
- 3
Finish & publish
Send the result into Genfire's editor, upscaler, and audio tools, then export platform-ready output.
A scene, not just a narrator
Most text-to-speech models take a sentence and give you that sentence, spoken. Seed Audio 1.0 is built to take a description of audio and give you that audio, which turns out to be a materially different tool. Prompt it with a line of copy and a preset voice and it behaves like a conventional TTS engine, cleanly and multilingually. Prompt it with a situation — a two-hander argument that resolves, an ambience bed with a voice buried in it, a scripted radio piece with a door and footsteps — and it renders the whole thing in one continuous take rather than making you generate three assets and mix them. That makes it the fastest route in Genfire from an idea for a moment of audio to a file you can drop on a timeline, and a genuinely different instrument from ElevenLabs, which is the better choice when you need one specific voice reading exactly your script.
The controls are deliberately small, and each one earns its place. The voice field is optional: pick one of 20 presets — Vivi, Mindy, Kian, Sophie, Magnus and the rest, each with a preview and its own language mix across English, Chinese, Japanese, Spanish, Indonesian and Portuguese — or leave it empty and let the model choose a narrator to fit the prompt. Note that these are Seed preset names, not ElevenLabs voice ids: a cloned or designed voice from elsewhere in Genfire does not apply here, and the picker only offers you what Seed actually accepts. References come in two mutually exclusive shapes. Up to three audio clips can be attached and cited directly in the prompt as @Audio1 through @Audio3, so you can say what to do with each one instead of hoping the model infers it. Or a single image can be attached instead, and the model voices what it sees. The prompt itself caps at 2,048 characters, which is a scene rather than a chapter, and billing is per second of audio produced.
On the output side there are more knobs than most speech models expose. Speed runs from 0.5x to 2x in fixed steps, pitch shifts up or down by as much as twelve semitones, and volume is adjustable — enough to fit a read to a cut without re-rolling it. Formats cover MP3 for the common case, WAV when the editor should not touch a lossy file, OGG Opus for the web, and raw PCM at sample rates from 8 kHz up to 48 kHz when something downstream wants the unwrapped stream. Seed Audio appears in the audio studio next to the ElevenLabs speech models, in the GenBar for a quick one-off, as a model on the audio node in the workflow editor, and on the REST API as speech.seed_audio_1_0, through the MCP server at mcp.genfire.ai, and from the @genfire/cli. Everything it makes lands in your library alongside the rest of your audio, ready for the video editor and the clip workflows. Genfire is pay-as-you-go: credit packs start at $19, purchased credits last 12 months, optional monthly plans start at $29, and every purchase includes a commercial license.
Loved by creators
“My agency saves 40+ hours per week using GENFIRE. ROI was instant. Absolutely essential tool.”
“The consistency in character generation is miles ahead of anything else I've used.”
“Finally, an AI tool that actually understands cinematic lighting. My portfolio looks amazing.”
Seed Audio 1.0 — frequently asked questions
What is Seed Audio 1.0?
Seed Audio 1.0 is ByteDance's general text-to-audio model. It does natural multilingual text to speech with preset voices, and it also generates whole audio scenes — dialogue, ambience, radio-drama style pieces — from a single prompt.
Which voices can Seed Audio use?
Twenty Seed presets, covering English, Chinese, Japanese, Spanish, Indonesian and Portuguese in various combinations, each with a preview in the picker. The voice is optional: leave it empty and the model narrates in a voice of its own choosing. Seed presets are their own thing — an ElevenLabs voice id or a cloned voice cannot be used with this model.
Can Seed Audio generate more than one speaker?
Yes, as part of a scene. Describe a conversation in the prompt and Seed Audio renders it as one continuous take with the speakers in it. It does not take a per-line voice assignment the way ElevenLabs Text to Dialogue does — you describe the scene, rather than casting it line by line.
Can I give Seed Audio a reference?
Yes, in one of two ways. Attach up to three reference audio clips and refer to them in the prompt as @Audio1, @Audio2 and @Audio3, or attach a single reference image for the model to voice. A request can use reference audio or a reference image, but not both at once.
How long can a Seed Audio prompt be?
Up to 2,048 characters. Because the model is billed per second of audio it produces, longer prompts cost proportionally more — a short scene is a cheap experiment.
What audio formats does Seed Audio return?
MP3, WAV, OGG Opus, or raw PCM, at sample rates from 8 kHz to 48 kHz. You can also set the speed from 0.5x to 2x, shift the pitch by up to twelve semitones in either direction, and adjust the volume.
How do I use Seed Audio on Genfire, and what does it cost?
Generate it in the browser audio studio, from the GenBar, as a node in the workflow editor, on the REST API as speech.seed_audio_1_0, through the MCP server, or from the CLI — no ByteDance account needed. Genfire is pay-as-you-go: credits are billed per second of generated audio, credit packs start at $19, purchased credits last 12 months, and every purchase includes a commercial license.
Compare with other models
From the blog
All posts →Explore Genfire
AI Tools
AI Models
Generate with Seed Audio 1.0 today
Seed Audio 1.0 plus 65+ other AI models, one credit balance, and a full editing pipeline — free plan included, no credit card required.
Start Free