What ElevenLabs is best at
Five speech models, one picker
Flash v2.5 is the default — roughly 75ms latency across 32 languages. ElevenLabs v3 is the expressive one, Multilingual v2 the high-fidelity 29-language option, Turbo v2.5 the legacy pin. Text to Dialogue renders up to 10 speakers as one file.
Voices you can actually find
Search the ElevenLabs Voice Library by language, gender, accent, age, use case and style, design a brand-new voice from a written description, or use a voice cloned from your own sample.
Music v2, up to ten minutes
Write a prompt, or hand it a composition plan with per-section lyrics and durations. Store a song for inpainting and regenerate one chunk instead of the whole track.
Sound effects and transcription
Sound effects from half a second to 30 seconds, loopable for beds. Scribe v2 transcribes audio, video or a YouTube URL in 90+ languages with word timestamps and speaker labels.
ElevenLabs at a glance
Voice, music, and sound prompts
Built to ElevenLabs’s strengths — steal these to get started.
How to use ElevenLabs on Genfire
- 1
Describe it
Write a prompt, and pick ElevenLabs in the model selector.
- 2
Generate & compare
Render the track, iterate on the prompt, or run the same brief on a sibling model to compare — all from one credit balance.
- 3
Finish & publish
Send the result into Genfire's editor, upscaler, and audio tools, then export platform-ready output.
Everything you can hear, on one balance
The five speech models are not a menu of near-identical options; they trade against each other in ways you can feel. Flash v2.5 is the default because it is the one you want most of the time — around 75ms of latency, 32 languages, and a 40,000-character ceiling, which is enough for a long-form narration in a single call. ElevenLabs v3 is the expressive model, the one that carries a performance rather than a read, and its 5,000-character limit reflects that: you send it a scene, not a chapter. Multilingual v2 sits between them at 29 languages and 10,000 characters when fidelity matters more than speed. Turbo v2.5 is deprecated upstream and kept pinnable only so old work still reproduces — Flash v2.5 replaces it at the same price and lower latency. Text to Dialogue is the odd one out and the most useful surprise: send an ordered array of lines with a voice on each, up to 10 distinct voices and 2,000 characters, and you get back a single file where the speakers actually respond to each other, with matched prosody across the cut and per-word timings if you ask for them. Lines take v3 audio tags like [whispers] and [laughs].
Finding the right voice is usually the harder half of the job, so the picker treats it as a real search. The Library tab queries the ElevenLabs Voice Library with facets for language, gender, accent, age, use case and style, and previews play one at a time so you can audition down a column. The Designed tab does the other thing: describe a voice in words — age, accent, texture, pace — and ElevenLabs returns several previews to choose from before you save one permanently to your account. Cloned voices live in the same tab, and once saved, any of these voices behaves identically everywhere else in Genfire. Music works the same way at a larger scale. Music v2 takes a prompt and a length anywhere from three seconds to ten minutes, or a composition plan: chunks with their own text, section markers, per-chunk styles and durations. Ask it to store the song and you get an id back, which lets a later plan regenerate a single range instead of rerolling the whole track. It will also score a video you hand it instead of a prompt. Genfire's Music Video Studio leans on exactly that structure — the section boundaries in the plan are what the cuts get timed to.
The point of putting all of this behind one balance is that audio in Genfire is rarely the finished product. A voice you pick here narrates faceless reels and explainers, becomes an influencer model's standing voice, and gets attached to a video in the editor. The speech you generate is the audio track a lip-sync engine drives a performance from. Voices are reusable as references too: hand a short clip to a video model that conditions on audio and an on-camera character speaks in that voice. Sound effects run from half a second to 30 seconds and can be generated loopable, which is what you want for a room tone or an engine bed rather than a one-shot. Scribe v2 closes the loop from the other end — audio, video or a YouTube URL in, transcript out, in 90+ languages, with a speaker id on every word for up to 32 speakers, keyterm biasing for names your transcript keeps getting wrong, and entity detection or redaction for PII. All of it runs in the browser studio, from the GenBar, as a node in the workflow editor, on the REST API, through the MCP server at mcp.genfire.ai, and from the @genfire/cli — no ElevenLabs account or separate subscription. Genfire is pay-as-you-go: credit packs start at $19, purchased credits last 12 months, optional monthly plans start at $29, and every purchase includes a commercial license.
Loved by creators
“I use GENFIRE for storyboarding and animatics. It communicates my vision perfectly to clients.”
“We scaled our TikTok output from 3 to 15 videos a week without hiring more editors.”
“The AI lip-sync for translations allowed us to enter 3 new markets flawlessly.”
ElevenLabs — frequently asked questions
What is ElevenLabs, and what does it do on Genfire?
ElevenLabs is a voice and audio AI company. On Genfire its models cover five jobs from one credit balance: text to speech across five models and dozens of languages, multi-speaker dialogue in a single file, music generation up to ten minutes with Music v2, sound effects up to 30 seconds, and speech-to-text with Scribe v2.
Which ElevenLabs speech model should I use?
Flash v2.5 is the default and the right answer most of the time — about 75ms latency, 32 languages, and the lowest cost. Choose ElevenLabs v3 when you need an expressive performance rather than a clean read, and Multilingual v2 when you want high fidelity across 29 languages. Turbo v2.5 is deprecated upstream and only worth pinning to reproduce older work.
How much text can I send in one request?
It depends on the model: up to 40,000 characters on Flash v2.5 and Turbo v2.5, 10,000 on Multilingual v2, and 5,000 on ElevenLabs v3. Text to Dialogue caps at 2,000 characters across all of its lines.
Can I clone my own voice, or design a new one?
Both. Voice design is an ElevenLabs feature: describe a voice in words and it returns several previews, and you save the one you like permanently to your account. Cloning from an audio sample runs on Genfire's own cloning model rather than through ElevenLabs, but the cloned voice lands in the same picker and works everywhere an ElevenLabs voice does — the studio, faceless reels, influencer models, and the API.
Can two or more speakers talk to each other in one file?
Yes. Text to Dialogue takes an ordered list of lines, each with its own voice id — up to 10 distinct voices and 2,000 characters total — and returns one audio file with matched prosody between lines. Lines accept v3 audio tags such as [whispers] or [laughs], and you can request per-word timings and per-line voice segments.
How long can an ElevenLabs music track be, and can I regenerate part of it?
Music v2 generates anywhere from 3 seconds to 10 minutes. Ask it to store the song for inpainting and you get a song id back, which a later composition plan can reference to regenerate a single range instead of the whole track. Music v2 can also score a video you supply instead of writing to a prompt.
Can Genfire transcribe audio with ElevenLabs Scribe v2?
Yes. Scribe v2 accepts an audio URL, a video URL or a YouTube URL and returns a transcript in 90+ languages with word-level timestamps, speaker diarization for up to 32 speakers, keyterm biasing for up to 1,000 terms, filler removal, and entity detection or redaction for categories such as PII and PHI.
Compare with other models
From the blog
All posts →Explore Genfire
AI Tools
AI Models
Generate with ElevenLabs today
ElevenLabs plus 65+ other AI models, one credit balance, and a full editing pipeline — free plan included, no credit card required.
Start Free