Explore the Fish Audio voice library — narration, support, and character voices across 30+ languages — and learn how to match a voice to a script.
The Fish Audio library hosts more than 2,000,000 voices. That number sounds overwhelming until you treat it as what it is: a catalogue you filter, not a list you read. Voices are grouped by the job they do, so most projects resolve in a few minutes rather than a few hundred auditions.
Three categories cover most of the work:
Every voice works across 30+ languages, so picking one is not a decision about which market you can serve.
The fastest way to choose is to audition with your own text rather than a demo paragraph. A voice that sounds impressive reading a sample can sound wrong reading your script, because pace and phrasing differ between genres.
Two rules of thumb help. First, match the delivery to the audience: a patient, reassuring, knowledgeable read suits support and health content, while a warm, immersive read suits narrative. Second, listen for pace rather than tone. If the voice rushes a sentence you wrote as a pause, the script and the voice disagree about rhythm, and a different voice will fix it faster than rewriting.
The library is a starting point, not a limit. If the exact voice you need does not exist — a specific presenter, a character from your own production, or the voice already in your archive — voice cloning builds it from a short recording.
Consistency is what makes a series feel like a series. The same narrator across every episode, the same assistant across every demo, the same character across every scene.
To hold that line, settle the voice before you produce volume. Generate a full paragraph, listen to it end to end, and only then commit to the rest of the script. If a later line goes wrong, regenerate the line rather than the voice, and keep the script as the single source of truth so a re-generation months later still matches.
When a project grows, the pattern is the same as for long-form work: split the script into paragraphs or chapters, keep each generation small, and regenerate only what changed. The Studio guide walks through that workflow.
A voice is not tied to the language it was recorded in. The same voice can read a script in a different language, which keeps a brand or character recognisable across markets instead of introducing a new narrator for each locale.
Pronunciation of names, brands, and technical terms is the one thing worth checking early. Read the script aloud in your head before generating: if a term is ambiguous, spell it the way it should sound, and re-generate that line rather than the whole file. It is a small habit that saves a lot of editing later.
Text to speech itself is covered in more detail in the text to speech guide, and everything here runs on the same engine, priced the same way: one credit per 1,000 characters, with packs starting at $9 on the pricing page.
The library hosts more than 2,000,000 voices, across narration, support, and character styles.
Yes. Voices work across the 30+ supported languages, so one voice can serve every locale in a project.
Clone it. A 15 second recording is enough to build a voice you can reuse, and cloning is included in every credit pack.