Fish Audio S2.1 Pro is available — create expressive voices now.Try Fish Audio
Product guide

Fish Audio Text to Speech Real-Time AI TTS

Generate speech from text with Fish Audio: 30+ languages, emotion control, low-latency streaming, and voice cloning from a 15-second clip.

2026/09/16

What Fish Audio text to speech does

Fish Audio turns written text into spoken audio. You paste a script into the workspace, pick a voice, and generate a file you can download and use. The model behind it is built for expressiveness rather than flat narration: it carries tone, emphasis, and pacing into the output instead of reading every sentence the same way.

Generation runs in real time. Short lines come back quickly enough to keep a session moving, which is what makes the tool usable for iteration: you hear a take, change a phrase, and hear it again. The same engine powers the on-page demo, the workspace, and the API, so what you hear while testing is what you get in production.

The workflow is deliberately short. There is no project setup, no voice training step, and no file format decisions to make before you can hear something. If you want to compare this against the rest of the product first, the voice cloning and studio guides cover the other two ways people use the same audio engine.

Emotion control and pacing

The difference between robotic text to speech and narration you would actually publish is control. Fish Audio exposes that in two ways.

The first is emotion tags: short markers inside the script that tell the voice how a line should land. The second is natural language descriptions, where you describe the delivery you want and let the model interpret it. Both are written into the script itself, so a take can be re-generated later with the same direction instead of being re-recorded from memory.

Multispeaker output is part of the same system. Dialogue between two voices is generated as one piece of audio rather than two files stitched together, which keeps interruptions and pauses sounding natural.

Languages and voices

Fish Audio speaks 30+ languages, including English, Japanese, Korean, Chinese, French, German, Arabic, and Spanish. A voice is not locked to the language it was recorded in: the same voice can carry a script in another language, which is what makes a series sound consistent across markets instead of like a different narrator in every locale.

Beneath that sits the voice library, which hosts more than 2,000,000 voices. You can browse it, shortlist a few, and hear them on your own script before committing to a direction. The voice library guide explains how to narrow that down quickly.

From script to finished audio

A practical pass looks like this:

  1. Write the script the way it should be spoken, with punctuation where you want pauses.
  2. Add emotion tags where the delivery changes.
  3. Choose a voice from the library, or clone one from a 15 second recording.
  4. Generate, listen, and revise the script rather than the settings.
  5. Download the finished audio and keep the script, so future revisions cost one generation instead of a rebuild.

Longer projects are best split into paragraphs or chapters. It keeps each generation small, makes a bad line cheap to redo, and keeps the credit spend visible while you work. The Studio guide covers that pattern in more detail.

Frequently asked questions

How much does text to speech cost?

Generation is paid for with credits: one credit covers 1,000 characters. Packs start at $9 for 1,000 credits and never renew. The pricing page lists all three packs and what each one includes.

Can I use the generated audio commercially?

Commercial usage rights come with the Creator and Studio packs. Personal projects are covered by any pack, including the $9 Starter pack.

Do I need an API to generate speech?

No. The workspace generates audio from a text box. API access is included with the Studio pack for teams that want to call the same engine from their own product.

Fish Audio Text to Speech | Real-Time AI TTS