Inference APIs

Text-to-speech by language

Kokoro 82M ships 54 voices across eight languages; Orpheus adds expressive English. Pick a language for its voices and a ready-to-run request.

How language selection works

There is no language parameter on the speech endpoint. The language is decided by the voice you pick: each Kokoro voice id starts with a two-letter prefix (a American English, b British English, j Japanese, z Mandarin, e Spanish, f French, h Hindi, i Italian, p Brazilian Portuguese) followed by f or m for the voice's gender. Send text in the voice's language; mixed-script input is read with that language's pronunciation rules.

Orpheus 3B is English-only and offers eight expressive voices with emotion tags. Languages not listed here (German, Korean, Arabic, Russian and others) have no Kokoro voice yet; if you need one, tell us which. Every language on this page is also supported in the other direction by the speech-to-text API.

All voices bill at the same rate, $5.20 per million characters for Kokoro and $19.50 for Orpheus, return mp3, wav or raw PCM, and are listed live by GET /v1/voices. Try any of them in the speech playground.