Inference APIs

Which text-to-speech formats and voices are available?

Output is mp3 (default), wav or raw PCM. Kokoro 82M has 54 voices across eight languages (ids like af_heart, bm_george, jf_alpha); Orpheus has eight expressive English voices. GET /v1/voices?model=hexgrad/Kokoro-82M returns the list.

Related

More questions