Inference APIs

Chat models USD per 1M tokens

ModelModel IDInputOutput
GPT-OSS 120B Popular openai/gpt-oss-120b $0.20 $0.80
DeepSeek V4 Flash Best value deepseek-ai/DeepSeek-V4-Flash $0.19 $0.38
DeepSeek V4.1 Flash deepseek-ai/DeepSeek-V4.1-Flash $0.40 $1.60
GLM 5.3 Flash zai-org/GLM-5.3-Flash $0.20 $0.66
Llama 3.3 70B Instruct meta-llama/Llama-3.3-70B-Instruct-Turbo $1.35 $1.35
GPT-OSS 20B openai/gpt-oss-20b $0.07 $0.27 Coming soon
Qwen3-VL 8B Qwen/Qwen3-VL-8B-Instruct $0.24 $0.90 Coming soon

Text to speech USD per 1M characters

ModelModel IDPrice
Kokoro 82M Popular hexgrad/Kokoro-82M $5.20
Orpheus 3B canopylabs/orpheus-3b-0.1-ft $19.50

Speech to text USD per audio minute

ModelModel IDPrice
Whisper Large v3 Popular openai/whisper-large-v3 $0.002
Parakeet TDT 0.6B v3 Fastest nvidia/parakeet-tdt-0.6b-v3 $0.002

How billing works

Prepaid credits

Add funds with a card; each request deducts its exact cost from your balance. New accounts start with free credit so you can test every model.

Metered per model

Chat is billed on prompt + completion tokens (reasoning tokens included), speech on input characters, transcription on audio minutes measured from your file.

No tiers, no caps

There is no free-tier ceiling to outgrow and no upgrade to be approved for. Rate limits exist only to protect the service and are far above typical use.

Prices are in USD and may change as upstream costs change; the model page always shows the current rate. Questions? Contact us from your account.