Inference APIs
Blog/pricing

How we price: upstream cost plus a fixed margin, published per model

Every model is priced at about 30% over what it costs us to serve, and the rate is on the model page. What that means in practice, and why we do not have tiers.

Taylor Hawkes · September 16, 2026

Inference pricing is usually opaque: a per-token number with no relation to cost, tiers that change the number when you spend more, and free tiers that stop working when you need them. We do it differently and would rather explain it than have you guess.

The rule

Each model's price is its upstream cost to us plus roughly 30%. Chat models are priced per million tokens, input and output separately; text-to-speech per million characters; transcription per audio minute. When an upstream cost changes, the price changes with it and the changelog says so.

ModelWhat we chargeUnit
DeepSeek V4 Flash$0.19 / $0.38per 1M tokens in / out
GPT-OSS 120B$0.20 / $0.80per 1M tokens in / out
GLM 5.3 Flash$0.20 / $0.66per 1M tokens in / out
DeepSeek V4.1 Flash$0.40 / $1.60per 1M tokens in / out
Llama 3.3 70B Turbo$1.35 / $1.35per 1M tokens in / out
Kokoro 82M$5.20per 1M characters
Orpheus 3B$19.50per 1M characters
Whisper Large v3$0.002per audio minute
Parakeet TDT 0.6B v3$0.002per audio minute

Why no tiers

Tiers exist to ration capacity and to push spend. We rent capacity elastically, so rationing is not needed, and a metered price already scales with use. The practical consequence: there is no daily cap to outgrow and no upgrade to be approved for — the thing that sent a lot of people here from Groq in August.

What "cost plus 30%" buys you

The margin pays for the gateway, metering, the playgrounds, the docs and reference, and support answered by a person. It does not pay for a sales team, which is why there is no enterprise tier either. If your volume is large enough that 30% is a lot of money, talk to us; a lower margin on committed volume is a conversation we are happy to have.

All rates: pricing.

More posts