Inference APIs

Chat

Text generation, reasoning, assistants and agents

GPT-OSS 120B

Available Popular
by OpenAI

OpenAI's open-weight 120B reasoning model. Strong at coding, tool use and agentic tasks; the drop-in replacement Groq recommends for Llama 3.3 70B.

ChatReasoningTool callingJSON mode131K context
Context
131K tokens
Modality
Text → Text
Input
$0.20 / 1M tokens
Output
$0.80 / 1M tokens
Model ID
openai/gpt-oss-120b

DeepSeek V4 Flash

Available Best value
by DeepSeek

DeepSeek's fast, inexpensive V4 model with a 1M-token context. Strong general assistant and coding model at the lowest price per token here.

ChatReasoningCodingTool calling1M context
Context
1M tokens
Modality
Text → Text
Input
$0.19 / 1M tokens
Output
$0.38 / 1M tokens
Model ID
deepseek-ai/DeepSeek-V4-Flash

The newest DeepSeek Flash model: better reasoning and tool use than V4 Flash, still with a 1M-token context.

ChatReasoningTool callingJSON mode1M context
Context
1M tokens
Modality
Text → Text
Input
$0.40 / 1M tokens
Output
$1.60 / 1M tokens
Model ID
deepseek-ai/DeepSeek-V4.1-Flash

Z.ai's fast GLM 5.3 model with tool calling, JSON mode and a 1M-token context, at a low per-token price.

ChatReasoningTool callingJSON mode1M context
Context
1M tokens
Modality
Text → Text
Input
$0.20 / 1M tokens
Output
$0.66 / 1M tokens
Model ID
zai-org/GLM-5.3-Flash

Meta's instruction-tuned 70B model. Still served here after its retirement on Groq — the same <code>llama-3.3-70b-versatile</code> id works unchanged.

ChatTool callingJSON modeMultilingual
Context
131K tokens
Modality
Text → Text
Input
$1.35 / 1M tokens
Output
$1.35 / 1M tokens
Model ID
meta-llama/Llama-3.3-70B-Instruct-Turbo

GPT-OSS 20B

Coming soon
by OpenAI

The small GPT-OSS model: fast and inexpensive for lightweight assistants and classification.

ChatReasoningTool calling
Context
131K tokens
Modality
Text → Text
Input
$0.07 / 1M tokens
Output
$0.27 / 1M tokens
Model ID
openai/gpt-oss-20b

Qwen3-VL 8B

Coming soon
by Alibaba Qwen

Compact vision-language model for image understanding, OCR and screenshot reasoning.

VisionChatOCRTool calling
Context
262K tokens
Modality
Image + Text → Text
Input
$0.24 / 1M tokens
Output
$0.90 / 1M tokens
Model ID
Qwen/Qwen3-VL-8B-Instruct

Text to speech

Natural voices on the OpenAI /v1/audio/speech endpoint

Kokoro 82M

Available Popular
by Hexgrad

Fast, natural open-weight text-to-speech with 54 voices across 8 languages, on the OpenAI <code>/v1/audio/speech</code> endpoint.

Text to speech54 voicesMP3 / WAV8 languages
Modality
Text → Audio
Price
$5.2 / 1M characters
Model ID
hexgrad/Kokoro-82M

Orpheus 3B

Available
by Canopy Labs

Expressive Llama-based text-to-speech with emotive tags and eight English voices.

Text to speechExpressive8 voicesEnglish
Modality
Text → Audio
Price
$19.5 / 1M characters
Model ID
canopylabs/orpheus-3b-0.1-ft

Speech to text

Transcription with timestamps on /v1/audio/transcriptions

Whisper Large v3

Available Popular
by OpenAI

OpenAI's multilingual speech recognition model on the standard <code>/v1/audio/transcriptions</code> endpoint — same request shape as OpenAI and Groq.

Speech to text99 languagesTimestampsSRT / VTT
Modality
Audio → Text
Price
$0.002 / audio minute
Model ID
openai/whisper-large-v3

Parakeet TDT 0.6B v3

Available Fastest
by NVIDIA

NVIDIA's very fast English-first transcription model with accurate timestamps; ideal for long recordings and batch jobs.

Speech to textVery fastTimestamps25 languages
Modality
Audio → Text
Price
$0.002 / audio minute
Model ID
nvidia/parakeet-tdt-0.6b-v3

Vision

Image understanding and screen parsing

OmniParser V2

Unavailable
by Microsoft

Parse screenshots into structured, labeled UI elements for computer-use agents.

VisionUI element detectionOCR
Modality
Image → Structured data
Price
$0.003 / request
Model ID
omniparser2