Inference APIs
Compare/Speech-to-text

Voxtral Small 24B vs Whisper Large v3

Two open-weight speech-to-text models on the same endpoint and API key, so switching between them is a one-word change. Below: what each costs for three realistic workloads, how they differ, and measured speed.

Inference APIs prices are read live from the price list.

Short answer
  • Cost: Whisper Large v3 is cheaper in all three workloads, by 50%.

Side by side

Voxtral Small 24BWhisper Large v3
Served byInference APIsInference APIs
Model authorMistral AIOpenAI
Model idmistralai/Voxtral-Small-24B-2507openai/whisper-large-v3
Price$0.004 / audio minute$0.002 / audio minute
CapabilitiesSpeech to text, 8 languages, Auto language, Text onlySpeech to text, 100 languages, Word timestamps, Speaker diarization, SRT / VTT
WeightsOpenOpen
StatusAvailableAvailable
LimitsNo per-minute or per-day request or token caps; usage draws on a prepaid balanceNo per-minute or per-day request or token caps; usage draws on a prepaid balance
Speed0.09× real time (23.1 s clip in 2.08 s)0.047× real time (19.5 s clip in 0.92 s)

What three workloads cost

List prices applied to the same work. The cheaper side of each row is in bold.

WorkloadVoxtral Small 24BWhisper Large v3Difference
1 hour of audio$0.24$0.1250%
100 hours of audio$24.00$12.0050%
1,000 hours of audio$240$12050%

Whisper Large v3 is cheaper in all three workloads, by 50%.

When to choose which

Choose Voxtral Small 24B if
  • Accuracy on messy real-world audio in English, Spanish, French, Portuguese, Hindi, German, Dutch or Italian
  • You want Mistral's Voxtral without running a 24B model yourself
  • You need live streaming transcription
Choose Whisper Large v3 if
  • Cost is the deciding factor; it is cheaper in every workload above
  • Broad language coverage with automatic language detection
  • Files up to 100 MB and SRT or VTT subtitles straight from the API
  • Speaker labels with diarize=true, at the same price
  • You need timestamps or subtitles (use Whisper or Parakeet)
  • Languages outside its eight; use Whisper
  • Speaker diarization

Measured speed

Median of three runs on 2026-09-16: a 23.1-second mp3 clip, timed from upload to response. Expect ±30% with time of day. Method and raw numbers: speed measurements.

Trying both

Same endpoint, same key. Change one string:

cURL
for MODEL in "mistralai/Voxtral-Small-24B-2507" "openai/whisper-large-v3"; do
  curl -s https://api.inferenceapis.com/v1/audio/transcriptions \
    -H "Authorization: Bearer $INFERENCE_API_KEY" \
    -F "model=$MODEL" -F "file=@meeting.mp3" | jq -r .text
done

FAQ

Which is cheaper, Voxtral Small 24B or Whisper Large v3?

Whisper Large v3 is cheaper in all three workloads, by 50%. The table above uses list prices: Voxtral Small 24B $0.004 / audio minute; Whisper Large v3 $0.002 / audio minute.

How do I switch between them?

They are on the same endpoint and key. Change the model field from mistralai/Voxtral-Small-24B-2507 to openai/whisper-large-v3 and nothing else.

Spotted a price or limit that has changed? Tell us and we will re-check the provider page.