Voxtral Small 24B vs Whisper Large v3
Two open-weight speech-to-text models on the same endpoint and API key, so switching between them is a one-word change. Below: what each costs for three realistic workloads, how they differ, and measured speed.
Inference APIs prices are read live from the price list.
- Cost: Whisper Large v3 is cheaper in all three workloads, by 50%.
Side by side
| Voxtral Small 24B | Whisper Large v3 | |
|---|---|---|
| Served by | Inference APIs | Inference APIs |
| Model author | Mistral AI | OpenAI |
| Model id | mistralai/Voxtral-Small-24B-2507 | openai/whisper-large-v3 |
| Price | $0.004 / audio minute | $0.002 / audio minute |
| Capabilities | Speech to text, 8 languages, Auto language, Text only | Speech to text, 100 languages, Word timestamps, Speaker diarization, SRT / VTT |
| Weights | Open | Open |
| Status | Available | Available |
| Limits | No per-minute or per-day request or token caps; usage draws on a prepaid balance | No per-minute or per-day request or token caps; usage draws on a prepaid balance |
| Speed | 0.09× real time (23.1 s clip in 2.08 s) | 0.047× real time (19.5 s clip in 0.92 s) |
What three workloads cost
List prices applied to the same work. The cheaper side of each row is in bold.
| Workload | Voxtral Small 24B | Whisper Large v3 | Difference |
|---|---|---|---|
| 1 hour of audio | $0.24 | $0.12 | 50% |
| 100 hours of audio | $24.00 | $12.00 | 50% |
| 1,000 hours of audio | $240 | $120 | 50% |
Whisper Large v3 is cheaper in all three workloads, by 50%.
When to choose which
- Accuracy on messy real-world audio in English, Spanish, French, Portuguese, Hindi, German, Dutch or Italian
- You want Mistral's Voxtral without running a 24B model yourself
- You need live streaming transcription
- Cost is the deciding factor; it is cheaper in every workload above
- Broad language coverage with automatic language detection
- Files up to 100 MB and SRT or VTT subtitles straight from the API
- Speaker labels with diarize=true, at the same price
- You need timestamps or subtitles (use Whisper or Parakeet)
- Languages outside its eight; use Whisper
- Speaker diarization
Measured speed
Median of three runs on 2026-09-16: a 23.1-second mp3 clip, timed from upload to response. Expect ±30% with time of day. Method and raw numbers: speed measurements.
Trying both
Same endpoint, same key. Change one string:
for MODEL in "mistralai/Voxtral-Small-24B-2507" "openai/whisper-large-v3"; do
curl -s https://api.inferenceapis.com/v1/audio/transcriptions \
-H "Authorization: Bearer $INFERENCE_API_KEY" \
-F "model=$MODEL" -F "file=@meeting.mp3" | jq -r .text
doneFAQ
Which is cheaper, Voxtral Small 24B or Whisper Large v3?
Whisper Large v3 is cheaper in all three workloads, by 50%. The table above uses list prices: Voxtral Small 24B $0.004 / audio minute; Whisper Large v3 $0.002 / audio minute.
How do I switch between them?
They are on the same endpoint and key. Change the model field from mistralai/Voxtral-Small-24B-2507 to openai/whisper-large-v3 and nothing else.
Spotted a price or limit that has changed? Tell us and we will re-check the provider page.
