Voxtral Small 24B vs AssemblyAI Universal
Voxtral Small 24B is an open-weight speech-to-text model you can call on this endpoint; AssemblyAI Universal is a closed model from AssemblyAI. Below: what each costs for three realistic workloads, the limits that apply, and when to choose which.
AssemblyAI figures are public list prices and documented limits, checked 2026-09-16 against the provider page. Inference APIs prices are read live from the price list.
- Cost: AssemblyAI Universal is cheaper in all three workloads, by 38%.
- Limits: AssemblyAI: Concurrency limits by plan. Here: no per-minute or per-day caps.
Side by side
| Voxtral Small 24B | AssemblyAI Universal | |
|---|---|---|
| Served by | Inference APIs | AssemblyAI |
| Model author | Mistral AI | AssemblyAI |
| Model id | mistralai/Voxtral-Small-24B-2507 | universal |
| Price | $0.004 / audio minute | $0.0025 / audio minute |
| Capabilities | Speech to text, 8 languages, Auto language, Text only | Speech to text, Diarization, Speech understanding |
| Weights | Open | Closed |
| Status | Available | Available |
| Limits | No per-minute or per-day request or token caps; usage draws on a prepaid balance | Concurrency limits by plan. |
| Speed | 0.09× real time (23.1 s clip in 2.08 s) | not measured by us |
What three workloads cost
List prices applied to the same work. The cheaper side of each row is in bold.
| Workload | Voxtral Small 24B | AssemblyAI Universal | Difference |
|---|---|---|---|
| 1 hour of audio | $0.24 | $0.15 | 38% |
| 100 hours of audio | $24.00 | $15.00 | 38% |
| 1,000 hours of audio | $240 | $150 | 38% |
AssemblyAI Universal is cheaper in all three workloads, by 38%.
When to choose which
- You want open weights: the same model can be moved to another host or self-hosted later
- Accuracy on messy real-world audio in English, Spanish, French, Portuguese, Hindi, German, Dutch or Italian
- You want Mistral's Voxtral without running a 24B model yourself
- You need an OpenAI-compatible endpoint; its API has its own request format
- Cost is the deciding factor; it is cheaper in every workload above
- You need speaker diarization
- You want built-in summaries, sentiment or PII redaction on top of the transcript
Measured speed
Median of three runs on 2026-09-16: a 23.1-second mp3 clip, timed from upload to response. Expect ±30% with time of day. We did not measure AssemblyAI and do not quote other people's numbers; read "not measured" as unknown, not slow. Method and raw numbers: speed measurements.
FAQ
Which is cheaper, Voxtral Small 24B or AssemblyAI Universal?
AssemblyAI Universal is cheaper in all three workloads, by 38%. The table above uses list prices: Voxtral Small 24B $0.004 / audio minute; AssemblyAI Universal $0.0025 / audio minute.
What limits apply to AssemblyAI Universal?
Concurrency limits by plan. On Inference APIs there are no per-minute or per-day request or token caps; usage draws on a prepaid balance.
Spotted a price or limit that has changed? Tell us and we will re-check the provider page.
