Inference APIs
Compare/Speech-to-text

Voxtral Small 24B vs AssemblyAI Universal

Voxtral Small 24B is an open-weight speech-to-text model you can call on this endpoint; AssemblyAI Universal is a closed model from AssemblyAI. Below: what each costs for three realistic workloads, the limits that apply, and when to choose which.

AssemblyAI figures are public list prices and documented limits, checked 2026-09-16 against the provider page. Inference APIs prices are read live from the price list.

Short answer
  • Cost: AssemblyAI Universal is cheaper in all three workloads, by 38%.
  • Limits: AssemblyAI: Concurrency limits by plan. Here: no per-minute or per-day caps.

Side by side

Voxtral Small 24BAssemblyAI Universal
Served byInference APIsAssemblyAI
Model authorMistral AIAssemblyAI
Model idmistralai/Voxtral-Small-24B-2507universal
Price$0.004 / audio minute$0.0025 / audio minute
CapabilitiesSpeech to text, 8 languages, Auto language, Text onlySpeech to text, Diarization, Speech understanding
WeightsOpenClosed
StatusAvailableAvailable
LimitsNo per-minute or per-day request or token caps; usage draws on a prepaid balanceConcurrency limits by plan.
Speed0.09× real time (23.1 s clip in 2.08 s)not measured by us

What three workloads cost

List prices applied to the same work. The cheaper side of each row is in bold.

WorkloadVoxtral Small 24BAssemblyAI UniversalDifference
1 hour of audio$0.24$0.1538%
100 hours of audio$24.00$15.0038%
1,000 hours of audio$240$15038%

AssemblyAI Universal is cheaper in all three workloads, by 38%.

When to choose which

Choose Voxtral Small 24B if
  • You want open weights: the same model can be moved to another host or self-hosted later
  • Accuracy on messy real-world audio in English, Spanish, French, Portuguese, Hindi, German, Dutch or Italian
  • You want Mistral's Voxtral without running a 24B model yourself
  • You need an OpenAI-compatible endpoint; its API has its own request format
Choose AssemblyAI Universal if
  • Cost is the deciding factor; it is cheaper in every workload above
  • You need speaker diarization
  • You want built-in summaries, sentiment or PII redaction on top of the transcript

Measured speed

Median of three runs on 2026-09-16: a 23.1-second mp3 clip, timed from upload to response. Expect ±30% with time of day. We did not measure AssemblyAI and do not quote other people's numbers; read "not measured" as unknown, not slow. Method and raw numbers: speed measurements.

FAQ

Which is cheaper, Voxtral Small 24B or AssemblyAI Universal?

AssemblyAI Universal is cheaper in all three workloads, by 38%. The table above uses list prices: Voxtral Small 24B $0.004 / audio minute; AssemblyAI Universal $0.0025 / audio minute.

What limits apply to AssemblyAI Universal?

Concurrency limits by plan. On Inference APIs there are no per-minute or per-day request or token caps; usage draws on a prepaid balance.

Spotted a price or limit that has changed? Tell us and we will re-check the provider page.