Inference APIs
Compare/Speech-to-text

Voxtral Mini 3B vs AssemblyAI Universal

Voxtral Mini 3B is an open-weight speech-to-text model you can call on this endpoint; AssemblyAI Universal is a closed model from AssemblyAI. Below: what each costs for three realistic workloads, the limits that apply, and when to choose which.

AssemblyAI figures are public list prices and documented limits, checked 2026-09-16 against the provider page. Inference APIs prices are read live from the price list.

Short answer
  • Cost: Voxtral Mini 3B is cheaper in all three workloads, by 48%.
  • Limits: AssemblyAI: Concurrency limits by plan. Here: no per-minute or per-day caps.

Side by side

Voxtral Mini 3BAssemblyAI Universal
Served byInference APIsAssemblyAI
Model authorMistral AIAssemblyAI
Model idmistralai/Voxtral-Mini-3B-2507universal
Price$0.0013 / audio minute$0.0025 / audio minute
CapabilitiesSpeech to text, 8 languages, Auto language, Lowest priceSpeech to text, Diarization, Speech understanding
WeightsOpenClosed
StatusAvailableAvailable
LimitsNo per-minute or per-day request or token caps; usage draws on a prepaid balanceConcurrency limits by plan.
Speed0.058× real time (23.1 s clip in 1.35 s)not measured by us

What three workloads cost

List prices applied to the same work. The cheaper side of each row is in bold.

WorkloadVoxtral Mini 3BAssemblyAI UniversalDifference
1 hour of audio$0.08$0.1548%
100 hours of audio$7.80$15.0048%
1,000 hours of audio$78.00$15048%

Voxtral Mini 3B is cheaper in all three workloads, by 48%.

When to choose which

Choose Voxtral Mini 3B if
  • Cost is the deciding factor; it is cheaper in every workload above
  • You want open weights: the same model can be moved to another host or self-hosted later
  • Cheap, fast multilingual transcription in Voxtral's eight languages
  • Batch jobs where cost per minute matters most
  • You need an OpenAI-compatible endpoint; its API has its own request format
Choose AssemblyAI Universal if
  • You need speaker diarization
  • You want built-in summaries, sentiment or PII redaction on top of the transcript

Measured speed

Median of three runs on 2026-09-16: a 23.1-second mp3 clip, timed from upload to response. Expect ±30% with time of day. We did not measure AssemblyAI and do not quote other people's numbers; read "not measured" as unknown, not slow. Method and raw numbers: speed measurements.

FAQ

Which is cheaper, Voxtral Mini 3B or AssemblyAI Universal?

Voxtral Mini 3B is cheaper in all three workloads, by 48%. The table above uses list prices: Voxtral Mini 3B $0.0013 / audio minute; AssemblyAI Universal $0.0025 / audio minute.

What limits apply to AssemblyAI Universal?

Concurrency limits by plan. On Inference APIs there are no per-minute or per-day request or token caps; usage draws on a prepaid balance.

Spotted a price or limit that has changed? Tell us and we will re-check the provider page.