Blog
Benchmarks, migration notes and how the service is run. RSS
We timed every chat model on the public endpoint with a fixed prompt: median time to first token, output tokens per second and total time, plus speech and transcription speed. Numbers, method and caveats.
On 16 August 2026 Groq shut down llama-3.3-70b-versatile and llama-3.1-8b-instant for free and developer tiers. A month later, hundreds of thousands of code files still pin the id. What the error looks like, who it hit, and the fix ladder.
Every model is priced at about 30% over what it costs us to serve, and the rate is on the model page. What that means in practice, and why we do not have tiers.
