Taylor Hawkes
Founder and engineer, Inference APIs.
Taylor builds and operates Inference APIs: the gateway, the model catalog, the playgrounds and the billing. Everything in the reference and documentation is written and verified against the live API and the providers' own documentation, with the verification date shown on each page.
Elsewhere: GitHub. Reach the company at hello@inferenceapis.com.
Writing
- Measured: TTFT and throughput of open chat models · 2026-09-16
We timed every chat model on the public endpoint with a fixed prompt: median time to first token, output tokens per second and total time, plus speech and transcription speed. Numbers, method and caveats. - Groq retired llama-3.3-70b-versatile: what to do · 2026-09-16
On 16 August 2026 Groq shut down llama-3.3-70b-versatile and llama-3.1-8b-instant for free and developer tiers. A month later, hundreds of thousands of code files still pin the id. What the error looks like, who it hit, and the fix ladder. - How we price inference: cost plus a fixed margin · 2026-09-16
Every model is priced at about 30% over what it costs us to serve, and the rate is on the model page. What that means in practice, and why we do not have tiers.
Also maintained: the provider error reference, the free-tier rate limits tracker and the model deprecations tracker, each re-verified against provider documentation on the date shown on the page.
