Inference APIs
Compare/Chat

DeepSeek V4 Pro vs Llama 3.3 70B on Groq

Llama 3.3 70B on Groq was retired for self-serve accounts, so the practical question is what replaces it. This page puts DeepSeek V4 Pro on this endpoint next to the retired offer: price, limits and what changes in your code.

Groq figures are public list prices and documented limits, checked 2026-09-16 against the provider page. Inference APIs prices are read live from the price list.

Short answer
  • Cost: Groq listed it at a lower price, but it can no longer be called there, so the comparison that matters is between the options that still exist.
  • Availability: Retired from Groq free and developer tiers on 2026-08-16; requests return 404 model_not_found. Enterprise contracts only.
  • Not tied to this model? DeepSeek V4 Flash costs less here for the same workloads; keep DeepSeek V4 Pro only if your prompts and tests depend on it.

Side by side

DeepSeek V4 ProLlama 3.3 70B on Groq
Served byInference APIsGroq
Model authorDeepSeekDeepSeek
Model iddeepseek-ai/DeepSeek-V4-Prollama-3.3-70b-versatile
Input price$1.72 / 1M tokens$0.59 / 1M tokens
Output price$5.15 / 1M tokens$0.79 / 1M tokens
Context window1M tokens131K tokens
CapabilitiesChat, Reasoning, Coding, Tool calling, JSON modeChat, Tool calling, JSON mode
WeightsOpenOpen
StatusAvailableRetired
LimitsNo per-minute or per-day request or token caps; usage draws on a prepaid balanceRetired from Groq free and developer tiers on 2026-08-16; requests return 404 model_not_found. Enterprise contracts only.
Time to first token0.64 snot measured by us
Output speed121.4 tokens / snot measured by us

What three workloads cost

List prices applied to the same work. The cheaper side of each row is in bold.

WorkloadDeepSeek V4 ProLlama 3.3 70B on GroqDifference
Chat assistant — 10,000 turns of 800 tokens in, 300 out$29.21$7.0976%
Long prompts or RAG — 10,000 requests of 8,000 in, 500 out$163$51.1569%
Generation-heavy — 10,000 requests of 500 in, 2,000 out$112$18.7583%

Groq listed it at a lower price, but it can no longer be called there, so the comparison that matters is between the options that still exist. Reasoning models bill their thinking as output tokens, so real output counts run higher than the visible answer; treat the generation-heavy row as a floor for them. The Llama 3.3 70B on Groq column is its last list price, shown for reference only.

When to choose which

Choose DeepSeek V4 Pro if
  • Llama 3.3 70B on Groq is no longer offered to self-serve accounts, so this is the one you can actually call
  • You need more than 131K tokens of context (it takes 1M)
  • Hard reasoning and coding tasks where the Flash models fall short
  • You want DeepSeek's largest V4 model without sending data to DeepSeek's own API
Llama 3.3 70B on Groq

Retired for self-serve accounts. There is no case for choosing it for new work.

Price, context and speed are measurable; answer quality on your task is not something a table can settle. Both take the same request, so the honest test is to run your own prompts through each.

Measured speed

Medians of three streamed runs on 2026-09-17, public endpoint, a prompt that asks for about 120 words. Time to first token is the first token of any kind, reasoning included; output speed counts every generated token from that point. Reasoning models then think before the visible answer starts: DeepSeek V4 Pro began answering after 1.25 s at default settings, which you can shorten with the reasoning controls on each model page. Expect ±30% with time of day. We did not measure Groq and do not quote other people's numbers; read "not measured" as unknown, not slow. Method and raw numbers: speed measurements.

Switching from Groq

Both speak the OpenAI wire format. The edit is the base URL, the key and the model id; then re-run your own prompts, because it is a different model.

Python · openai SDK
import os
from openai import OpenAI

# before: Groq
# client = OpenAI(base_url="https://api.groq.com/openai/v1", api_key=os.environ["GROQ_API_KEY"])
# MODEL = "llama-3.3-70b-versatile"

# after: Inference APIs
client = OpenAI(base_url="https://api.inferenceapis.com/v1", api_key=os.environ["INFERENCE_API_KEY"])
MODEL = "deepseek-ai/DeepSeek-V4-Pro"

resp = client.chat.completions.create(model=MODEL, messages=[{"role": "user", "content": "Hello"}], max_tokens=500)

More detail, including Node, LangChain and LiteLLM: switching OpenAI-compatible providers.

FAQ

Which is cheaper, DeepSeek V4 Pro or Llama 3.3 70B on Groq?

Groq listed it at a lower price, but it can no longer be called there, so the comparison that matters is between the options that still exist. The table above uses list prices: DeepSeek V4 Pro $1.72 / 1M tokens input, $5.15 / 1M tokens output; Llama 3.3 70B on Groq $0.59 / 1M tokens input, $0.79 / 1M tokens output.

Can I still use Llama 3.3 70B on Groq?

Not on a self-serve account. Retired from Groq free and developer tiers on 2026-08-16; requests return 404 model_not_found. Enterprise contracts only.

Why does DeepSeek V4 Pro use more output tokens than the visible answer?

It is a reasoning model: it thinks before it answers and the thinking is billed as output tokens. Budget for that in output-heavy workloads, and set max_tokens to a few hundred or more so the answer is not cut off.

Spotted a price or limit that has changed? Tell us and we will re-check the provider page.