Kimi K2.7 Code vs Llama 3.3 70B Instruct
Two open-weight chat models on the same endpoint and API key, so switching between them is a one-word change. Below: what each costs for three realistic workloads, how they differ, and measured speed.
Inference APIs prices are read live from the price list.
- Cost: It depends on the shape of the workload: Kimi K2.7 Code is cheaper for long prompts or rag, Llama 3.3 70B Instruct for chat assistant and generation-heavy.
- Speed (measured): Kimi K2.7 Code 0.87 s to first token and 41 tok/s; Llama 3.3 70B Instruct 0.62 s and 98 tok/s.
Side by side
| Kimi K2.7 Code | Llama 3.3 70B Instruct | |
|---|---|---|
| Served by | Inference APIs | Inference APIs |
| Model author | Moonshot AI | Meta |
| Model id | moonshotai/Kimi-K2.7-Code | meta-llama/Llama-3.3-70B-Instruct-Turbo |
| Input price | $0.89 / 1M tokens | $1.35 / 1M tokens |
| Output price | $4.42 / 1M tokens | $1.35 / 1M tokens |
| Context window | 262K tokens | 131K tokens |
| Capabilities | Chat, Reasoning, Coding, Tool calling | Chat, Tool calling, JSON mode, Multilingual |
| Weights | Open | Open |
| Status | Available | Available |
| Limits | No per-minute or per-day request or token caps; usage draws on a prepaid balance | No per-minute or per-day request or token caps; usage draws on a prepaid balance |
| Time to first token | 0.87 s | 0.62 s |
| Output speed | 40.7 tokens / s | 98.2 tokens / s |
What three workloads cost
List prices applied to the same work. The cheaper side of each row is in bold.
| Workload | Kimi K2.7 Code | Llama 3.3 70B Instruct | Difference |
|---|---|---|---|
| Chat assistant — 10,000 turns of 800 tokens in, 300 out | $20.38 | $14.85 | 27% |
| Long prompts or RAG — 10,000 requests of 8,000 in, 500 out | $93.30 | $115 | 19% |
| Generation-heavy — 10,000 requests of 500 in, 2,000 out | $92.85 | $33.75 | 64% |
It depends on the shape of the workload: Kimi K2.7 Code is cheaper for long prompts or rag, Llama 3.3 70B Instruct for chat assistant and generation-heavy. Reasoning models bill their thinking as output tokens, so real output counts run higher than the visible answer; treat the generation-heavy row as a floor for them.
When to choose which
- You need more than 131K tokens of context (it takes 262K)
- Coding agents: it is the coding-tuned model of Moonshot's Kimi family
- You are starting fresh with no Llama-tuned prompts to preserve
- Time to first token matters: 0.62 s against 0.87 s in our measurement
- You stream long answers: 98 tokens per second against 41
- Your prompts, evals or output formats were tuned on Llama 3.3 70B and you do not want to re-tune
- You were on Groq: the llama-3.3-70b-versatile id is accepted unchanged
- You need full-precision weights; our upstream serves this model at FP4
Price, context and speed are measurable; answer quality on your task is not something a table can settle. Both take the same request, so the honest test is to run your own prompts through each in the playground.
Measured speed
Medians of three streamed runs on 2026-09-17, public endpoint, a prompt that asks for about 120 words. Time to first token is the first token of any kind, reasoning included; output speed counts every generated token from that point. Reasoning models then think before the visible answer starts: Kimi K2.7 Code began answering after 37.34 s and Llama 3.3 70B Instruct after 0.62 s at default settings, which you can shorten with the reasoning controls on each model page. Expect ±30% with time of day. Method and raw numbers: speed measurements.
Trying both
Same endpoint, same key. Change one string:
for MODEL in "moonshotai/Kimi-K2.7-Code" "meta-llama/Llama-3.3-70B-Instruct-Turbo"; do
curl -s https://api.inferenceapis.com/v1/chat/completions \
-H "Authorization: Bearer $INFERENCE_API_KEY" -H "Content-Type: application/json" \
-d "{\"model\": \"$MODEL\", \"max_tokens\": 500, \"messages\": [{\"role\": \"user\", \"content\": \"Summarise the plot of Hamlet in two sentences.\"}]}" \
| jq -r '.model, .choices[0].message.content, .usage'
doneFAQ
Which is cheaper, Kimi K2.7 Code or Llama 3.3 70B Instruct?
It depends on the shape of the workload: Kimi K2.7 Code is cheaper for long prompts or rag, Llama 3.3 70B Instruct for chat assistant and generation-heavy. The table above uses list prices: Kimi K2.7 Code $0.89 / 1M tokens input, $4.42 / 1M tokens output; Llama 3.3 70B Instruct $1.35 / 1M tokens input, $1.35 / 1M tokens output.
How do I switch between them?
They are on the same endpoint and key. Change the model field from moonshotai/Kimi-K2.7-Code to meta-llama/Llama-3.3-70B-Instruct-Turbo and nothing else.
Why does Kimi K2.7 Code use more output tokens than the visible answer?
It is a reasoning model: it thinks before it answers and the thinking is billed as output tokens. Budget for that in output-heavy workloads, and set max_tokens to a few hundred or more so the answer is not cut off.
Spotted a price or limit that has changed? Tell us and we will re-check the provider page.
