Inference APIs
Compare/Chat

DeepSeek V4 Flash vs Qwen3.8 Flash

Two open-weight chat models on the same endpoint and API key, so switching between them is a one-word change. Below: what each costs for three realistic workloads, how they differ, and measured speed.

Inference APIs prices are read live from the price list.

Short answer
  • Cost: It depends on the shape of the workload: DeepSeek V4 Flash is cheaper for chat assistant and generation-heavy, Qwen3.8 Flash for long prompts or rag.

Side by side

DeepSeek V4 FlashQwen3.8 Flash
Served byInference APIsInference APIs
Model authorDeepSeekAlibaba Qwen
Model iddeepseek-ai/DeepSeek-V4-FlashQwen/Qwen3.8-Flash
Input price$0.19 / 1M tokens$0.15 / 1M tokens
Output price$0.38 / 1M tokens$0.50 / 1M tokens
Context window1M tokens1M tokens
CapabilitiesChat, Reasoning, Coding, Tool callingChat, Reasoning, Tool calling, JSON mode, Cached input
WeightsOpenOpen
StatusAvailableAvailable
LimitsNo per-minute or per-day request or token caps; usage draws on a prepaid balanceNo per-minute or per-day request or token caps; usage draws on a prepaid balance
Time to first token0.85 s
Output speed54 tokens / s

What three workloads cost

List prices applied to the same work. The cheaper side of each row is in bold.

WorkloadDeepSeek V4 FlashQwen3.8 FlashDifference
Chat assistant — 10,000 turns of 800 tokens in, 300 out$2.66$2.701%
Long prompts or RAG — 10,000 requests of 8,000 in, 500 out$17.10$14.5015%
Generation-heavy — 10,000 requests of 500 in, 2,000 out$8.55$10.7520%

It depends on the shape of the workload: DeepSeek V4 Flash is cheaper for chat assistant and generation-heavy, Qwen3.8 Flash for long prompts or rag. Reasoning models bill their thinking as output tokens, so real output counts run higher than the visible answer; treat the generation-heavy row as a floor for them.

When to choose which

Choose DeepSeek V4 Flash if
  • Cost per token matters most; it is the cheapest chat model here
  • You process long documents and want a 1M-token context
  • You need vision; use Qwen3-VL 235B
  • You need the strongest coding model; GLM 5.3 and DeepSeek V4 Pro rank higher
Choose Qwen3.8 Flash if
  • You want a fast, inexpensive model with a 1M-token window and prompt caching
  • Long-document summarisation and extraction where the whole file goes in one request
  • A person is waiting for the answer and you leave reasoning on; at default settings it thought for about twenty seconds in our measurement

Price, context and speed are measurable; answer quality on your task is not something a table can settle. Both take the same request, so the honest test is to run your own prompts through each in the playground.

Measured speed

Medians of three streamed runs on 2026-09-17, public endpoint, a prompt that asks for about 120 words. Time to first token is the first token of any kind, reasoning included; output speed counts every generated token from that point. Reasoning models then think before the visible answer starts: DeepSeek V4 Flash began answering after 20.1 s at default settings, which you can shorten with the reasoning controls on each model page. Expect ±30% with time of day. Method and raw numbers: speed measurements.

Trying both

Same endpoint, same key. Change one string:

cURL
for MODEL in "deepseek-ai/DeepSeek-V4-Flash" "Qwen/Qwen3.8-Flash"; do
  curl -s https://api.inferenceapis.com/v1/chat/completions \
    -H "Authorization: Bearer $INFERENCE_API_KEY" -H "Content-Type: application/json" \
    -d "{\"model\": \"$MODEL\", \"max_tokens\": 500, \"messages\": [{\"role\": \"user\", \"content\": \"Summarise the plot of Hamlet in two sentences.\"}]}" \
    | jq -r '.model, .choices[0].message.content, .usage'
done

FAQ

Which is cheaper, DeepSeek V4 Flash or Qwen3.8 Flash?

It depends on the shape of the workload: DeepSeek V4 Flash is cheaper for chat assistant and generation-heavy, Qwen3.8 Flash for long prompts or rag. The table above uses list prices: DeepSeek V4 Flash $0.19 / 1M tokens input, $0.38 / 1M tokens output; Qwen3.8 Flash $0.15 / 1M tokens input, $0.50 / 1M tokens output.

How do I switch between them?

They are on the same endpoint and key. Change the model field from deepseek-ai/DeepSeek-V4-Flash to Qwen/Qwen3.8-Flash and nothing else.

Why does DeepSeek V4 Flash use more output tokens than the visible answer?

It is a reasoning model: it thinks before it answers and the thinking is billed as output tokens. Budget for that in output-heavy workloads, and set max_tokens to a few hundred or more so the answer is not cut off.

Spotted a price or limit that has changed? Tell us and we will re-check the provider page.