Skip to content

DeepSeek V4-Flash

A 284B-parameter MoE with only 13B active, which is why it runs fast and prices low. Notable for a 384K maximum output — far beyond the 128K most frontier models cap at — and for supporting both thinking and non-thinking modes, so you can switch reasoning off on easy requests rather than paying for it.

Pricing checked against provider documentation on . How we verify

stable DeepSeek DeepSeek V4 deepseek-v4-flash reasoning open weights

Input

$0.22

/ 1M tokens

Output

$0.66

/ 1M tokens

Context

1M

1,000,000 tokens

Max output

384K

tokens per request

Pricing caveat: Off-peak rate shown. Peak rates double ($0.44 / $1.32) during 01:00–04:00 and 06:00–10:00 UTC on weekdays. Superseded the earlier flat-rate pricing of $0.14 / $0.28 that many comparison sites still quote.

What it costs in practice

A workload of 10M input and 2M output tokens per month — roughly a moderately busy production assistant — runs $3.52/month on DeepSeek V4-Flash.

Run your own numbers in the calculator →

Where DeepSeek V4-Flash fits

  • High-volume classification and extraction
  • Very long generated outputs
  • Budget coding agents
  • Local deployment on modest hardware

Specifications

API model ID
deepseek-v4-flash
Provider
DeepSeek
Model family
DeepSeek V4
Released
June 2026
Knowledge cutoff
March 2026
Context window
1,000,000 tokens
Max output
384,000 tokens
Native reasoning
Yes
Relative latency
low
Input modalities
text
Output modalities
text
Open weights
MIT
Cached input
$0.007 / 1M tokens
Batch discount

Frequently asked

How much does DeepSeek V4-Flash cost?

DeepSeek V4-Flash costs $0.22 per million input tokens and $0.66 per million output tokens, with cached input at $0.01.

What is DeepSeek V4-Flash's context window?

DeepSeek V4-Flash accepts up to 1,000,000 tokens of context and can generate up to 384,000 output tokens per request.

Is DeepSeek V4-Flash a reasoning model?

Yes. DeepSeek V4-Flash performs native chain-of-thought before answering. Those thinking tokens are billed at the output rate, so budget above the sticker price.

What is DeepSeek V4-Flash best for?

The cheapest capable open-weight option. A 284B-parameter MoE with only 13B active, which is why it runs fast and prices low. Notable for a 384K maximum output — far beyond the 128K most frontier models cap at — and for supporting both thinking and non-thinking modes, so you can switch reasoning off on easy requests rather than paying for it.