DeepSeek V4-Flash

DeepSeek has released V4-Flash, a faster and cheaper version of its flagship artificial intelligence model.

The headline number is hard to ignore. Independent testing by Artificial Analysis found that DeepSeek V4-Flash costs an average of just $0.03 per benchmark test, making it the cheapest widely known AI model included in the research firm’s comparison.

That does not mean it is the smartest model available. It does mean that developers can run a large number of tasks without watching API bills climb quite so quickly.

DeepSeek V4-Flash Pricing Undercuts Major Rivals

DeepSeek charges $0.14 per million input tokens and $0.28 per million output tokens through its API.

Artificial Analysis calculated a blended price of roughly $0.06 per million tokens based on a mix of cached inputs, new inputs and generated outputs. Prices may vary between hosting providers, but the model remains unusually cheap compared with other major reasoning systems.

Benchmark testing produced an even wider gap.

DeepSeek V4-Flash averaged three cents per test. Moonshot AI’s Kimi K3 reportedly cost $0.86 for the same type of evaluation, while OpenAI’s GPT-5.6 Sol reached $1.86. Anthropic’s Claude Fable 5 came in at $3.15.

That makes DeepSeek’s model more than 100 times cheaper to run than Claude Fable 5 under the test conditions used by Artificial Analysis.

Token prices alone rarely tell the whole story. Some models generate longer internal reasoning chains, call more tools or produce far more output to reach an answer. Measuring the cost of completing the entire benchmark gives developers a clearer picture of what a model may cost in practice.

The Model Trades Some Performance for Lower Costs

DeepSeek V4-Flash scored 50 out of 100 on the Artificial Analysis Intelligence Index.

The index combines nine evaluations covering reasoning, coding and workplace tasks. Its score placed the model alongside Google’s Gemini 3.6 Flash and slightly behind several competing systems from Meta and Z.AI.

More expensive frontier models still finished higher.

Moonshot AI’s Kimi K3 scored 57, while top models from Anthropic and OpenAI held a lead of at least nine points, according to the comparison.

That difference matters for difficult coding problems, long agent workflows and tasks where accuracy is more important than cost.

For customer support, document processing, content classification, data extraction or routine coding work, the calculation looks different. A model does not always need to lead every benchmark. It needs to perform the job reliably enough at a price that works at scale.

DeepSeek appears to be targeting exactly that market.

Fast Output and a One-Million-Token Context Window

Price is only part of the V4-Flash pitch.

Artificial Analysis measured output speeds of more than 100 tokens per second through DeepSeek’s own API. The model also supports a context window of up to one million tokens, giving it room to process large documents, long conversations and substantial codebases in a single request.

V4-Flash is a text-based reasoning model and its weights are available for self-hosting. That gives businesses another option beyond paying for access through a closed API.

Running a large model internally still brings infrastructure costs, engineering work and security responsibilities. Open weights do not make deployment effortless. They do give companies more control over where their data goes and how the model is adjusted.

DeepSeek Is Competing on Efficiency Again

DeepSeek became one of the most closely watched Chinese AI companies after releasing its R1 reasoning model. The company challenged the idea that competitive AI systems must always require enormous development and operating budgets.

V4-Flash continues that strategy, but the market is much more crowded now.

Moonshot AI, Alibaba, MiniMax, Z.AI and ByteDance are all releasing models aimed at developers and enterprise users. Many are competing on open weights, lower API fees and performance that sits just below the most expensive Western systems.

The result is a price war that reaches well beyond China.

Software companies can route simple requests to cheaper models and reserve premium systems for the small number of tasks that genuinely require them. A support question might go to DeepSeek. A complicated autonomous coding job could be sent somewhere else.

That model-routing approach could become more common as the performance gap continues to narrow.

Cheap AI Changes How Companies Build Products

A few cents saved on one request means very little.

Multiply that saving across millions of customer conversations, searches, document reviews or automated workflows and the economics change quickly.

Lower inference costs allow developers to experiment more freely. Features that once looked too expensive to run inside a consumer app may suddenly become practical. Startups also get more room to compete without raising huge amounts of money simply to pay model providers.

There is still a catch.

The cheapest model is not automatically the best choice. Businesses need to examine reliability, privacy, output quality, latency, content restrictions and hosting options before moving production workloads.

DeepSeek V4-Flash will not replace every premium AI system. It does not need to.

At three cents per benchmark run, it puts pressure on every major model provider to explain why its intelligence should cost more.

Sources