Drawpie Explainers & tools

DeepSeek V4 Flash vs Gemini, GPT and Claude: Similar AI Scores, Vastly Different Token Prices

DeepSeek V4 Flash vs Gemini, GPT and Claude: Similar AI Scores, Vastly Different Token Prices
Photo by Conny Schneider on Unsplash
Key takeaways
  • Matched on measured capability rather than tier name, DeepSeek V4 Flash’s closest peer is Gemini 3.6 Flash: 49.93 against 50.07 on the Artificial Analysis Intelligence Index, a gap of 0.14 points. Google charges 10.7x more for input and 26.8x more for output.
  • Two models commonly put in this comparison do not belong. Gemini 3.5 Flash-Lite scores 13.45 points lower, and Claude Haiku 4.5 scores 29.58 with reasoning — 20.35 below. Anthropic’s nearest model by score is Sonnet 4.6 at 47.21, but it costs more than Sonnet 5 while scoring six points less, so the tables use Sonnet 5.
  • GPT-5.6 Luna is the real competitor. It scores 1.31 points higher for 1.4x the input price and 4.3x the output price — a far narrower gap than either Google or Anthropic offers.
  • DeepSeek V4 Pro costs about three times V4 Flash and scores 5.66 points lower on the same index. Within DeepSeek’s own range, the cheaper model is the stronger one on this measure.
  • Two dated changes sit in the vendors’ own pages: DeepSeek says peak pricing at 2x is coming for seven hours a day, and Claude Sonnet 5 rises from $2/$10 to $3/$15 on 1 September 2026, taking its output multiple to 53.6x.

On Artificial Analysis’s cross-vendor index, DeepSeek V4 Flash scores 49.93 and Gemini 3.6 Flash scores 50.07. Google charges 10.7 times more for input and 26.8 times more for output.

Those are list prices per token. What a given task costs depends on how many tokens each model spends reaching an answer, which no price table shows — so read 27x as the gap on the output rate, not a promise about your bill.

That is the comparison worth making, and it is not the one you get by matching tier names. Matched on measured capability instead, DeepSeek’s peers are Gemini 3.6 Flash, GPT-5.6 Luna and Claude Sonnet 5 — not the models with “Flash” or “Haiku” in the name.

What does DeepSeek V4 Flash actually cost?

Three numbers, from DeepSeek’s own pricing page.

DeepSeek V4 FlashPer 1M tokens
Input, cache miss$0.14
Input, cache hit$0.0028
Output$0.28

The model version is DeepSeek-V4-Flash-0731. It carries a 1M-token context window and a maximum output of 384K tokens, supports both thinking and non-thinking modes with thinking as the default, and has a concurrency limit of 2,500.

For reference, the tier above it — DeepSeek V4 Pro — costs $0.435 input, $0.87 output and $0.003625 on a cache hit, with a concurrency limit of 500.

Those Pro figures have a history worth knowing, because it shows how movable these numbers are. On 1 May 2026 the same page listed them as a 75% discount: $0.435 against a standard $1.74, and $0.87 against $3.48. Today the discount markers and the higher figures are gone and the reduced rates are simply the price. V4 Flash carried no such marker — $0.14 and $0.28 have been the plain price throughout.

Which models is it actually competing with?

Not the ones with matching names. That is the easy comparison and it is the wrong one.

Matching by tier label — DeepSeek’s “Flash” against everyone else’s “Flash”, “Luna” or “Haiku” — assumes vendors use those words to mean the same thing. They do not. To place these models on one scale we used the Artificial Analysis Intelligence Index , which runs its own evaluations across vendors on a single methodology, rather than each vendor’s self-reported benchmarks, which are not comparable with each other.

On that index, DeepSeek V4 Flash scores 49.93. Here is where it sits, with Gemini 3.5 Flash-Lite kept in the bottom row for contrast rather than as a peer:

ModelIntelligence Indexvs DeepSeek V4 Flash
Claude Sonnet 5 (max)53.35+3.42
GPT-5.6 Luna (max)51.24+1.31
Gemini 3.6 Flash50.07+0.14
DeepSeek V4 Flash 0731 (max)49.93
Gemini 3.5 Flash-Lite36.48−13.45

Two corrections fall out of that, and both were in an earlier version of this article.

Gemini 3.5 Flash-Lite is not a peer. It sits 13.45 points below DeepSeek V4 Flash — a larger gap than separates V4 Flash from Claude Opus 5’s tier. Including it because its price was close was exactly the error of matching on the wrong axis.

Claude Haiku 4.5 is not a peer either. Artificial Analysis scores it at 29.58 with reasoning — 20.35 points below DeepSeek V4 Flash. Its non-reasoning figure of 23.71 is labelled on the same site as an estimate pending independent evaluation, so only the 29.58 is a measured score. It is a real model with a real score, and that score is nowhere near this comparison.

Anthropic is the awkward case, and worth setting out honestly. Two of its models bracket DeepSeek V4 Flash:

Anthropic modelIndexGapInputOutput
Claude Sonnet 5 (max)53.35+3.42$2.00$10.00
Claude Sonnet 4.6 (max)47.21−2.72$3.00$15.00

Sonnet 4.6 is closer on score. It is also more expensive than Sonnet 5 — $3 and $15 against $2 and $10 — while scoring six points lower, because Sonnet 5’s current rates are introductory and expire on 31 August. So the nearest Anthropic model by score is the worse buy on both axes, and the tables below use Sonnet 5, which is what anyone choosing today would actually pick.

At matched capability, what does each one cost?

ModelIndexInputOutputCheapest cached input
DeepSeek V4 Flash49.93$0.14$0.28$0.0028
Gemini 3.6 Flash50.07$1.50$7.50$0.15
GPT-5.6 Luna51.24$0.20$1.20$0.02
Claude Sonnet 5, to 31 Aug53.35$2.00$10.00$0.20
Claude Sonnet 5, from 1 Sep53.35$3.00$15.00$0.30

As multiples of DeepSeek V4 Flash:

ModelIndex gapInputOutputCached input
Gemini 3.6 Flash+0.1410.7x26.8x53.6x
GPT-5.6 Luna+1.311.4x4.3x7.1x
Claude Sonnet 5, to 31 Aug+3.4214.3x35.7x71.4x
Claude Sonnet 5, from 1 Sep+3.4221.4x53.6x107.1x

The Gemini row is the one to sit with. Gemini 3.6 Flash scores 0.14 index points above DeepSeek V4 Flash and costs 10.7 times as much on input and 26.8 times as much on output. Whatever that 0.14 is worth, it is not obviously worth 27x.

GPT-5.6 Luna is the genuinely competitive answer: 1.31 points higher for 1.4x input and 4.3x output. Claude Sonnet 5 buys 3.42 points at 14.3x and 35.7x — and after 1 September, at 21.4x and 53.6x.

The oddity inside DeepSeek’s own range

DeepSeek V4 Pro costs more than V4 Flash and scores lower on this index: 44.27 against 49.93, at $0.435 input and $0.87 output against $0.14 and $0.28.

That is a 5.66-point deficit for roughly three times the price. On this measure there is no reading in which Pro is the better buy, which is worth knowing before the name persuades anyone otherwise. It may well win on things the index does not capture — but on the index, it does not.

What the capability match does not settle

Three things, and they matter enough to state before anyone acts on the tables.

Reasoning effort is not held constant. DeepSeek V4 Flash, GPT-5.6 Luna and Claude Sonnet 5 are all measured at max; Gemini 3.6 Flash is measured at high, as its own model page states. Different settings, not a missing one. Scores taken at different effort settings are not strictly like for like — and effort drives output tokens, which is the expensive side of every price above. A model that reaches its score by thinking longer pays for it twice.

This is one index. Artificial Analysis is a single evaluator with a single methodology. It is the right kind of source for this question — the same harness across vendors, rather than four sets of self-reported numbers — but a different evaluator would order these models differently, and none of them measures your workload.

Index points are not tokens. A model that scores a point higher but needs half as many attempts to get a task right is cheaper in practice than the per-token table suggests. Cost per finished task is the number that decides a bill. Artificial Analysis does publish a version of it — a cost per Intelligence Index task, and output tokens per task, which fold reasoning tokens and cache behaviour into the total — but the figures sit in charts we could not extract reliably, so none are quoted here. That metric, not the rate card, is where this question is actually settled.

Two dated changes that are not in most comparison tables

Both are published by the vendors themselves, and both change the answer.

DeepSeek is introducing peak pricing at 2x. The footnote on its own pricing page reads that the service “will soon adopt a peak/off-peak pricing policy. During peak hours, prices will be 2x the regular prices, applicable to all billing items”, with peak hours given as 09:00–12:00 and 14:00–18:00 Beijing time daily. That is seven hours a day. The effective date is “subject to the official announcement” and had not been announced when this was written.

The consequence is worth spelling out. At 2x, V4 Flash costs $0.28 input and $0.56 output. Its output is still less than half GPT-5.6 Luna’s $1.20 — but its input becomes more expensive than Luna’s $0.20. The headline “cheaper on both sides” survives off-peak and does not survive peak.

Nor is DeepSeek the only one moving. The $0.20 and $1.20 used here are the figures published on OpenAI’s pricing page on 31 July 2026; that page carries no change log, so a comparison against an older capture of it may not reconcile, and the date this article was read is the only anchor we can offer.

Claude Sonnet 5 gets more expensive on 1 September 2026. Anthropic’s pricing page lists two rows for the same model: $2 input and $10 output per MTok through 31 August 2026, and $3 and $15 from 1 September. That is a 50% rise on both sides, published in advance and dated. That row is already in the comparison above, and it is what takes Anthropic’s output multiple from 35.7x to 53.6x.

One other thing on Anthropic’s page that surprises people: Claude Fable 5 is among its most expensive models at $10 input and $50 output, double Opus 5’s $5 and $25. It is not alone at that price — Claude Mythos 5, listed as limited availability, matches it — and the deprecated Opus 4.1 and retired Opus 4 are still listed higher at $15 and $75.

What these numbers do not tell you

Four things, in rough order of how much they move a real bill.

Cache mechanics are not comparable. DeepSeek prices a cache hit at $0.0028 and applies it automatically. Anthropic charges separately for writing to cache — on Sonnet 5, $2.50 per MTok for a 5-minute cache and $4 for an hour — before you get the $0.20 hit rate, all three rising by half on 1 September. Google charges $0.15 for cached context plus a storage fee of $1.00 per million tokens per hour. Three different billing models wearing the same word. The single “cheapest cached input” column above is the only piece of them that can be lined up.

Context tiers. OpenAI charges a different rate once a request crosses into its long-context band — GPT-5.6 Luna goes from $0.20/$1.20 to $0.40/$1.80. DeepSeek publishes one price for its full 1M-token window. For long-document work that difference compounds quietly.

The free tier. Google publishes a free tier for Gemini alongside the paid one, which no amount of per-token comparison captures.

List price is not the bill. Reasoning tokens you never see, cache-hit rates in the real world, retries, and per-request overheads all sit between the rate card and the invoice. We took that apart separately in what Opus 5, Fable 5 and GPT-5.6 Sol actually cost , and the reasoning there applies to every model here.

And one thing this article deliberately does not do: rank these models on quality. Price per token says nothing about how many tokens a model needs to get an answer right, which is the number that actually decides cost per finished task.

So is DeepSeek V4 Flash the cheapest?

On list prices, off-peak, yes — by a wide margin on output and a narrow one on input.

Off-peak and on output, nothing here is close to it. On input, GPT-5.6 Luna is within 43%, and once DeepSeek’s peak pricing takes effect, Luna is cheaper on input for seven hours a day. On cached input its peers charge 7.1x to 71.4x more today, and 7.1x to 107.1x once Sonnet 5’s introductory pricing ends on 1 September. That matters most for agent-style workloads that re-send the same context repeatedly.

The honest summary is narrower than the headline. DeepSeek V4 Flash is priced like a tier below the models it actually scores alongside, and its advantage is widest on output tokens and cache hits — which is where reasoning-heavy and context-resending workloads tend to spend, though not where every workload spends.

Sources

Each vendor’s own pricing page, all read on 31 July 2026.

SourceWhat it supports here
Artificial Analysis model indexThe Intelligence Index scores used to match models on measured capability rather than tier name. The listing defaults to 25 of 589 models, so per-model pages carry scores it omits — including Claude Haiku 4.5 and Claude Sonnet 4.6
DeepSeek API pricing, archived 1 May 2026The earlier V4 Pro rates and their “75% off” markers — $0.435 against $1.74 and $0.87 against $3.48 — which the current page no longer shows
DeepSeek API pricingV4 Flash at $0.14 / $0.28 / $0.0028, V4 Pro’s rates, the DeepSeek-V4-Flash-0731 version, the 1M context and 384K max output, concurrency limits, and the peak-pricing footnote with its Beijing-time hours
OpenAI API pricingGPT-5.6 Luna at $0.20 / $1.20 short-context and $0.40 / $1.80 long-context on the Standard tier, and the Standard / Batch / Flex / Fast mode split
Anthropic API pricingClaude Haiku 4.5 at $1 / $5 with $0.10 cache hits, the Sonnet 5 change dated 1 September 2026, and Fable 5 at $10 / $50
Google Gemini API pricingGemini 3.6 Flash at $1.50 / $7.50, Gemini 3.5 Flash-Lite at $0.30 / $2.50, the caching and storage rates, the free tier, and the note that output prices include thinking tokens

Anthropic’s pricing documentation is reached through redirects from docs.anthropic.com and docs.claude.com; the final URL is the one cited. Vendor pricing pages change without notice — every figure here is as published on 31 July 2026.

How we verified this

Every price in this article comes from the vendor’s own published pricing page, read on 31 July 2026, not from an aggregator or a price-tracking site. Several sites that rank well for these queries are not the vendors — deepseek.ai and chat-deep.ai are not DeepSeek’s domain, which is deepseek.com — and nothing from them is used here.

Prices are list prices per million tokens in US dollars, for the standard service tier. That qualifier matters more than it looks. OpenAI’s page carries four tiers behind a toggle — Standard, Batch, Flex and Fast mode — and we read the Standard table, confirmed by checking which radio was selected rather than assuming the first table shown was the default. OpenAI also splits every model into short-context and long-context columns at different prices; the short-context figures are used and labelled. Google lists a free tier alongside the paid tier, and its output price is stated to include thinking tokens. Anthropic quotes per MTok and lists cache writes separately from cache hits.

Because of those differences, the only figures compared directly here are base input, output, and the cheapest cached-input rate each vendor publishes. Caching is not otherwise comparable across the four: DeepSeek prices a cache hit automatically, Anthropic charges separately for 5-minute and 1-hour cache writes, and Google charges a storage fee of $1.00 per million tokens per hour on top of its cache rate. Anywhere those mechanisms would need to be added together to compare, we have not compared them.

Model selection is the part of this article most likely to be wrong if done carelessly, so it is not done by tier name. Candidates are chosen by score on the Artificial Analysis Intelligence Index, an independent evaluator running one methodology across vendors, and one judgement is layered on top of that: where two models from the same vendor bracket DeepSeek V4 Flash, we use the one a buyer would actually choose today. That applies once, to Anthropic — Sonnet 4.6 is nearer on score but costs more than Sonnet 5 while scoring lower, so the tables use Sonnet 5 and the article shows both. Vendor-published benchmark numbers are not used, because they are produced on different harnesses at different reasoning settings and cannot be compared with each other. An earlier version of this article matched by name, which put Gemini 3.5 Flash-Lite in the peer set at 13.45 points below DeepSeek V4 Flash, and Claude Haiku 4.5 in it at 20.35 points below. A later version claimed the index carried no score for Haiku 4.5 at all; it does, on the model’s own page. That claim came from scraping the index’s default view, which the page itself labels “25 of 589 models”, and mistaking it for the whole catalogue. Neither Haiku 4.5 nor Sonnet 4.6 appears in that default view; both have scores on their own pages.

The index has a limit we state rather than smooth over: reasoning effort is not constant across its entries. DeepSeek V4 Flash, GPT-5.6 Luna and Claude Sonnet 5 are labelled “(max)” while Gemini 3.6 Flash’s own page labels it “(high)”, so those scores are not strictly like for like — and because effort drives output tokens, that caveat cuts directly into the price comparison rather than sitting beside it.

These are list prices, not a bill. Reasoning tokens, cache-hit rates, context growth and per-request overheads all move the real number, and our separate piece on what LLMs actually cost covers that. This article does use a capability measure — the Artificial Analysis index — but only to decide which models belong in the same comparison. It does not run or reproduce any benchmark of its own, does not rank these models on quality beyond that one index, and is not investment or purchasing advice.