DeepSeek V4 Flash vs Gemini, GPT and Claude: Similar AI Scores, Vastly Different Token Prices

- Matched on measured capability rather than tier name, DeepSeek V4 Flash’s closest peer is Gemini 3.6 Flash: both score 52 on the Artificial Analysis Intelligence Index, exactly level. Against DeepSeek’s off-peak rates Google charges 3.4x more for input and 5.7x more for output; against its peak rates, 1.7x and 2.8x.
- Two models commonly put in this comparison do not belong. Gemini 3.5 Flash-Lite scores 37 against V4 Flash’s 52, fifteen points lower; Claude Haiku 4.5 sat far below the whole peer set when we last read its page. Anthropic’s nearest model is Sonnet 5 at 55 — Sonnet 4.6 scores 37 and costs half again as much, so the tables use Sonnet 5.
- GPT-5.6 Luna is the real competitor, and on input it is now the cheaper of the two. It scores 52, level with V4 Flash, at 0.9x DeepSeek’s off-peak input price and 1.8x its output price — a far narrower gap than either Google or Anthropic offers.
- DeepSeek V4 Pro costs about three times V4 Flash and scores one point higher on the same index, 53 against 52. Within DeepSeek’s own range, three times the money buys a single index point.
- The two dated changes this article flagged in July have both resolved, in opposite directions. DeepSeek’s 2x peak pricing is now in force, seven hours a day from Monday to Friday. Anthropic cancelled the Sonnet 5 increase and made $2/$10 the standard price. The dated change still ahead belongs to Google: Gemini 3.6 Flash goes from $0.75/$3.75 to $1.50/$7.50 on 1 January 2027.
On Artificial Analysis’s cross-vendor index, DeepSeek V4 Flash and Gemini 3.6 Flash both score 52. Against DeepSeek’s off-peak rates, Google charges 3.4 times more for input and 5.7 times more for output.
Those are list prices per token. What a given task costs depends on how many tokens each model spends reaching an answer, which no price table shows — so read 5.7x as the gap on the output rate, not a promise about your bill.
That is the comparison worth making, and it is not the one you get by matching tier names. Matched on measured capability instead, DeepSeek’s peers are Gemini 3.6 Flash, GPT-5.6 Luna and Claude Sonnet 5 — not the models with “Flash” or “Haiku” in the name.
What does DeepSeek V4 Flash actually cost?
Six numbers now, from DeepSeek’s own pricing page, because its rates are published in two bands. All are per 1M tokens.
| DeepSeek V4 Flash | Off-peak | Peak |
|---|---|---|
| Input, cache miss | $0.22 | $0.44 |
| Input, cache hit | $0.007 | $0.014 |
| Output | $0.66 | $1.32 |
The model version is DeepSeek-V4-Flash-0731. It carries a 1M-token context window and a maximum output of 384K tokens, supports both thinking and non-thinking modes with thinking as the default, and has a concurrency limit of 2,500.
For reference, the tier above it — DeepSeek V4 Pro, now at version DeepSeek-V4-Pro-0813 — costs $0.66 input, $1.98 output and $0.022 on a cache hit off-peak, with a concurrency limit of 500.
Those figures have a history worth knowing, because it shows how movable these numbers are. On 1 May 2026 the same page listed Pro at a 75% discount: $0.435 against a standard $1.74, and $0.87 against $3.48. The discount markers went, and the rates have moved again since — Pro is now $0.66 and $1.98 off-peak. V4 Flash has moved too. It was a flat $0.14 and $0.28 when this article was first written on 31 July, with no discount marker on it at all; it is $0.22 and $0.66 off-peak today.
Which models is it actually competing with?
Not the ones with matching names. That is the easy comparison and it is the wrong one.
Matching by tier label — DeepSeek’s “Flash” against everyone else’s “Flash”, “Luna” or “Haiku” — assumes vendors use those words to mean the same thing. They do not. To place these models on one scale we used the Artificial Analysis Intelligence Index , which runs its own evaluations across vendors on a single methodology, rather than each vendor’s self-reported benchmarks, which are not comparable with each other.
On that index, DeepSeek V4 Flash scores 52. Here is where it sits, with Gemini 3.5 Flash-Lite kept in the bottom row for contrast rather than as a peer:
| Option | Intelligence Index | vs DeepSeek V4 Flash | Checked |
|---|---|---|---|
| Claude Sonnet 5 (max) | 55 | +3 | 2026-08-26 |
| GPT-5.6 Luna (max) | 52 | level | 2026-08-26 |
| Gemini 3.6 Flash | 52 | level | 2026-08-26 |
| DeepSeek V4 Flash 0731 (max) | 52 | — | 2026-08-26 |
| Gemini 3.5 Flash-Lite | 37 | −15 | 2026-08-26 |
Read on each model's own page on 26 August 2026, under Intelligence Index v4.1.1.
Two corrections fall out of that, and both were in an earlier version of this article.
Gemini 3.5 Flash-Lite is not a peer. It scores 37 against V4 Flash’s 52 — fifteen points below, five times the three-point gap that separates V4 Flash from Claude Sonnet 5 at the top of these tables. Including it because its price was close was exactly the error of matching on the wrong axis.
Claude Haiku 4.5 is not a peer either. The index does carry a score for it, on the model’s own page — an earlier version of this article wrongly said it did not — and when we read that page in July, the score sat far below every model in this comparison. It has not been re-read under the new index version, so no figure for it is quoted here: it is a real model with a real score, and that score is nowhere near this comparison.
Anthropic is the awkward case, and worth setting out honestly. The two models in these tables sit eighteen points apart, and neither lands just below DeepSeek V4 Flash:
| Option | Index | Gap | Input | Output | Checked |
|---|---|---|---|---|---|
| Claude Sonnet 5 (max) | 55 | +3 | $2.00 | $10.00 | 2026-08-26 |
| Claude Sonnet 4.6 (max) | 37 | −15 | $3.00 | $15.00 | 2026-08-26 |
Gap is against DeepSeek V4 Flash at 52. Prices from Anthropic's pricing page, scores from each model's own page under Intelligence Index v4.1.1, both read 26 August 2026.
Sonnet 5 is the nearer model on both axes at once: three points above V4 Flash, and cheaper than Sonnet 4.6 at $2 and $10 against $3 and $15. Sonnet 4.6 is fifteen points below V4 Flash and half again the price of the better model above it, which leaves it out of the comparison on both counts rather than only one. Sonnet 5’s $2 and $10 were announced as introductory rates expiring on 31 August 2026; Anthropic has since made them the standard price. So the tables below use Sonnet 5, which is what anyone choosing today would actually pick — and nothing in Anthropic’s range now sits just under DeepSeek V4 Flash. On the index version this article was first written against, Sonnet 4.6 did; the re-versioned index moved it a long way down.
At matched capability, what does each one cost?
| Option | Index | Input | Output | Cheapest cached input | Checked |
|---|---|---|---|---|---|
| DeepSeek V4 Flash, off-peak | 52 | $0.22 | $0.66 | $0.007 | 2026-08-26 |
| Gemini 3.6 Flash, to 31 Dec | 52 | $0.75 | $3.75 | $0.075 | 2026-08-26 |
| Gemini 3.6 Flash, from 1 Jan 2027 | 52 | $1.50 | $7.50 | $0.15 | 2026-08-26 |
| GPT-5.6 Luna | 52 | $0.20 | $1.20 | $0.02 | 2026-08-26 |
| Claude Sonnet 5 | 55 | $2.00 | $10.00 | $0.20 | 2026-08-26 |
As multiples of DeepSeek V4 Flash’s off-peak rates:
| Option | Index gap | Input | Output | Cached input | Checked |
|---|---|---|---|---|---|
| Gemini 3.6 Flash, to 31 Dec | level | 3.4x | 5.7x | 10.7x | 2026-08-26 |
| Gemini 3.6 Flash, from 1 Jan 2027 | level | 6.8x | 11.4x | 21.4x | 2026-08-26 |
| GPT-5.6 Luna | level | 0.9x | 1.8x | 2.9x | 2026-08-26 |
| Claude Sonnet 5 | +3 | 9.1x | 15.2x | 28.6x | 2026-08-26 |
The Gemini row is the one to sit with. Gemini 3.6 Flash scores 52, exactly level with DeepSeek V4 Flash, and costs 3.4 times as much on input and 5.7 times as much on output off-peak. Level on one index is not level on everything — but there is nothing in these two scores to pay 5.7x for, and Google’s own page has that gap widening on 1 January 2027, to 6.8x and 11.4x.
GPT-5.6 Luna is the genuinely competitive answer, and it has moved past DeepSeek on one side: level on the index at 52, at 0.9x input — cheaper, in other words — and 1.8x output. Claude Sonnet 5 buys three index points at 9.1x and 15.2x, and those are now its permanent rates rather than introductory ones.
What the step up to V4 Pro buys
DeepSeek V4 Pro costs three times V4 Flash and scores one point above it on this index: 53 against 52, at $0.66 input and $1.98 output against $0.22 and $0.66, off-peak on both.
That is one index point for exactly three times the price, in both bands. And one point is as small a difference as this index can now express — v4.1.1 publishes whole numbers, so what sits behind that point may be narrower than it looks. Pro may well win on things the index does not capture — but on the index, a single point is all it wins, and it is not obviously worth 3x. Earlier versions of this article had Pro scoring below Flash on the old index and called it the worse buy outright. On the current reading that reasoning was wrong — Pro is ahead, barely — and the case for staying on Flash now rests on the price rather than on the score.
What the capability match does not settle
Three things, and they matter enough to state before anyone acts on the tables.
Reasoning effort is not held constant. DeepSeek V4 Flash, GPT-5.6 Luna and Claude Sonnet 5 are all measured at max; Gemini 3.6 Flash is measured at high, as its own model page states. Different settings, not a missing one. Scores taken at different effort settings are not strictly like for like — and effort drives output tokens, which is the expensive side of every price above. A model that reaches its score by thinking longer pays for it twice.
This is one index. Artificial Analysis is a single evaluator with a single methodology. It is the right kind of source for this question — the same harness across vendors, rather than four sets of self-reported numbers — but a different evaluator would order these models differently, and none of them measures your workload.
Index points are not tokens. A model that scores a point higher but needs half as many attempts to get a task right is cheaper in practice than the per-token table suggests. Cost per finished task is the number that decides a bill. Artificial Analysis does publish a version of it — a cost per Intelligence Index task, and output tokens per task, which fold reasoning tokens and cache behaviour into the total — but the figures sit in charts we could not extract reliably, so none are quoted here. That metric, not the rate card, is where this question is actually settled.
Three dated changes, and where they landed
All three are published by the vendors themselves. Two of them have already happened, and in opposite directions.
DeepSeek’s peak pricing is in force. In July the footnote on its pricing page said a peak/off-peak policy was coming at 2x the regular rates, with the effective date still “subject to the official announcement”. It has since arrived. The page now publishes every rate twice, and gives peak hours as 01:00–04:00 and 06:00–10:00 UTC, Monday through Friday, with all other hours off-peak — the same seven hours a day the footnote described, now on weekdays only.
The consequence is worth spelling out. At peak, V4 Flash costs $0.44 input and $1.32 output, which puts it above GPT-5.6 Luna on both sides. Off-peak it is well under Luna on output, $0.66 against $1.20, but its $0.22 input is already above Luna’s $0.20. The “cheaper on both sides” line this article ran in July does not survive either band.
Nor is DeepSeek the only one moving. The $0.20 and $1.20 used here are the figures published on OpenAI’s pricing page on 26 August 2026; that page carries no change log, so a comparison against an older capture of it may not reconcile, and the date this article was read is the only anchor we can offer.
Claude Sonnet 5’s increase was cancelled. Anthropic’s pricing page used to carry two rows for the same model: $2 input and $10 output per MTok through 31 August 2026, and $3 and $15 from 1 September. The second row has gone, replaced by a note saying the scheduled increase “will not occur” and that the $2/$10 launch pricing is now the standard price. Anthropic’s output multiple therefore stays at 15.2x rather than rising by half, and the Sonnet 5 figures in the tables above are permanent rather than introductory.
Google’s is the dated change still ahead. Gemini 3.6 Flash’s published rates run to a deadline: $0.75 input and $3.75 output through 31 December 2026, then $1.50 and $7.50 from 1 January 2027. Cached context doubles with them, $0.075 to $0.15, and so does the storage fee, $0.50 to $1.00 per million tokens per hour. It is the one dated row still inside the comparison table.
One other thing on Anthropic’s page that surprises people: Claude Fable 5 is among its most expensive models at $10 input and $50 output, double Opus 5’s $5 and $25. It is not alone at that price — Claude Mythos 5, listed as limited availability, matches it — and the retired Opus 4.1 and Opus 4 are still listed higher at $15 and $75.
What these numbers do not tell you
Four things, in rough order of how much they move a real bill.
Cache mechanics are not comparable. DeepSeek prices a cache hit at $0.007 off-peak and $0.014 at peak, and applies it automatically. Anthropic charges separately for writing to cache — on Sonnet 5, $2.50 per MTok for a 5-minute cache and $4 for an hour — before you get the $0.20 hit rate, and all three are now permanent rather than introductory. Google charges $0.075 for cached context plus a storage fee of $0.50 per million tokens per hour, both doubling on 1 January 2027. Three different billing models wearing the same word. The single “cheapest cached input” column above is the only piece of them that can be lined up.
Context tiers. OpenAI charges a different rate once a request crosses into its long-context band — GPT-5.6 Luna goes from $0.20/$1.20 to $0.40/$1.80. DeepSeek publishes one rate for its full 1M-token window, whichever band you are in. For long-document work that difference compounds quietly.
The free tier. Google publishes a free tier for Gemini alongside the paid one, which no amount of per-token comparison captures.
List price is not the bill. Reasoning tokens you never see, cache-hit rates in the real world, retries, and per-request overheads all sit between the rate card and the invoice. We took that apart separately in what Opus 5, Fable 5 and GPT-5.6 Sol actually cost , and the reasoning there applies to every model here.
And one thing this article deliberately does not do: rank these models on quality. Price per token says nothing about how many tokens a model needs to get an answer right, which is the number that actually decides cost per finished task.
So is DeepSeek V4 Flash the cheapest?
On output and on cached input, off-peak, yes — by a wide margin. On plain input, no: GPT-5.6 Luna’s $0.20 now undercuts DeepSeek’s $0.22.
Off-peak and on output, nothing here is close to it: Luna, the nearest, is 1.8x. On input Luna is now the cheaper of the two in both of DeepSeek’s bands. On cached input its peers charge 2.9x to 28.6x more off-peak, and 1.4x to 14.3x against DeepSeek’s peak rate. That matters most for agent-style workloads that re-send the same context repeatedly.
The honest summary is narrower than the headline. DeepSeek V4 Flash is priced like a tier below the models it actually scores alongside, and its advantage is widest on output tokens and cache hits — which is where reasoning-heavy and context-resending workloads tend to spend, though not where every workload spends.
Sources
Each vendor’s own pricing page, all read on 26 August 2026.
| Source | What it supports here |
|---|---|
| Artificial Analysis model index | The Intelligence Index scores used to match models on measured capability rather than tier name, read on each model’s own page under index version v4.1.1. The listing defaults to 25 of 589 models, so per-model pages carry scores it omits — including Claude Haiku 4.5 and Claude Sonnet 4.6 |
| DeepSeek API pricing, archived 1 May 2026 | The earlier V4 Pro rates and their “75% off” markers — $0.435 against $1.74 and $0.87 against $3.48 — which the current page no longer shows |
| DeepSeek API pricing, archived 31 July 2026 | V4 Flash at $0.14 / $0.28 / $0.0028 with no discount marker, as the page stood on the day this article was first published, and the peak-pricing footnote as it then read |
| DeepSeek API pricing | V4 Flash at $0.22 / $0.66 / $0.007 off-peak and $0.44 / $1.32 / $0.014 at peak, V4 Pro’s rates, the DeepSeek-V4-Flash-0731 and DeepSeek-V4-Pro-0813 versions, the 1M context and 384K max output, concurrency limits, and the peak/off-peak bands with their UTC hours |
| OpenAI API pricing | GPT-5.6 Luna at $0.20 / $1.20 short-context and $0.40 / $1.80 long-context on the Standard tier, and the Standard / Batch / Flex / Fast mode split |
| Anthropic API pricing | Claude Haiku 4.5 at $1 / $5 with $0.10 cache hits, the note cancelling the Sonnet 5 increase that had been dated 1 September 2026, and Fable 5 at $10 / $50 |
| Google Gemini API pricing | Gemini 3.6 Flash at $0.75 / $3.75 through 31 December 2026 and $1.50 / $7.50 from 1 January 2027, Gemini 3.5 Flash-Lite at $0.30 / $2.50, the caching and storage rates, the free tier, and the note that output prices include thinking tokens |
Anthropic’s pricing documentation is reached through redirects from docs.anthropic.com and docs.claude.com; the final URL is the one cited. Vendor pricing pages change without notice — every current price here is as published on 26 August 2026, and the two earlier sets of DeepSeek rates are cited to the dated archive captures above. The Intelligence Index scores carry the same date: they were re-read on 26 August 2026, under index version v4.1.1, and replace the 31 July figures this article previously carried. Every table carries the date its own figures were read.
How we verified this
Every price in this article comes from the vendor’s own published pricing page, read on 26 August 2026, not from an aggregator or a price-tracking site. Several sites that rank well for these queries are not the vendors — deepseek.ai and chat-deep.ai are not DeepSeek’s domain, which is deepseek.com — and nothing from them is used here. The Intelligence Index scores were re-read on the same day, on each model’s own page. Artificial Analysis has re-versioned the index to v4.1.1 since this article was first written, and that version publishes whole numbers where the old one carried two decimals; every score here is the v4.1.1 reading, and the 31 July figures this article previously carried have been replaced rather than converted.
Prices are list prices per million tokens in US dollars, for the standard service tier. That qualifier matters more than it looks. OpenAI’s page carries four tiers behind a toggle — Standard, Batch, Flex and Fast mode — and we read the Standard table, confirmed by checking which radio was selected rather than assuming the first table shown was the default. OpenAI also splits every model into short-context and long-context columns at different prices; the short-context figures are used and labelled. Google lists a free tier alongside the paid tier, and its output price is stated to include thinking tokens. Anthropic quotes per MTok and lists cache writes separately from cache hits.
Because of those differences, the only figures compared directly here are base input, output, and the cheapest cached-input rate each vendor publishes. Caching is not otherwise comparable across the four: DeepSeek prices a cache hit automatically, Anthropic charges separately for 5-minute and 1-hour cache writes, and Google charges a storage fee of $0.50 per million tokens per hour on top of its cache rate for Gemini 3.6 Flash, rising to $1.00 on 1 January 2027. Anywhere those mechanisms would need to be added together to compare, we have not compared them.
Model selection is the part of this article most likely to be wrong if done carelessly, so it is not done by tier name. Candidates are chosen by score on the Artificial Analysis Intelligence Index, an independent evaluator running one methodology across vendors. On the v4.1.1 reading that choice needs no judgement layered on top: Anthropic’s nearest model to DeepSeek V4 Flash is Sonnet 5, which is also the cheaper of the two Sonnets, so the tables use Sonnet 5 and the article shows Sonnet 4.6 beside it. On the older version of the index the two Sonnets sat either side of V4 Flash and choosing between them took an argument; they no longer do, and that argument has been removed rather than restated. Vendor-published benchmark numbers are not used, because they are produced on different harnesses at different reasoning settings and cannot be compared with each other. An earlier version of this article matched by name, which put Gemini 3.5 Flash-Lite in the peer set — fifteen points below DeepSeek V4 Flash on the current reading — and Claude Haiku 4.5 in it, which sat far below the whole peer set on the reading taken then. A later version claimed the index carried no score for Haiku 4.5 at all; it does, on the model’s own page. That claim came from scraping the index’s default view, which the page itself labels “25 of 589 models”, and mistaking it for the whole catalogue. Neither Haiku 4.5 nor Sonnet 4.6 appears in that default view; both have scores on their own pages.
The index has a limit we state rather than smooth over: reasoning effort is not constant across its entries. DeepSeek V4 Flash, GPT-5.6 Luna and Claude Sonnet 5 are labelled “(max)” while Gemini 3.6 Flash’s own page labels it “(high)”, so those scores are not strictly like for like — and because effort drives output tokens, that caveat cuts directly into the price comparison rather than sitting beside it.
These are list prices, not a bill. Reasoning tokens, cache-hit rates, context growth and per-request overheads all move the real number, and our separate piece on what LLMs actually cost covers that. This article does use a capability measure — the Artificial Analysis index — but only to decide which models belong in the same comparison. It does not run or reproduce any benchmark of its own, does not rank these models on quality beyond that one index, and is not investment or purchasing advice.