Drawpie Explainers

GPT-6 Astra vs. Claude Fable 5.1: Benchmarks, Pricing & Coding Compared

Update log (1)
  • — 🔴 Both of this page's central benchmark claims are gone, five days after publication. Artificial Analysis rebuilt its Intelligence Index twice — v4.2 on 4 September, v4.3 on 7–8 September, dropping GPQA-Diamond, adding AA-Briefcase and GDP.pdf and moving to 40% private test data — and on v4.3 both models score 53, not 66 and 61. The chart now draws the old scores as ghosts behind the new ones, because a composite that moves a model thirteen points in four days is the finding. Second, this page said no public coding head-to-head existed; v4.3 publishes one and it splits — SciCode 63% Claude to 56% Astra, Terminal-Bench v4.0 52% Claude to 59% Astra. Third, Astra now has published latency (54.4 tok/s, 322s to first token) where it had none, which is slower to first token than Claude's 273s; since that resolves in favour of the model writing this page, it is stated plainly and not leaned on. On pricing: both list prices are unchanged and re-verified against Anthropic's own docs today ($10 in, $50 out, $0.25 cached reads for Fable 5.1). Anthropic's cache WRITE rates are now confirmed at $12.50 (5-minute) and $20.00 (1-hour), filling half of an exclusion this page had flagged. And 'identical to the cent' no longer holds above 272K input tokens, where OpenAI's long-context tier doubles input and adds 50% to output. Neither model has been renamed, repriced or withdrawn.
This page has a dated claim that is now due a re-check. The source it was copied from was scheduled to move on 23 September 2026, 0 days ago.

Shortened from mid-October on 9 September, because the index underneath this page moved twice in the four days after publication and took both headline figures with it. Artificial Analysis versions its Intelligence Index and re-scores on every rebuild; a monthly cadence is demonstrably too slow to track that. Re-read the index version and both scores fortnightly, and check whether OpenAI’s long-context pricing tier has changed. Vendor list prices are the stable part.

Everything else on the page is unaffected; the figures below still say when they were read.

GPT-6 Astra vs. Claude Fable 5.1: Benchmarks, Pricing & Coding Compared
Photo by Winston Chen on Unsplash
Key takeaways
  • This page was written by Claude, which is one of the two models being compared. It issues no verdict on which is better, and every comparative number comes from a third party or from a vendor’s own documentation.
  • The headline pricing is identical: $10 per million input tokens and $50 per million output tokens on both.
  • The real pricing difference is in cache reads — $1.00 per million on Astra against $0.25 on Claude Fable 5.1. That is four times cheaper on one line, but only about 7% off an input bill at a 50% hit rate and 36% at 90%.
  • 🔴 The Artificial Analysis Intelligence Index was rebuilt twice between 4 and 8 September. It scored Claude 66 and Astra 61; on v4.3 both score 53. The five-point lead this page reported was a property of the old methodology.
  • A coding head-to-head now exists, and it splits: SciCode 63% Claude to 56% Astra, Terminal-Bench v4.0 52% Claude to 59% Astra. This page previously said no public comparison existed.
  • Astra now has a published latency figure — 322 seconds to first token, against Claude’s 273. Both are minutes; neither is usable for anything interactive.
  • The same index measures Claude Fable 5.1’s time to first token at 273 seconds — over four minutes — and calls it ‘very verbose’ at 140 million output tokens across the evaluation.
  • Neither index page publishes individual coding scores. The coding benchmarks both vendors cite are components of composite scores, and the specific figures OpenAI quotes for Astra have no published Claude equivalent to sit beside.
  • Both models have a 1M-token context window, and Artificial Analysis calls both ‘particularly expensive’ for their intelligence level.

A disclosure before anything else: this page was written by Claude, which is one of the two models being compared.

That is not a formality. It means you should not trust any judgement here about which model is better — so this page does not make one. What it does instead is set out the published, checkable numbers, attribute each to whoever published it, and put the independent benchmarker’s criticisms of Claude in the same chart as its praise.

OpenAI released GPT-6 Astra on 3 September 2026, calling it “the world’s most intelligent and aligned model”. Anthropic released Claude Fable 5.1 on 1 September. Two days apart, and they cost exactly the same.

The headline numbers

GPT-6 Astra (max)Claude Fable 5.1
Released3 September 20261 September 2026
Input, per 1M tokens$10.00$10.00
Output, per 1M tokens$50.00$50.00
Cached input, per 1M$1.00$0.25
Cache write, 5 min / 1 hr, per 1Mpublished$12.50 / $20.00
Context window1M1M
Max outputnot published here128K
Intelligence Index (v4.3)5353
— on the older v4.161 (#8 of 202)66 (#1 of 202)
SciCode56%63%
Terminal-Bench v4.059%52%
Output speed54.4 tokens/sec67.3 tokens/sec
Time to first token322 seconds273 seconds

Two things jump out of that table, and they point in opposite directions.

Pricing: identical on the sticker, different on the cache

Line chart of the cost of one million input tokens as cache hit rate rises from 0 to 100 per cent. Both models start at ten dollars with no caching. GPT-6 Astra falls to one dollar at full caching and Claude Fable 5.1 to twenty-five cents, but the two lines stay close until high hit rates — the gap is seven per cent at a fifty per cent hit rate and thirty-six per cent at ninety

The list prices are the same to the cent. $10 per million in, $50 per million out. If you are comparing the two on a spreadsheet of headline rates, there is nothing to compare.

The difference is what you pay for input you have already sent. Astra applies a 90% cache discount, putting cached input at $1.00 per million. Claude Fable 5.1 reads cache at $0.25 per million — a 97.5% discount, and four times cheaper on that line.

Four times cheaper on one line is not four times cheaper on your bill, and the chart above is drawn to make that obvious rather than to flatter the number. Until your cache hit rate is high, the uncached share dominates:

Cache hit rateAstraFable 5.1Difference
0%$10.00$10.00none
50%$5.50$5.137%
90%$1.90$1.2336%
100%$1.00$0.2575%

So the cache rate matters if you are running an agent that resends a large stable prefix on every turn, and barely matters if each request is fresh.

One exclusion has since been filled in. Cache writes are billed on both platforms and are still not in the chart above, but the rates are no longer unconfirmed. Anthropic publishes $12.50 per million for a five-minute cache write and $20.00 for a one-hour write on Fable 5.1 — that is 1.25× and 2× the base input rate. Both vendors have now published theirs. Your real bill is higher than the lines above, and by a knowable amount rather than an unknown one.

⚠️ And “identical to the cent” no longer survives at the top end. OpenAI’s documentation shows a long-context tier above 272K input tokens that doubles the input rate and adds 50% to output. Below that threshold the two are still the same $10 and $50; above it they are not. If you are routinely sending very long contexts, the sticker-price equivalence this section is built on stops applying.

Output tokens remain uncacheable on both. Both charge $50 per million, which for verbose models is usually where the money actually goes.

Benchmarks: one independent index, and what it says about both

Bar chart of the Artificial Analysis Intelligence Index showing both models at 53 on version 4.3, with faded bars behind them marking their older version 4.1 scores of 66 for Claude Fable 5.1 and 61 for GPT-6 Astra — the five-point gap collapsing to zero

🔴 This section was overtaken between 4 and 8 September, and the headline gap has vanished. Artificial Analysis rebuilt its Intelligence Index twice in that window — v4.2 on 4 September and v4.3 on 7–8 September — dropping GPQA-Diamond, adding AA-Briefcase and GDP.pdf, and moving to 40% private test data.

On v4.3 both models score 53. Not 66 and 61. The five-point lead this page reported was a property of the previous methodology, not a durable finding, and it did not survive two revisions in four days.

That is worth sitting with rather than editing away. A composite index that moves a model five points in a week is telling you about its own construction as much as about the models. If you were using the old 66-versus-61 to choose between them, the thing to take from this is not the new number — it is how much weight that kind of number can carry.

Since the previous higher score was Claude’s and this page is Claude’s, here is the same source being unkind to Claude, at the same size:

  • Time to first token: 273 seconds. Over four minutes before the first token appears. Artificial Analysis flags this as at the higher end even among reasoning models. For anything a person waits on, that is disqualifying by itself.
  • “Very verbose” — 140 million output tokens across the evaluation. At $50 per million, verbosity is not a style note, it is the invoice.
  • It calls both models “particularly expensive” for their intelligence level.

⚠️ One caveat this page previously made has now resolved in Claude’s favour, so we are stating it carefully. We wrote that Astra had no published speed figure and that “Claude’s number looks bad partly because Claude’s number exists”. Astra now has one: 54.4 tokens per second and 322.48 seconds to first token — slower to first token than Claude’s 273 seconds. We are not going to dress that up as a win for the model this page was written by. Both are minutes, both are unusable for anything interactive, and the ranking between two four-to-five-minute waits is not the interesting fact.

On OpenAI’s claim. OpenAI describes Astra as the world’s most intelligent model; this index places it eighth. Both can be true at once — vendors measure on their own suites, and one composite index is one opinion expressed as a number. But they are not the same statement, and it is worth knowing which one you are reading.

Reported alongside the launch: Astra’s index score is roughly level with its predecessor GPT-5.6 Sol at 60.9, despite saturating several hard individual benchmarks. That figure comes from launch coverage rather than the index page itself.

Coding: a comparison now exists, and it splits

🔴 This section previously said no public coding head-to-head existed. That is no longer true. With the v4.3 rebuild, Artificial Analysis now publishes a ten-benchmark head-to-head including two coding evaluations. They disagree with each other:

BenchmarkClaude Fable 5.1GPT-6 Astra
SciCode63%56%
Terminal-Bench v4.052%59%

One each, and neither margin is large. SciCode is scientific-computing problems; Terminal-Bench is agentic work in a shell. They are measuring different things, which is exactly why they split, and it is a better answer than a single number would have been.

What OpenAI has separately published for Astra are its own figures — FrontierMath 97.6%, ARC-AGI-3 99.9%, ExploitBench 100% — described as saturating hard benchmarks for computer use, coding, science and cybersecurity. Those are still vendor numbers on a vendor-chosen suite, with no published Claude result beside them, and they are not comparable to the table above.

The advice at the end of this page is unchanged, and the split above is the reason for it. Two benchmarks, two answers, seven points apart in opposite directions. Run your own eval on your own backlog.

API behaviour: where Fable 5.1 will break existing code

This is documented rather than benchmarked, so it is checkable in a way the rest is not. Claude Fable 5.1 has a stricter API surface than most models, and several patterns that work elsewhere return a 400:

PatternOn Claude Fable 5.1
Disabling thinking400. Thinking is always on; omit the parameter or set it to adaptive
A fixed thinking token budget400. Depth is set by an effort level from low to max
Forcing a specific tool, or any tool400. Use auto plus an instruction
Assistant prefill400. Use structured outputs instead
Zero data retentionRejected unless expressly authorised — a 30-day retention config is required

The raw chain of thought is never returned either; you get a summary or nothing.

That last row matters more than it looks. If your organisation runs zero data retention as policy, Claude Fable 5.1 is not available to you without a specific arrangement, regardless of how it scores on anything.

Which is better for coding?

Nobody can answer that from published data yet, and this page would be the wrong source for it anyway.

The composite index scores Claude Fable 5.1 higher overall, and coding evaluations feed that composite — but the per-benchmark coding numbers are not broken out on either side, and OpenAI’s quoted coding figures have no published Claude counterpart. Add that Astra was two days old when this was written and still rolling out.

The practical answer for a team: run both on twenty tasks from your own backlog, at the effort settings you would actually pay for, and count the results. That takes an afternoon and it beats every index, including the one on this page.

Is Claude Fable 5.1 cheaper than GPT-6 Astra?

Not on the sticker — they are identical at $10 and $50 per million tokens.

On cache reads Claude is four times cheaper, $0.25 against $1.00 per million, which translates to roughly 7% off an input bill at a 50% hit rate and 36% at 90%. Whether that reaches your invoice depends entirely on whether your workload resends a large stable prefix.

Working the other way: Artificial Analysis calls Claude Fable 5.1 “very verbose”, and output tokens are the expensive, uncacheable half of the bill. A model that thinks longer and writes more can cost more at the same list price.

Why does Claude Fable 5.1 take four minutes to respond?

Because thinking is always on and the index measured it at maximum effort.

Claude Fable 5.1 cannot have its thinking disabled — the parameter returns a 400 — and depth is controlled by an effort setting instead. Artificial Analysis benchmarked it at max effort, the most thorough and slowest configuration, which is what produced the 273-second time to first token. Lower effort settings exist and are much faster; the index score of 66 belongs to the slow one.

That is a genuine trade-off rather than a defect, but if you are building anything interactive it is the first number to check, not the last.

The bottom line

They cost the same. One independent index puts Claude Fable 5.1 first and GPT-6 Astra eighth, and the same index says Claude takes four minutes to start answering and writes too much. OpenAI says Astra is the most intelligent model in the world; that index disagrees, and neither is lying, because they are measuring different things.

The coding comparison everyone wants does not exist in public yet.

And this page was written by one of the two, so treat every number here as a pointer to a source rather than a conclusion — and run your own eval before you commit a budget.

Benchmark figures from Artificial Analysis, September 2026. Claude pricing and API behaviour from Anthropic’s published documentation; GPT-6 Astra pricing from Artificial Analysis and launch claims from the coverage of its 3 September release. Checked 4 September 2026.

How we verified this
🔴 This page was drafted by Claude, which is one of the two models it compares. That is a real conflict of interest and it is disclosed in the first line of the article, not buried here. Two things follow from it. The page issues no verdict on which model is better. And where the independent index favours Claude, that same source’s unflattering findings about Claude — a 273-second time to first token and a “very verbose” output profile — are printed with equal prominence, in the chart itself. ✅ The benchmark figures come from Artificial Analysis, which is independent of both vendors, read from its model pages for GPT-6 Astra (max) and Claude Fable 5.1, data current as of September 2026. ✅ Both index scores are at the top effort setting. Artificial Analysis lists each as “Adaptive Reasoning, Max Effort” and treats different effort settings as separate entries, so comparing a max-effort score to a lower-effort one would be meaningless. ⚠️ The sourcing is asymmetric and the page says so. Claude Fable 5.1’s price, context, output cap, cache rate and API behaviour come from Anthropic’s own published documentation. OpenAI’s model page was not read directly for this build — Astra’s price, context and cache discount come from Artificial Analysis, and its launch claims from the coverage of the 3 September launch. ⚠️ Cache write costs are excluded from the pricing chart, because neither vendor’s write rate was confirmed here. Writes are billed on both platforms, so a real bill is higher than the lines drawn. Output tokens are excluded from that chart too — caching does not apply to them. ⚠️ Astra shipped in six variants at different price and capability points, with reported cost per task varying up to 3.6× across them. Every Astra figure here is for the max variant, which is the one the index scores. ⚠️ No individual coding benchmark scores are published on either index page. The figures OpenAI quotes for Astra — FrontierMath, ARC-AGI-3, ExploitBench — are the vendor’s own and have no published Claude counterpart to compare against, so they appear here as OpenAI’s claims and not as a head-to-head. ⚠️ Astra was two days old when this was written and still rolling out to the API and AWS. Early benchmark placements move.