Drawpie Explainers

Claude Opus 5.5 vs GPT-6 Sol: One Shared Benchmark, Two Price Sheets

Claude Opus 5.5 vs GPT-6 Sol: One Shared Benchmark, Two Price Sheets
Photo by Ludovico Ceroseis on Unsplash
Key takeaways
  • 🔑 Claude Opus 5.5 and GPT-6 Sol both launched on Tuesday 22 September 2026, and neither launch post compares against the other. The only eval both vendors report is FrontierCode v1.1 Main, and its maintainer, Cognition, lists both models on one leaderboard.
  • On Cognition’s board, read 23 September 2026, Opus 5.5 scores 54.6% at its default medium effort against GPT-6 Sol’s 45.9%, for a mean spend per rollout of $0.80 against $0.77. At max effort it is 54.4% against 49.3%, at $6.19 against $2.07.
  • GPT-6 Sol lists at exactly half Opus 5.5’s per-token price, $2 and $10 per million input and output tokens against $4 and $20, and both charge $0.20 per million for cache reads. Above 272,000 input tokens, OpenAI bills the whole Sol request at $4 input, $0.40 cached and $15 output.
  • 🔴 Half the per-token price is not half the bill. Measured per task, the gap runs from roughly level (FrontierCode at medium) to about 5.6 times (Artificial Analysis’s index at max effort, $5.98 against $1.06), because the two models spend very different numbers of tokens.
  • Opus 5.5’s thinking cannot be switched off and most cybersecurity tasks sent to it are re-routed to Opus 4.8. GPT-6 Sol accepts a reasoning effort of none, and takes at most 922,000 input tokens inside its 1,050,000-token window.
  • This page was researched and drafted with the help of Claude, the model family that includes Opus 5.5. It names no overall winner, and it labels every vendor-reported figure as vendor-reported.

Claude Opus 5.5 and GPT-6 Sol launched on the same day, Tuesday 22 September 2026, and neither vendor’s launch post measures its model against the other’s. The one test both report is FrontierCode v1.1 Main, where the benchmark’s own leaderboard has Opus 5.5 at 54.6% and GPT-6 Sol at 45.9% at their default medium effort, for about the same spend. On list price, GPT-6 Sol costs exactly half as much per token, up to 272,000 input tokens.

This page was researched and drafted with the help of Claude, an Anthropic model family that includes one of the two models compared, so every figure here is sourced to a vendor page or an independent evaluator, and neither vendor’s self-reported numbers are presented as independent. We name no overall winner.

Is there a head-to-head benchmark for Opus 5.5 and GPT-6 Sol?

Only one so far: FrontierCode v1.1 Main, a coding benchmark maintained by Cognition, which added both models to its public leaderboard on 22 September 2026. Read on 23 September, the board has Opus 5.5 at 54.6% and GPT-6 Sol at 45.9% at their default medium effort, with a mean spend per rollout of $0.80 and $0.77.

EffortOpus 5.5GPT-6 Sol
Low47.3%, $0.4037.3%, $0.43
Medium54.6%, $0.8045.9%, $0.77
High54.0%, $1.0947.7%, $1.04
Xhigh51.4%, $2.2548.4%, $1.32
Max54.4%, $6.1949.3%, $2.07

The rows ran in different harnesses, which Cognition names as claude-code for Opus 5.5 and codex for GPT-6 Sol, so this partly tests each vendor’s own coding agent. On best runs, the board ranks Opus 5.5 first and GPT-6 Sol seventh.

Opus 5.5 gains nothing above medium: max costs nearly eight times as much per rollout and scores 0.2 points lower. GPT-6 Sol improves at every step, and its max run costs about a third of Opus 5.5’s. It is one benchmark, measuring one kind of work.

What do Anthropic and OpenAI each claim?

Each vendor reports its own set of tests, and apart from FrontierCode they do not overlap. Anthropic’s headline figures for Opus 5.5 include 66.4% on Terminal-Bench 4.0; OpenAI’s for GPT-6 Sol include 68.8% on DeepSWE v1.1. Anthropic’s table leaves GPT-6 Sol out, OpenAI’s charts compare against Claude Opus 5, Fable 5.1 and Fable 5, and every figure here is vendor-reported.

TestOpus 5.5GPT-6 Sol
Terminal-Bench 4.066.4% at xhighNot reported
CursorBench 4.057.8% at maxNot reported
Humanity’s Last Exam, with tools67.7%Not reported
Terminal-Bench-Science 0.158.7%Not reported
DeepSWE v1.1Not reported68.8% at max
Agents’ Last Exam V1Not reported56.4% at max

Anthropic’s figures are at max effort unless marked; CursorBench falls to 52.5% at the default medium. Anthropic also ran its benchmarks with production safeguards on, and says that when they intervened, Opus 4.8 finished cyber tasks and Opus 5 finished bio and frontier-LLM-development tasks. Some Opus 5.5 scores therefore include work done by other models.

Both posts also report OSWorld 2.0 scores, left out here on purpose: the same older model, Opus 5, scores 74.0% in Anthropic’s table and 70.2% in OpenAI’s chart, so the setups differ.

What do independent evaluators say?

Artificial Analysis, which runs both models through its own harness, scores Opus 5.5 at 58 and GPT-6 Sol at 48 on its Intelligence Index v4.3.2, both at max effort, and puts the cost per Intelligence Index task at $5.98 against $1.06. Opus 5.5 leads on all ten component evals shown on its comparison page, three of them by a single point. Zapier has also scored both models; ARC Prize, so far, only Opus 5.5.

EvalOpus 5.5GPT-6 Sol
Terminal-Bench 4.060%44%
Humanity’s Last Exam61%48%
SciCode67%58%
AutomationBench-AA70%62%
GDPval-AA v2.11846 Elo1487 Elo
AA-Briefcase18221483
AA-Omniscience4627
AA-LCR v1.185%84%
CritPt32%31%
GDP.pdf26%25%

The same data shows Opus 5.5 used 260 million output tokens across the index, GPT-6 Sol 77 million. Artificial Analysis puts GPT-6 Sol level with GPT-5.6 on the index at about half the cost per task, with its GDPval-AA score down about 100 Elo. The index was revised four times in September 2026, and in our earlier comparison of GPT-6 Astra and Claude Fable 5.1 we watched both headline scores move within days. Treat the version number as part of the score.

Zapier, which runs AutomationBench, has both models on its 1.0.6 leaderboard, priced the same way. Opus 5.5 reaches 40.0% at max for $1.28 a task, 34.4% at xhigh for $0.80 and 32.0% at high for $0.65. GPT-6 Sol reaches 33.2% at xhigh for $0.27 and 32.0% at max for $0.34. GPT-6 Astra tops the board at 41.4%. Anthropic notes that Zapier ran without fallback models, so any safeguard intervention counted as a failure.

ARC Prize has a verified page for Opus 5.5, dated 22 September: 98.5% on ARC-AGI-1 and 93.3% on ARC-AGI-2’s semi-private set, both at high effort, with no ARC-AGI-3 score. It has no GPT-6 Sol page yet. On 23 September, neither model was ranked on arena.ai’s text leaderboard or listed on the official Terminal-Bench 4.0 leaderboard at tbench.ai.

How much do Opus 5.5 and GPT-6 Sol cost per token?

GPT-6 Sol lists at exactly half Opus 5.5’s per-token price: $2 per million input tokens and $10 per million output, against $4 and $20. Cache reads cost $0.20 per million on both. That holds only up to 272,000 input tokens; above it, OpenAI bills the whole Sol request at $4 input, $0.40 cached and $15 output, while Anthropic charges Opus 5.5 one rate across its full 1M-token window.

RateOpus 5.5GPT-6 Sol
Input$4$2; $4 above 272K
Output$20$10; $15 above 272K
Cache read$0.20$0.20; $0.40 above 272K
Cache write$5 for 5 min; $8 for 1 hr$2.50; $5 above 272K
Batch$2 in, $10 out$1 in, $5 out, also Flex
Fast mode$8 in, $40 outTwice the standard rate
Regional1.1x for US-only inferencePlus 10%

Opus 5.5’s fast mode is a research preview, on Anthropic’s own API only and not with Batch. Its cache reads are 0.05 times the input price; Anthropic’s other models charge 0.1 times. Reasoning tokens bill as output on both, as our July breakdown of true LLM cost explains.

Both discount claims use each vendor’s own earlier price. OpenAI’s 50% cut is against GPT-5.6 Sol’s promotional $4 and $20, which OpenAI says runs at least through 21 November 2026. Anthropic says Opus 5.5 at default settings will cost 40% less than Opus 5 ($5 and $25) on typical workloads; we have not tested that.

What does a real workload cost on each model?

With identical token counts, Opus 5.5 costs twice as much as GPT-6 Sol for ordinary prompts and 1.75 times as much for a cache-heavy agent turn, or $7.00 against $4.00 over 50 turns. Above 272,000 input tokens the gap almost closes, and a repeat question against a large cached document is cheaper on Opus 5.5.

JobOpus 5.5GPT-6 Sol
Chat: 2K in, 1K out$0.028$0.014
Agent turn: 100K cached, 10K fresh, 4K out$0.14$0.08
272,000-token prompt, 5K out$1.188$0.594
272,001-token prompt, 5K out$1.188$1.163
500K uncached, 5K out$2.10$2.075
Question on cached 500K doc: 2K fresh, 5K out$0.208$0.283
Batch: 1M jobs, 1K in, 200 out$4,000$2,000

One extra input token past 272,000 nearly doubles the Sol bill, from $0.594 to $1.163.

The equal-token assumption is the weak point. Neither vendor publishes how the same text tokenizes on the other’s model, and the two spend very different amounts of output. Measured per task instead, the gap runs from roughly level to more than fivefold:

  • FrontierCode at medium: $0.80 against $0.77 per rollout.
  • AutomationBench at near-equal scores: Opus 5.5 at high (32.0%) for $0.65, GPT-6 Sol at xhigh (33.2%) for $0.27, about 2.4 times.
  • FrontierCode at max: $6.19 against $2.07, about three times.
  • Artificial Analysis index at max: $5.98 against $1.06, about 5.6 times.

What API differences will developers notice?

The one most likely to bite is that Opus 5.5 cannot run with thinking switched off. Anthropic lists its thinking as adaptive and always on, with medium as the default effort, while GPT-6 Sol accepts a reasoning effort of none. Sol also caps input at 922,000 tokens inside its 1,050,000-token window.

SpecOpus 5.5GPT-6 Sol
API IDclaude-opus-5-5gpt-6-sol
Context1M tokens1,050,000 tokens
Max output128K tokens128,000 tokens
EffortLow, medium, high, xhigh, maxNone, low, medium, high, xhigh, max
DefaultMediumMedium
CutoffJune 202620 April 2026
  • On Chat Completions, GPT-6 Sol supports function calling only at reasoning effort none.
  • Anthropic says most cybersecurity tasks sent to Opus 5.5 will be re-routed to Opus 4.8. OpenAI’s Preparedness assessment treats GPT-6 Sol as High capability in cybersecurity and in biological and chemical, in an appendix added to the GPT-6 Astra system card on 22 September.
  • Preserved thinking, which ties Opus 5.5’s reasoning blocks to the model and conversation that produced them, is enforced for API accounts created on or after 31 August 2026. Zero data retention is available.
  • Opus 5.5 is on the Claude Platform, AWS, Google Cloud and Microsoft Azure; GPT-6 Sol is in the API, in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise and Edu users (not yet in Chat), and in GitHub Copilot since 22 September.

What should you do?

Test on your own workload; no published number settles it.

  • Set effort explicitly on both. On FrontierCode, Opus 5.5 scored as well at medium as at max for about an eighth of the spend, while GPT-6 Sol gained at every step, so test Sol at more than one setting.
  • Count billed tokens, not prompt length: each model’s tokenizer and reasoning set the real bill.
  • Check where your prompts sit against 272,000 tokens. Above it, GPT-6 Sol’s price advantage mostly disappears, and repeated queries over one large cached document cost less on Opus 5.5 at list price.
  • Pick Sol if you need calls with no reasoning at all. Opus 5.5 does not offer that mode.
  • Re-check around 7 October 2026, when the independent boards may have moved or caught up.
How we verified this

🔴 This page was researched and drafted with the help of Claude, an Anthropic model family that includes Claude Opus 5.5, one of the two models it compares. That is a conflict of interest, and it is stated in the body before the first heading as well as here. What follows from it: every figure is sourced to a vendor page or an independent evaluator, neither vendor’s self-reported numbers are presented as independent, and no overall winner is named. Where independent data favors Opus 5.5 on score, the same sources’ findings that it costs more per task and spends far more output tokens are printed alongside, not after.

The FrontierCode figures were read from Cognition’s own leaderboard at cognition.com/frontiercode on 23 September 2026, not spliced from the two vendors’ launch charts. Our research pass first compared chart points taken from each vendor’s post and described both medium-effort runs as costing about $0.80. The adversarial re-check went to the benchmark’s operator instead, whose changelog records both models added on 22 September 2026, and corrected three things: GPT-6 Sol’s medium run is $0.77, not $0.80; Cognition’s unit is mean USD spend per rollout, not per task; and the two rows ran in different harnesses, claude-code for Opus 5.5 and codex for GPT-6 Sol. It also found why the vendor charts drift. Anthropic’s chart matches Cognition exactly for Opus 5.5 but carries older competitor costs that appear to predate Cognition’s 10 September pricing fixes; OpenAI’s chart shows its own GPT-6 Sol about 3% to 5% costlier than the leaderboard does, with identical scores.

Prices and specs were checked on 23 September 2026 against Anthropic’s pricing page, models overview and launch post, and OpenAI’s GPT-6 Sol model page, pricing page, API changelog and launch post. OpenAI’s launch post returned 403 to automated fetches and was read in a browser. GPT-6 Sol’s 922,000-token maximum input appears only on the Markdown version of OpenAI’s model page. ⚠️ Artificial Analysis lists GPT-6 Sol’s context window as 872k, a figure OpenAI’s pages do not support and nobody has explained; this page uses OpenAI’s 1,050,000 total and 922,000 input.

⚠️ The worked costs were computed from official list prices only, assuming identical token counts on both models. That assumption is the weak point. Neither vendor publishes how the same text tokenizes on the other’s model; Anthropic’s own docs say its tokenizer since Claude 4.7 produces roughly 30% more tokens than its previous one, which is a comparison with itself, not with OpenAI. OpenAI lists a $2.50 cache-write rate for GPT-6 Sol but does not say whether it applies to all uncached input or only to explicit cache breakpoints; computing the agent turn both ways moves the ratio from 1.75x to 1.76x. The long-document examples exclude the one-time cost of writing the document to cache.

Artificial Analysis figures are from Intelligence Index v4.3.2, read from its model pages and its Opus 5.5 vs GPT-6 Sol comparison page on 23 September 2026, both models at max effort. The Opus 5.5 entry is labeled Adaptive Reasoning, Max Effort, Default Fallback. The $5.98 and $1.06 figures are Artificial Analysis’s cost per Intelligence Index task, which its methodology recomputes from live cache-hit rates, so they can drift without a re-run. ⚠️ The index was revised four times in September 2026 (v4.2, v4.3, v4.3.1, v4.3.2), and v4.3.2 re-anchored the GDPval-AA Elo scale. The re-check also corrected a label: the 1846 GDPval-AA score in Anthropic’s launch table is Artificial Analysis’s figure, which Anthropic quotes, not Anthropic’s own measurement.

⚠️ OSWorld 2.0 appears in both launch posts and is deliberately kept out of the comparison tables. Anthropic reports Opus 5.5 at 81.8%, marked partial, at max effort, with no test set or version given. OpenAI reports GPT-6 Sol at 60.5% at xhigh in its text and 64.4% at max in its chart data, on the partial reward of the offline set from the v2026.08.08 release. The same older model, Opus 5, scores 74.0% in Anthropic’s table and 70.2% in OpenAI’s chart, so the setups differ and the numbers cannot be set side by side.

Two smaller attributions were corrected by the re-check. Anthropic’s statement that production safeguards were on during its benchmark runs sits in the paragraph under its table, not in its footnotes. And Anthropic’s table values are at max effort unless marked, so its 57.8% on CursorBench 4.0 is a max-effort score; at the default medium, Anthropic’s chart note gives 52.5%.

⚠️ What we could not confirm is left out. The exact publication times of both launch posts could not be established, so no launch time is given. The OSWorld-v2, DeepSWE and Agents’ Last Exam leaderboards render only with JavaScript and could not be checked for either model. Secondary reports of a deadline and a percentage attached to Anthropic’s subscriber rate-limit changes do not appear in Anthropic’s post and are not repeated. Anthropic’s claim that Opus 5.5 costs 40% less than Opus 5 on typical workloads is reported as its claim.

🔴 This page names no overall winner and carries no forecast. It makes no prediction about either vendor’s future pricing or future leaderboard placings, and it quotes the end date of GPT-5.6 Sol’s promotional pricing only as OpenAI states it.