Claude Opus 5.5 vs GPT-6.1 Sol: Effort Flips the Winner
Opus 5.5 wins at max effort, GPT-6.1 Sol at medium, and Sol costs less per task in every published row. List prices, a cost-per-turn model, and the catch.

Two flagship models shipped seven days apart. A comparison published 1 October argues the winner flips depending on one setting: reasoning effort.
Claude Opus 5.5 arrived 22 September 2026 from Anthropic. GPT-6.1 Sol, OpenAI's upgraded mid-tier model, landed at DevDay on 29 September 2026. The guide's short answer: pick Sol for high-volume, cost-sensitive work, and Opus 5.5 for long agentic coding and tasks where a wrong answer costs more than the tokens.
One disclosure first, because it shapes how to read the numbers below. The author states plainly: "I work on apimodels.app, a multi-model API gateway that serves both models." List prices are taken from the vendors' own pages on 1 October 2026; gateway prices are that gateway's own.
The spec sheet
| Claude Opus 5.5 | GPT-6.1 Sol | |
|---|---|---|
| Released | 22 Sep 2026 (Anthropic) | 29 Sep 2026 (OpenAI DevDay) |
| List price, input / output | $4 / $20 | $2 / $10 |
| Cached input (list) | $0.20 read, $5 write | $0.10 read, $2.50 write |
| Context / max output | 1M / 128K | 1.05M / 128K |
| Knowledge cutoff | June 2026 | 30 April 2026 |
| Reasoning control | effort low to max, default medium, always thinks | reasoning_effort low to max, default medium, no none |
| Long-context surcharge | — | above 272K input: input and cache 2×, output 1.5× |
Prices are per 1M tokens. Three rows here change code rather than budgets.
The first is "always thinks." Opus 5.5 has no off switch for reasoning, so its reasoning tokens land in the output bill even on short answers. The second is Sol's missing none setting — if you have code that relies on reasoning_effort: none, the guide names that as the one case Sol is explicitly not for. The third is the 272K threshold: cross it and Sol's input and cache double while output rises by half, which quietly erases its price advantage on very long contexts.
The benchmarks, and who ran them
OpenAI's GPT-6.1 Sol launch included Opus 5.5 in several rows. These are OpenAI's runs and OpenAI's choice of benchmarks, so the guide tells readers to treat them as a vendor's claim. On that basis:
| Benchmark (OpenAI launch) | Effort | GPT-6.1 Sol | Claude Opus 5.5 |
|---|---|---|---|
| AutomationBench | medium | 31.7% | 29.5% |
| AutomationBench | max | 36.1% ($0.30/task) | 42.5% ($1.44/task) |
| GDP.pdf | medium | 30.0% ($0.34/task) | 25.6% ($0.80/task) |
| Terminal-Bench Science 0.1 | max | 57.0% ($5.47/task) | 63.3% ($23.21/task) |
The pattern is consistent across all four rows. At medium effort Sol edges ahead. At max effort Opus 5.5 pulls ahead. And Sol costs less per task in every row, including the ones it loses.
That is the whole decision in one shape: effort level, not model name, decides which one wins. A team that runs everything at default medium and a team that runs everything at max are not choosing between the same two models.
Independent aggregates point the same way on capability. The guide reports max-effort figures of 58 for Opus 5.5 and 52 for GPT-6.1 Sol on the Artificial Analysis Intelligence Index, with cost per index task of $5.98 and $0.72 respectively — a six-point capability gap against roughly an eight-fold cost gap.
Anthropic's own launch numbers for Opus 5.5 — 66.4% on Terminal-Bench 4.0 and 57.8% on CursorBench 4.0 — carry no GPT-6.1 Sol row, so as the guide notes, they do not settle the head-to-head. Treat them as a floor on what Anthropic claims, not as a comparison.
What one agent turn actually costs
Benchmark cost-per-task figures are hard to map onto a real bill. The guide models a coding-agent turn instead: 50,000 input tokens, 40,000 of them cached, returning 2,000 output tokens.
| Per turn | Claude Opus 5.5 | GPT-6.1 Sol |
|---|---|---|
| At list price | $0.088 | $0.044 |
| On apimodels.app | $0.053 | $0.022 |
The second row is the author's own gateway pricing, published by a party with an interest in it. The list-price row is the one to plan against: a clean 2× gap on this particular shape of call.
Two things move those numbers in practice, per the guide. Opus 5.5 always thinks, so reasoning tokens show up in the output bill even on short answers. GPT-6.1 Sol at low effort spends far fewer reasoning tokens, and that is where most of its cost-per-task advantage comes from.
Note what this implies: the published 2:1 input-price ratio is the floor of the gap, not the ceiling. The real spread depends on how much each model thinks on your traffic, which no price table can tell you.
Effort is doing more work than the spec sheet shows
A separate write-up published the same day runs three autonomous GPT-6.1 Sol builds — a sheet-cutting layout tool, a document-comparison tool and a traffic solver — and notes that batch ran at "Extra High reasoning."
The interesting part is the author's own limit: "These portfolios cannot establish a model's taste or superiority." What the review changed in every case was not whether the output was correct, but whether a person could inspect and use it — a cut sequence that distinguishes identical sheets, navigation that is not offscreen, a solver that exposes a legible network instead of raw JSON.
That is a useful corrective to the benchmark tables. At the top of the effort range, differences between these models show up as qualities a percentage score does not capture.
Run the test yourself in ten minutes
The guide's closing argument is the right one: benchmarks are someone else's tasks. Its recommended method is to send the same twenty prompts to both models and compare answers and cost side by side.
The published recipe uses one OpenAI-compatible client against both models, tested with Python 3.11 and openai 1.x, reading a per-response billed amount from an x-apimodels-cost header. That header is gateway-specific; if you call the vendors directly, read token counts from the usage object and price them yourself.
Our suggested adjustment, not the guide's: run the twenty prompts at two effort levels, not one. Given that every benchmark row above flips between medium and max, a single-effort test measures the wrong variable.
What to watch
Two things will settle questions this comparison leaves open. First, whether independent evaluators publish Opus 5.5 and GPT-6.1 Sol rows on the same suite at matched effort — every head-to-head table above is currently OpenAI's.
Second, whether Sol's 272K surcharge threshold holds. It is the single line that decides whether the cheaper model stays cheaper for teams running million-token agent contexts.
More from DangMua