AI News
Claude Sonnet 5.5 vs Opus 5.5: Which to Use for Coding
Abhishek Bahukhandi

Claude Sonnet 5.5 vs Opus 5.5 is the first model-choice question in a while where the cheaper option wins a headline coding benchmark. Anthropic shipped Sonnet 5.5 on 28 September 2026 at $2 and $10 per million tokens — half of Opus 5.5 — and it scores higher on Terminal-Bench 4.0. "Just use Opus" became a worse default that day.
We run Claude on both tiers behind Taqari, so this is a comparison we had to actually make rather than a rewrite of a launch post. Below: what the published table says, the effort-level footnote that makes half of it misleading, and why the price gap on a real agent loop is nearer 43% than 50%.
What Anthropic shipped on 28 September
Sonnet 5.5 replaces Sonnet 5 as the mid-tier model, arriving six days after Claude Opus 5.5 on 22 September. The framing in Anthropic's own Claude Sonnet 5.5 announcement is deliberately modest: a clear upgrade over Sonnet 5 that "runs 30%+ faster and costs up to 30% less for most work." The API ID is claude-sonnet-5-5, and it is available on Amazon Web Services, Google Cloud and Microsoft Azure alongside the first-party API.
The rate card, across the whole lineup
Prices are per million tokens, from Anthropic's published models overview:
| Model | Input | Output | Cache read |
|---|---|---|---|
| Claude Fable 5.1 | $10 | $50 | $0.25 |
| Claude Opus 5.5 | $4 | $20 | $0.20 |
| Claude Sonnet 5.5 | $2 | $10 | $0.20 |
| Claude Haiku 4.5 | $1 | $5 | $0.10 |
Look at the cache-read column before anything else. Anthropic prices cache hits at 10% of the base input rate on most models, but at 5% on Opus 5.5. Ten percent of Sonnet's $2 and five percent of Opus's $4 are the same number. A cache hit costs $0.20 per million tokens on both models. That single coincidence is what stops Sonnet from being half the price in practice, and we come back to it below.
What did not change
- Context window: one million input tokens, same as Opus 5.5.
- Max output: 128,000 tokens on the synchronous Messages API, with a 300k beta header available on both.
- Knowledge cutoff: June 2026 across the current lineup.
- Thinking: adaptive on both, steered by the effort parameter rather than a token budget.
- Batch discount: 50% off, unchanged.
So context is not the axis to choose on. Cost, latency and sustained judgement are.
Claude Sonnet 5.5 vs Opus 5.5 on the benchmarks
These are Anthropic's published figures from the Sonnet 5.5 comparison table, which is the one place all three models are scored together under the same conditions. Sonnet 5 is included because it is what most teams are upgrading from.
| Benchmark | Sonnet 5.5 | Opus 5.5 | Sonnet 5 |
|---|---|---|---|
| Terminal-Bench 4.0 | 70.6% | 66.4% | 10.3% |
| FrontierCode 1.1 (Main) | 46.2% | 54.4% | 42.4% |
| CursorBench 4.0 | 55.5% | 57.8% | 34.1% |
| OSWorld 2.1 (computer use) | 80.1% | 81.8% | 57.0% |
| GDPval-AA v2.1 (Elo) | 1844 | 1846 | 1449 |
| AA-Briefcase v1.1 (Elo) | 1811 | 1822 | 1359 |
| Humanity's Last Exam | 64.5% | 67.7% | 54.9% |
Where Sonnet 5.5 wins
One benchmark outright, and it happens to be the one that looks most like a coding agent doing its job. Terminal-Bench 4.0 measures work in a shell — run the test suite, read the failure, edit a file, run it again. Sonnet 5.5 takes it 70.6% to 66.4%, which is a four-point lead at half the price.
The Sonnet 5 column is the other story in that row: 10.3%. A jump from 10.3% to 70.6% in one generation is not a normal increment, and it suggests Sonnet 5 was failing the harness structurally rather than reasoning badly. Treat that cell as "the old model could not do this" rather than as a smooth sixty-point gain.
The Elo rows are the quieter win. On GDPval-AA v2.1 the gap is two points, 1844 against 1846. On AA-Briefcase it is eleven, 1811 against 1822. For a model costing half as much, landing inside the noise on broad professional work is the more useful result than the Terminal-Bench headline.
Where Opus 5.5 still wins
Opus leads everywhere else, and the pattern in where it leads matters more than the margins:
- FrontierCode 1.1 — an 8.2-point gap, the widest on the board. This is the hard, open-ended coding set. If a task has no obvious decomposition, Opus is still worth double.
- CursorBench 4.0 — 2.3 points. Built from real IDE agent sessions, so it rewards multi-file edits that apply cleanly.
- Humanity's Last Exam — 3.2 points. Multidisciplinary reasoning, not coding, but a reasonable proxy for "can it hold a hard problem in its head."
- OSWorld 2.1 — 1.7 points. Close enough that computer-use automation is now a genuine toss-up.
Read together: the gap narrows as tasks get better specified and widens as they get vaguer — a more actionable rule than any single score.
The effort-level footnote that changes the table
Here is the part most comparisons skip, and it is the reason we would not take the table at face value.
Two ways the published table is not apples-to-apples
The two models have different default effort
On the Claude API, Sonnet 5.5 defaults to high effort. Opus 5.5 defaults to medium. If you send the same prompt to both and compare results, you are comparing one model thinking hard against another thinking moderately — and then comparing invoices that reflect that, not the models. Set effort explicitly on both before you conclude anything.
Some published cells are max-effort runs
Anthropic marks the Sonnet 5.5 FrontierCode figure as a maximum-effort result. A max-effort number and a default-effort number in adjacent columns do not describe the same experiment. This is not sleight of hand, but it does mean the table is a direction, not a verdict.
The practical consequence: when Sonnet 5.5 comes up short on your work, raise effort before you move to Opus. Sonnet at high or maximum effort is a genuinely different data point from Sonnet at a lower setting, and it is still cheaper per token than Opus. We made the same argument about effort when we worked through what Claude Opus 5.5 actually costs to run, and it applies with more force one tier down.
Why the price gap is not 2x
Sticker rates say Sonnet 5.5 is exactly half of Opus 5.5. A realistic agent loop says otherwise, because of that identical $0.20 cache-read line.
Take a plausible coding agent: a 60,000-token cached prefix of system prompt and repository context, 20 turns, 4,000 fresh input tokens and 2,000 output tokens per turn, one cache write at the start.
Claude Sonnet 5.5
cache write 60,000 x $2.50 /1M = $0.15
cache reads 1,200,000 x $0.20 /1M = $0.24
fresh input 80,000 x $2.00 /1M = $0.16
output 40,000 x $10.00/1M = $0.40
total = $0.95
Claude Opus 5.5
cache write 60,000 x $5.00 /1M = $0.30
cache reads 1,200,000 x $0.20 /1M = $0.24
fresh input 80,000 x $4.00 /1M = $0.32
output 40,000 x $20.00/1M = $0.80
total = $1.66
That is 43% cheaper, not 50%. The cache-read line is identical on both models and accounts for a quarter of Sonnet's bill, so it dilutes every other saving. The shape of your workload decides where in that range you land: output-dominated work approaches the full 50% saving, while a loop that mostly re-reads a large cached prefix converges toward no saving at all.
The one thing to take away: cache hits cost $0.20 per million tokens on both Sonnet 5.5 and Opus 5.5. The more effectively you cache, the less switching down the tier saves you — so measure the saving on your own traces before you plan a budget around "half price."
How we would choose between them
- Default to Sonnet 5.5 for anything well specified. Bug fixes, scoped features, terminal-heavy loops, high-volume API work. It wins Terminal-Bench and costs half as much.
- Raise effort before changing model. Sonnet at high or max effort is cheaper than Opus at medium and often enough.
- Reach for Opus 5.5 when the task is vague. Long multi-file refactors, changes that must merge without human edits, anything where FrontierCode's 8.2-point gap is the relevant signal.
- Batch anything non-interactive. Half price on either model, stacks with caching.
- Only then consider Fable 5.1 at $10 and $50, and only if your evals on Opus 5.5 at higher effort still fall short.
If you are wiring either model into a product rather than a terminal, the latency work tends to matter more than the token arithmetic. Sonnet 5.5's "30%+ faster" claim is the spec most likely to change your architecture, and the lessons we wrote up while building voice agents over WebRTC transfer directly: the model is rarely the slow part.
What this changes for a working developer
- The mid tier is now a real default, not a downgrade. For a year the advice was "use the biggest model you can afford." On well-scoped coding that advice is now wrong.
- Two-model routing gets easier to justify. With the same context window, the same output ceiling and the same API shape, routing by task type is a config change rather than a rewrite.
- Pin your model IDs. Every current Claude ID is a pinned snapshot, including dateless ones like
claude-sonnet-5-5. Anthropic commits to keeping it available until at least 28 September 2027. We learned the value of pinning the hard way, which is why we wrote up the Opus 5.5 breaking changes. - Cache stability is still the biggest lever. Anything that mutates the head of your prompt invalidates the prefix and moves you from $0.20 to $2.00 per million.
What it changes for interview preparation
This is the part launch coverage skips. Agentic coding competence at $2 per million input tokens is cheap enough to leave running. When a model closes a well-scoped ticket for pennies, the market value of reciting a known algorithm from memory keeps falling.
What holds value is what the model cannot be trusted to do alone: deciding whether a change is actually correct, noticing the abstraction is wrong, and explaining a trade-off to someone who will have to maintain it. Interview formats are following. The questions worth practising look less like "implement a trie" and more like "here is a diff, tell me what breaks," "this query got slow after a million rows, walk me through it," or "defend this schema."
Our guide to preparing for interviews using AI tools goes further on drilling that muscle, and a Taqari mock interview is built around being asked to justify an answer rather than just produce one.
One caution worth stating plainly: 70.6% on Terminal-Bench is a vendor's report of a vendor's harness, run at a vendor-chosen effort level. Useful for direction, useless as a substitute for running the model on your own problems. The scepticism you would apply to a candidate's self-assessment applies here too.
Frequently asked questions
Is Claude Sonnet 5.5 better than Opus 5.5 for coding?
+
For well-scoped coding, often yes. Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0 against Opus 5.5's 66.4%, at half the token price. Opus 5.5 still leads on FrontierCode 1.1 (54.4% to 46.2%) and CursorBench 4.0 (57.8% to 55.5%), which reward long multi-file changes.
How much does Claude Sonnet 5.5 cost?
+
Two dollars per million input tokens and ten dollars per million output tokens on the Claude API. Cache reads are $0.20 per million and a cache write is $2.50. Batch requests are half price. That is exactly half Opus 5.5's $4 and $20 base rates.
Why is Claude Sonnet 5.5 only 43% cheaper than Opus 5.5, not 50%?
+
Because cache reads cost $0.20 per million on both models. Sonnet prices cache hits at 10% of its input rate, Opus at 5% of its higher one, and the two land on the same number. On a cache-heavy agent loop that identical line dilutes the 2x gap on everything else.
Does Claude Sonnet 5.5 have the same context window as Opus 5.5?
+
Yes. Both take one million input tokens and emit up to 128,000 output tokens on the synchronous Messages API, with a June 2026 knowledge cutoff. Context is not the reason to pick between them; cost, latency and sustained judgement are.
What is the catch in the Sonnet 5.5 benchmark table?
+
Effort levels differ. Sonnet 5.5 defaults to high effort on the Claude API while Opus 5.5 defaults to medium, and some published cells are max-effort runs. A comparison at each model's default is not a comparison at equal compute, so run your own evals before concluding.
What does a cheaper mid-tier coding model change for interview prep?
+
It shifts what gets tested. When a $2-per-million model can close well-scoped tickets, writing a known algorithm from memory proves less than reading a diff, spotting a wrong abstraction, or defending a trade-off out loud. Expect more debugging and design questions.