Claude Sonnet 5.5 vs Sonnet 5: Same Price, 80% Less per Task
Contents
- What changed in Claude Sonnet 5.5?
- What do the release benchmarks say?
- How did we measure it?
- Is Sonnet 5.5 cheaper than Sonnet 5 per task?
- What does each effort level cost?
- Why does Sonnet 5.5 put its working in the answer?
- What happens in a tool loop?
- Is Sonnet 5.5 faster?
- Do Sonnet 5 token budgets carry over?
- Which should you use?
- FAQ
Claude Sonnet 5.5 has the same token prices as Claude Sonnet 5, $2 per million input tokens and $10 per million output, so every saving comes from using fewer tokens. On 13 single-shot tasks at the API defaults it cost $0.0041 per task against Sonnet 5’s $0.021, 80% less, and both models got every task right. The problem we found is output format: below the xhigh effort level, Sonnet 5.5 sometimes skips its hidden reasoning and writes the working into the reply, even when the prompt asks for the answer and nothing else.
TL;DR
- At the API default, Sonnet 5.5 cost $0.0041 per task against Sonnet 5’s $0.021, both correct on all 39 calls.
- Sonnet 5.5 cost about the same at
low,mediumandhigheffort;maxcost 3.5x the default. - Below
xhigh, Sonnet 5.5 put its working in the reply on 68 of 156 answer-only prompts; a system prompt did not fix it. - In a four-question tool loop, Sonnet 5.5 at
maxcost $0.042 per run, more than Opus 5.5 at its default ($0.033).
Anthropic released Sonnet 5.5 on 2026-09-28, claiming it is 30% faster and “up to 30% less per task” than Sonnet 5. We measured it the next day.
What changed in Claude Sonnet 5.5?
The price stayed the same; the thinking controls and several request parameters changed. Sonnet 5.5 uses adaptive thinking: the model decides how much hidden reasoning to do before answering, and those reasoning tokens are billed as output. You steer it with effort (output_config.effort on the Messages API), a request parameter with five levels from low to max. The API default is high.
| Sonnet 5.5 | Sonnet 5 | Opus 5.5 | |
|---|---|---|---|
| Released | 2026-09-28 | 2026-06-30 | 2026-09-22 |
| Input / output, per 1M tokens | $2 / $10 | $2 / $10 | $4 / $20 |
| Cache read, per 1M tokens | $0.20 | $0.20 | $0.20 |
| Context window / max output | 1M / 128K | 1M / 128K | 1M / 128K |
| Knowledge cutoff (month only published) | June 2026 | January 2026 | June 2026 |
| Default effort on the API | high | high | medium |
| Lowest thinking setting | between_tools (at high or below) | disabled | no off switch; thinking is always on |
| Minimum cacheable prompt | 512 tokens | 1,024 tokens | 512 tokens |
Sonnet 5’s $2 / $10 launched as an introductory price, but Anthropic’s pricing page now lists it as standard; the planned rise to $3 / $15 did not happen.
According to the migration guide, these requests worked on Sonnet 5 and return HTTP 400 on Sonnet 5.5:
thinking: {"type": "disabled"}. Usethinking: {"type": "between_tools"}, which turns off reasoning before the first answer and is accepted only athigheffort or below.- Forced tool use (
tool_choiceset toanyor a named tool). Keepautoand say in the prompt when the tool applies. - A manual thinking budget (
budget_tokens), non-defaulttemperature,top_portop_k, and a prefilled assistant turn (starting the model’s reply for it). - Replaying a Sonnet 5.5 thinking block after changing the system prompt, the tools or an earlier message, on accounts created from 2026-08-31.
- The older
computer_20251124computer use tool on the Claude API and Google Cloud.
One change raises no error: short notes the model writes between tool calls now arrive as thinking blocks, empty at the default display setting, so an interface that streams them goes quiet.
What do the release benchmarks say?
Anthropic’s own table puts Sonnet 5.5 within a few points of Opus 5.5 at half its token price, and far ahead of Sonnet 5.
| Benchmark (what it tests) | Sonnet 5.5 | Sonnet 5 | Opus 5.5 | GPT-6 Sol |
|---|---|---|---|---|
| Terminal-Bench 4.0 (agentic coding in a terminal) | 70.6% | 10.3% | 66.4% (xhigh) | not reported |
| CursorBench 4.0 (coding in an editor) | 55.5% | 34.1% | 57.8% | not reported |
| GDPval-AA v2.1 (knowledge-work documents, Elo rating, higher is better) | 1844 | 1449 | 1846 | 1487 |
| OSWorld 2.1 (computer use, partial credit) | 80.1% | 57.0% | 81.8% | not reported |
| Humanity’s Last Exam, with tools | 64.5% | 54.9% | 67.7% | not reported |
Artificial Analysis adds a cost warning: at max effort Sonnet 5.5 scored 56 on its Intelligence Index, two points behind Opus 5.5 at max, but used about 193K output tokens per task, the most it has measured and about 60% more than Opus 5.5 or Sonnet 5 at max, for about $7.60 per index task.
How did we measure it?
We sent 13 single-shot tasks (one prompt, one reply, no tools) with known answers to all three models: eight short ones (sum the primes below 60, count the digit 7 from 1 to 500) and five that need real multi-step work (a 10-item knapsack, paths through a blocked 8x8 grid, 13 to the power 1,001 modulo 10,007). Every answer key was brute-forced locally and every prompt ended with “Reply with a single integer, nothing else.” Each task ran 3 times at the API default and at each of the five effort levels, 702 calls in total, on the native Messages API. Each prompt carried a unique random string so no response came from a cache, and cost was computed from Anthropic list prices. We graded whether the final answer was right and whether the reply was the answer only.
Accuracy did not separate the models. Sonnet 5.5 was correct on all 234 calls and Sonnet 5 on all 233 that returned (one got a server error); Opus 5.5 missed one task at low and one at the default.
Is Sonnet 5.5 cheaper than Sonnet 5 per task?
Sonnet 5.5 cost 80% less than Sonnet 5 at the default and 77% to 78% less at low, medium and high, because it wrote about a fifth as many output tokens. The list prices are identical, so all of the saving is token efficiency. Sonnet 5 wrote about 1,800 to 2,100 tokens per task at every effort level; Sonnet 5.5 wrote about 400 until xhigh.
| All 13 tasks, per task | Sonnet 5.5 | Sonnet 5 | Opus 5.5 |
|---|---|---|---|
| API default | $0.0041 (396 tokens) | $0.0212 (2,105) | $0.0070 (334) |
low | $0.0043 (409) | $0.0184 (1,817) | $0.0065 (309) |
medium | $0.0040 (383) | $0.0180 (1,780) | $0.0082 (393) |
high | $0.0043 (416) | $0.0188 (1,858) | $0.0088 (421) |
xhigh | $0.0059 (574) | $0.0202 (2,001) | $0.0105 (507) |
max | $0.0146 (1,438) | $0.0193 (1,909) | $0.0234 (1,149) |
Token counts are mean output tokens per call, reasoning included. On the five hard tasks the saving at the default is 84% ($0.0057 against $0.0348). In our tool loop, where the whole transcript is re-sent every turn and input tokens dominate, it is 8%, so Anthropic’s “up to 30%” is conservative for single prompts like these and generous for a short loop like ours.
With only 13 tasks, the 80% saving at the default has a wide range: resampling the tasks gives a 95% interval of 66% to 85% less, and the lower end stays above 50% at every setting from low to xhigh.
What does each effort level cost?
On Sonnet 5.5, low, medium and high all cost about $0.0042 per task, so the high default costs nothing extra here; xhigh cost about 40% more and max 3.5x the default. At max, Sonnet 5.5 cost twice as much per task as Opus 5.5 at its default ($0.0146 against $0.0070), with no gain in accuracy.
The flat region will not hold everywhere: on a longer DevOps task, another tester saw high use roughly twice the output tokens of medium. A workload where effort matters needs its own sweep.
Why does Sonnet 5.5 put its working in the answer?
Below xhigh, Sonnet 5.5 does not always use a thinking block, and when it skips one it reasons in the reply. A thinking block is the separate part of the response that holds the model’s reasoning; its text is empty by default, and the answer follows in a text block.
Of the 234 Sonnet 5.5 replies, all 166 with a thinking block were the bare answer. All 68 without one wrote the working first, for example “09:47 + 3:46 = 13:33 … + 1:39 = 15:40” followed by “15:40”. The final answer was right every time, but the reply was not what the prompt asked for.
| Answer-only replies | Sonnet 5.5 | Sonnet 5 | Opus 5.5 |
|---|---|---|---|
| API default | 22 / 39 | 38 / 39 | 38 / 39 |
low | 12 / 39 | 37 / 39 | 38 / 39 |
medium | 24 / 39 | 37 / 39 | 39 / 39 |
high | 30 / 39 | 38 / 39 | 39 / 39 |
xhigh | 39 / 39 | 35 / 38 | 39 / 39 |
max | 39 / 39 | 37 / 39 | 39 / 39 |
The behavior is per task: where it happened, Sonnet 5.5 showed its working on all three repeats, except one task at the default where it did so twice. That happened on 9 of the 13 tasks at low, 6 at the default, 5 at medium and 3 at high (all short arithmetic tasks). Counted by task, with only 13 tasks, none of these gaps clears a Holm correction, so read the counts as what we observed, not as a rate to plan around. Sonnet 5’s misses look different: a correct answer in bold followed by a short explanation. Its xhigh row has 38 calls because one got a server error.
A system prompt asking for “the final result only”, with any working done silently, did not help: Sonnet 5.5 gave the bare answer on 8 of 24 short tasks at low and 9 of 24 at medium, against 6 and 9 without it. xhigh fixed it, with a thinking block and a bare answer every time, at about 40% more cost than low to high. For code that parses model output:
- Run at
xhighwhen one line must be the answer. - Or read the answer from the last non-empty line, which was right in all 68 cases here.
- Anthropic’s structured outputs can also constrain the reply to a schema; we did not test this path.
import anthropic
client = anthropic.Anthropic()
prompt = "Compute 7 raised to the power 222, modulo 1000. Reply with a single integer, nothing else."
resp = client.messages.create(
model="claude-sonnet-5-5",
max_tokens=16000,
output_config={"effort": "xhigh"}, # low, medium, high (API default), xhigh, max
messages=[{"role": "user", "content": prompt}],
)
text = "".join(block.text for block in resp.content if block.type == "text")
lines = text.strip().splitlines()
answer = lines[-1] if lines else None # below xhigh, working can precede the answer
print(answer, resp.usage.output_tokens) # output_tokens includes the reasoning
What happens in a tool loop?
At the default, low and medium, Sonnet 5.5 solved every run of a four-question code-reading loop for about $0.011, a third of Opus 5.5’s cost. At max it made three times as many tool calls and cost more than Opus 5.5. Each model gets three tools (list files, read a file, search) over a small synthetic codebase, and each answer needs 3 to 6 lookups across files.
| Tool loop, 12 runs each | Solved | Median turns | Tool calls per run | Cost per run | Median wall time |
|---|---|---|---|---|---|
| Sonnet 5.5, default | 12 / 12 | 3 | 4.8 | $0.0112 | 6.7 s |
Sonnet 5.5, low | 12 / 12 | 3 | 4.9 | $0.0112 | 7.4 s |
Sonnet 5.5, medium | 12 / 12 | 3 | 5.1 | $0.0113 | 6.8 s |
Sonnet 5.5, max | 12 / 12 | 4 | 14.7 | $0.0422 | 20.1 s |
| Opus 5.5, default | 12 / 12 | 4 | 5.3 | $0.0332 | 25.3 s |
| Sonnet 5, default | 11 / 12 | 3.5 | 4.2 | $0.0122 | 15.8 s |
Customer quotes reported by VentureBeat describe fewer tool calls than Sonnet 5 (one third fewer at Lovable). This loop is too short to show that: Sonnet 5.5 made 4.8 calls per run against Sonnet 5’s 4.2, while costing 8% less and finishing more than twice as fast.
Is Sonnet 5.5 faster?
It finished sooner mainly because it wrote less. The median wall time (from sending the request to receiving the full response) per single-shot task at the default was 3.9 seconds for Sonnet 5.5, 8.1 seconds for Sonnet 5 and 5.5 seconds for Opus 5.5. Output tokens per second of wall time were about the same for the two Sonnets (88 against 91), so the wait was half as long because Sonnet 5.5 wrote a fifth of the tokens, not because it generated them faster. On the hard tasks the gap is 5.8 seconds against 21.0.
Do Sonnet 5 token budgets carry over?
Context budgets and max_tokens limits written for Sonnet 5 should still fit. Anthropic’s migration notes say Sonnet 5.5 uses the same tokenizer, and four fixed texts we tried (English prose, Python code, JSON tool arguments, Chinese prose) came to identical counts on both models and on Opus 5.5, for example 1,270 tokens for the English and 493 for the Chinese; Sonnet 5.5 and Opus 5.5 add a fixed 2 tokens per request. Requests with tools get slightly cheaper: Anthropic’s pricing page lists the hidden tool-use system prompt at 286 tokens on Sonnet 5.5 against 354 on Sonnet 5.
Which should you use?
For most Sonnet 5 workloads, switch; what remains is picking the effort level.
| Workload | Watch for | Recommendation | Numbers |
|---|---|---|---|
| Output parsed by code: extraction, classification, single values | Working written into the reply below xhigh | Sonnet 5.5 at xhigh, or medium and parse the last line | answer-only 39 / 39 at xhigh, 24 / 39 at medium; xhigh about 40% more |
| Chat and user-facing text | Latency and cost | Sonnet 5.5 at medium | $0.0040 per task, 3.9 s median |
| Agent loops with tools | max multiplying tool calls | Sonnet 5.5 at the default or medium; move to Opus 5.5 before Sonnet 5.5 at max | $0.011 per run at the default, $0.042 at max, $0.033 on Opus 5.5 |
| Sonnet 5 code that turns thinking off or forces a tool | HTTP 400 after the model swap | between_tools (at high or below) and tool_choice: auto | Anthropic’s migration guide |
Whatever the effort level, keep two checks in your code: that the reply has the shape your parser expects, and a per-request output token cap, since max raised cost per task 3.5x on single prompts and tool calls 3x in the loop.
FAQ
Is Sonnet 5.5 cheaper than Opus 5.5?
At their defaults, Sonnet 5.5 cost 41% less than Opus 5.5 per single-shot task and a third as much per tool-loop run. At max the order flips: Sonnet 5.5 cost twice Opus 5.5’s default on single-shot tasks and 27% more per tool-loop run.
Which effort level should I use on Sonnet 5.5?
Sonnet 5.5 cost about the same at low, medium and high on our tasks ($0.0040 to $0.0043), so medium is a safe start for chat and tool loops. Use xhigh when code parses the reply as a bare value; max cost 3.5x the default.
Can I still turn thinking off on Sonnet 5.5?
Sonnet 5.5 rejects thinking: {"type": "disabled"} with HTTP 400. Send thinking: {"type": "between_tools"} instead, accepted at low, medium or high effort.
Related measurements: Claude Opus 5.5 vs Opus 5, the Claude Sonnet 5 tokenizer, and thinking controls across vendors.
Measured 2026-09-29, one day after release, through a gateway to the Anthropic Messages API. 702 graded single-shot calls (13 tasks with brute-forced answer keys, 3 repeats, 6 effort settings, 3 models), 48 calls testing a system instruction on output format, 72 tool-loop runs (4 multi-hop questions, 3 repeats, 6 settings), and token counts for four fixed texts per model. Prompts salted; costs computed from Anthropic list prices.