Neu Kostenlos registrieren, 10 Aufrufe gratis. Bis zu 1 $, ohne Karte.

Claude Opus 5.5 vs Claude Sonnet 5.5

vs

Welches Modell, wann

Dies sind zwei Stufen derselben Claude-Linie, keine Zwillinge: Anthropic positioniert claude-opus-5-5 als die höhere Stufe, entwickelt für lang laufendes, agentenbasiertes Programmieren und Wissensarbeit, und claude-sonnet-5-5 als die niedrigere Stufe, die als schneller aufgeführt wird. Auf dem Papier teilen sie sich einen 1000000-Token-Kontext, 128000 maximalen Output, Text- und Bildeingabe sowie $0.2 für Cache-Lesezugriffe, sodass die Preisliste den Stufenunterschied widerspiegelt: claude-opus-5-5 liegt bei $4 für Input und $20 für Output, 2x claude-sonnet-5-5 bei $2 und $10. Leiten Sie hochvolumigen und latenzempfindlichen Traffic über claude-sonnet-5-5 und senden Sie die langen, komplexen agentenbasierten Aufgaben an claude-opus-5-5.

Benchmarks

Claude Sonnet 5.5: Der Anbieter hat keine Benchmark-Werte veröffentlicht.

Über dem DurchschnittKeins besserClaude Opus 5.59 / 97 / 9
Claude Opus 5.5 Claude Sonnet 5.5 weitere gemessene Modelle Durchschnitt der Vergleichsmodelle ★ kein anderes Modell war besser
Terminal-bench 4.0
kein anderes Modell war besser 66.4%
N/A
OSWorld 2.0 partial
kein anderes Modell war besser 81.8%
N/A
Terminal-Bench-Science 0.1
58.7%
N/A
Humanity's Last Exam with tools
kein anderes Modell war besser 67.7%
N/A
AutomationBench
40%
N/A
Chartography with tools
kein anderes Modell war besser 89%
N/A

Herstellerangaben: Alibaba (Qwen) Anthropic DeepSeek Google Moonshot OpenAI Z.ai

Preise

Claude Opus 5.5 Claude Sonnet 5.5 Δ
Input / 1M Token $4 $2 2×
Output / 1M Token $20 $10 2×
Cache-Read / 1M Token $0.2 $0.2 =
Cache-Schreiben 1.25x (5m) / 2x (1h) 1.25x (5m) / 2x (1h) -

Preise aus dem Live-Katalog zum Zeitpunkt des Builds; jede Modellseite enthält die aktuelle Übersicht.

Wo sie stehen - Eingabepreis pro 1M Tokens über alle 76 Chat-Modelle mit dieser Abrechnungseinheit (logarithmische Skala)

Fähigkeiten

Claude Opus 5.5 Claude Sonnet 5.5
Tool-Nutzung ja ja
Thinking-Kontrolle immer aktiv immer aktiv
Strukturierte Ausgabe ja ja
Prompt-Caching explizit (Sie markieren das Präfix) explizit (Sie markieren das Präfix)
Cache-Lebensdauer 5m default, 1h option 5m default, 1h option
Minimales gecachtes Präfix 1024 Tokens 1024 Tokens

Spezifikationen

Claude Opus 5.5 Claude Sonnet 5.5
Input-Modalitäten Text Bild Text Bild
Ausgabemodalitäten Text Text
Veröffentlicht 2026-09-22 2026-09-28
Wissensgrenze 2026-06 2026-06
Kontextfenster 1M 1M
Maximaler Output 128K 128K
Thinking-Parameter output_config.effort thinking.type
Zulässige Werte
output_config.effort
  • low
  • medium
  • high
  • xhigh
  • max
thinking.type
  • adaptive (default)
  • between_tools
Standardwert medium adaptive, effort high

Die Spezifikationen sind aus der Dokumentation der jeweiligen Anbieter übernommen; eine Zeile, die ein Anbieter nicht veröffentlicht, wird weggelassen und nicht abgeleitet. Vollständige Quellen: Claude Opus 5.5 · Claude Sonnet 5.5

Ein Prompt, beide Modelle - gemessen über das Gateway

PROMPT Explain to a CFO, in exactly three sentences, why our LLM bill doubled after we enabled reasoning. PRÜFEN genau 3 Sätze

Claude Opus 5.5 bestanden · 3 sentences

When we turned on reasoning, the model began "thinking out loud" before each answer, generating thousands of internal working tokens that users never see but that we are billed for just like the final response. Those hidden tokens are charged at the output rate, typically the most expensive rate, so the same question with the same visible answer now costs roughly twice as much or more. We can bring the bill back down without losing the quality gains by capping the reasoning budget and reserving reasoning mode for the complex tasks that actually benefit from it.

Ausgabe 263 tok (+102 Denken) Latenz 5.7 s

Claude Sonnet 5.5 bestanden · 3 sentences

Enabling reasoning makes the model "think" before it answers, and that internal thinking is generated as extra text (reasoning tokens) that we are billed for at the same rate as output tokens, even though users never see it. Output tokens also cost several times more than input tokens, so a single request that used to produce a short answer can now generate many times more billable text behind the scenes. Because our request volume and prompts didn't change, the doubling comes almost entirely from this hidden thinking, and we can bring it down by limiting reasoning effort, turning it off for simple tasks, and reserving it for the workloads where the accuracy gain justifies the cost.

Ausgabe 205 tok Latenz 3.6 s

Befolgen der Anweisung (genau drei Sätze - abzählbar), Zuschnitt auf das Publikum (CFO-Register), und die Abrechnungslücke durch verstecktes Denken, die die Token-Zähler unten offenlegen.

PROMPT This function is supposed to return the longest run of consecutive equal items, but callers report it is off by one on single-element inputs and misses runs that end at the last index. Fix it and explain the bug in one paragraph. PRÜFEN Fix besteht die Tests

Claude Opus 5.5 bestanden · 8/8 cases

```python def longest_run(xs): if not xs: return 0 best = 1 cur = 1 for i in range(1, len(xs)): if xs[i] == xs[i - 1]: cur += 1 best = max(best, cur) else: cur = 1 return best ``` **Explanation:** The original code counted *matching adjacent pairs* rather than *items in a run*. Both `cur` and `best` started at 0, and `cur` reset to 0 on a mismatch, so a run of length *n* was reported as *n − 1*. For a single-element list the loop never executes, so it returned 0 instead of 1. For `[5, 5]` it returned 1 instead of 2. The fix treats every element as the start of a run of length 1: `best` and `cur` start at 1, and `cur` resets to 1 (not 0) when the value changes. An explicit check returns 0 for an empty list. The "misses runs at the last index" reports are the same off-by-one bug rather than a separate one. […]

Ausgabe 723 tok (+303 Denken) Latenz 10.4 s

Claude Sonnet 5.5 bestanden · 8/8 cases

**The bug:** `cur` counts the number of *adjacent equal pairs* rather than the number of items in the run. A run of length k contains only k−1 such pairs, so every result is one too small. The counter also starts at 0 and the loop starts at index 1, so a single-element list never updates `best` and returns 0 instead of 1. (An empty list should return 0.) In addition, `cur` resets to 0 on a mismatch, when it should reset to 1, because the item at `xs[i]` begins a new run of length 1. The claim that it "misses runs ending at the last index" is really the same off-by-one: the final run is counted, but as one less than its true length. Because `best` is updated inside the loop on every extension, the last run isn't skipped. […]

Ausgabe 444 tok Latenz 6.7 s

Ob der Fix tatsächlich korrekt ist (ausführbar), die Dichte der Erklärung, und die Token-Effizienz bei einer klar begrenzten Aufgabe.

PROMPT Extract a JSON object with fields {company, quarter, revenue_usd_m, growth_pct, guidance} from this text. Use null for anything not stated; add no other fields. PRÜFEN gültiges JSON, Schema exakt

Claude Opus 5.5 bestanden · 5/5 fields, guidance "Withheld pending Q4 close"

```json { "company": "Meridian Robotics", "quarter": "September quarter", "revenue_usd_m": 412, "growth_pct": 11, "guidance": "Withheld pending Q4 close" } ``` **Judgment calls to check:** - **quarter**: Kept as "September quarter" rather than "Q3", since the fiscal calendar isn't stated. The text mentions an upcoming "Q4 close," so it may be fiscal Q3, but that's not certain. - **revenue_usd_m**: The text says only "412 for the topline." It doesn't state currency or units, and the figure is on a constant-currency basis, not reported USD. USD millions is an assumption. Use `null` if you need strictly stated values. - **growth_pct**: The source says "eleven and change," so the true figure is slightly above 11. 11 is an approximation. - **guidance**: Guidance was explicitly withheld. I recorded that rather than `null`, since the text addresses it directly.

Ausgabe 829 tok (+488 Denken) Latenz 10.9 s

Claude Sonnet 5.5 bestanden · 5/5 fields, guidance null

Here is the extracted JSON: ```json { "company": "Meridian Robotics", "quarter": "September quarter", "revenue_usd_m": 412, "growth_pct": 11, "guidance": null } ``` Notes on the choices: - **quarter**: The text says "September quarter" without specifying a fiscal year or Q-number, so I kept the wording as stated. - **revenue_usd_m**: The text gives "412" for the topline, with no currency or unit stated explicitly. I've taken it as 412 million USD per the field definition, but the text does not confirm that, and the figure is constant-currency. - **growth_pct**: "Eleven and change" means slightly above 11, so 11 is an approximation. The exact figure isn't given. - **guidance**: Guidance is withheld pending the Q4 close, so no guidance is stated, hence null.

Ausgabe 300 tok Latenz 3.5 s

Schema-Treue (keine erfundenen Felder), Halluzinationsdruck (guidance wird ausdrücklich zurückgehalten), und Unterschiede im Pfad für strukturierte Ausgabe.

PROMPT Write a 120-word product announcement for an API gateway feature that caches prompts across providers. Forbidden words: "seamless", "unlock", "game-changer", "revolutionize", "empower". Exactly one sentence must be a question. PRÜFEN 120 Wörter, 0 verbotene Wörter

Claude Opus 5.5 bestanden · 120 words, 0 banned, 1 question

Today we're launching Cross-Provider Prompt Caching for our API gateway. Why pay twice for the same context? Our gateway now stores repeated prompt prefixes, system instructions, and document chunks, then reuses them whether your request goes to OpenAI, Anthropic, Google, or an open-source model. Teams running retrieval pipelines or long agent loops can cut token costs by up to 60 percent and trim latency on every cached call. Setup takes one configuration flag; no code changes are required. Cache rules respect tenant boundaries, expire on your schedule, and log every hit for audit review. Dashboards show savings by provider, route, and application in real time. Cross-Provider Prompt Caching is available on all Pro and Enterprise plans. Turn it on today.

Ausgabe 1263 tok (+1007 Denken) Latenz 13.5 s

Claude Sonnet 5.5 bestanden · 122 words, 0 banned, 1 question

**Introducing Cross-Provider Prompt Caching** Your API gateway now caches prompts across every connected LLM provider, whether you route traffic to OpenAI, Anthropic, Google, or a self-hosted model. Here's how it works: the gateway fingerprints each prompt prefix, stores it centrally, and applies the right provider-specific caching mechanism automatically. When you switch providers or fail over during an outage, your cached context follows you. Early testing shows input token costs dropping by up to 70% on repeated system prompts and long documents, with latency improvements of 40% or more on cache hits. Tired of paying to send the same instructions thousands of times a day? Enable it with one configuration flag. Cache rules, TTLs, and per-team isolation are available in the dashboard today.

Ausgabe 266 tok Latenz 3.8 s

Einhaltung der Vorgaben (Wortbudget, Liste verbotener Wörter, die eine Frage), Stil-Fingerabdruck, und Längensteuerung.

Mit einer Zeile zwischen ihnen wechseln

Beide IDs befinden sich in jedem Tab unten - das hervorgehobene Zeilenpaar ist die einzige Änderung. Gleicher Endpunkt, gleicher Schlüssel, gleiche Request-Struktur.

from openai import OpenAI

client = OpenAI(
    base_url="https://synthorai.io/v1",
    api_key="sk-syn-...",
)

resp = client.chat.completions.create(
    model="claude-opus-5-5",
    # model="claude-sonnet-5-5",  # diese Zeile einkommentieren, die darüberliegende auskommentieren
    messages=[{"role": "user", "content": "Summarize this diff"}],
    reasoning_effort="medium",
)
print(resp.choices[0].message.content)

API-Schlüssel abrufen →

FAQ

Welches ist günstiger, Claude Opus 5.5 oder Claude Sonnet 5.5?

Claude Sonnet 5.5 ist günstiger bei input / 1m token ($2 vs. $4, 2.0× Unterschied). Andere Zeilen können in die andere Richtung deuten - die obige Tabelle enthält alle Daten, und die tatsächlichen Kosten hängen von Ihrem Mix ab.

Kann ich Claude Opus 5.5 gegen Claude Sonnet 5.5 ohne zwei Integrationen A/B-testen?

Ja. Beide werden über denselben OpenAI-kompatiblen Endpunkt mit einem API-Schlüssel bereitgestellt - der Wechsel ist eine einzeilige Änderung des Modell-Strings, sodass Sie einen Bruchteil des Traffics an jedes Modell leiten und die Rechnungen direkt vergleichen können.

Unterstützen Claude Opus 5.5 und Claude Sonnet 5.5 Prompt-Caching?

Ja - beide berechnen Cache-Reads günstiger als ihre Eingaberate, sodass Warm-Prefix-Workloads weniger kosten, als die Listenpreise vermuten lassen. Die genauen Zeilen für Cache-Reads befinden sich in der obigen Preistabelle.

Verwandte Vergleiche

Aus unseren gemessenen Studien