Neu Kostenlos registrieren, 10 Aufrufe gratis. Bis zu 1 $, ohne Karte.

Claude Sonnet 5.5

Veröffentlicht 2026-09-28

chatCodeReasoningTool-AufrufeVisionPrompt Caching

Claude Sonnet 5.5 ist Anthropics Modell für die beste Kombination aus Geschwindigkeit und Intelligenz, veröffentlicht am 28.

Eingabe
Text Bild $2/M
Ausgabe
Text $10/M
Cache-Read
$0.2/M
Kontext
1M
vs. GPT-4o
~60% günstiger
Knowledge Cutoff
2026-06

Preis im Vergleich

Preisposition unter 68 vergleichbaren Modellen

Eingabe$2/M
$0.05 · Qwen3 VL Flash GPT-5.4 Pro · $30
Ausgabe$10/M
$0.275 · DeepSeek V4 Flash GPT-5.4 Pro · $180
Cache-Lesen$0.2/M
$0.0028 · DeepSeek V4 Flash GPT-5.4 Pro · $15

Der Balken zeigt, wo der Preis dieses Modells unter allen Modellen derselben Art auf Synthorai liegt. An beiden Enden stehen das günstigste und das teuerste Modell. Es sind Grundpreise; Rabatte für Batch, Region und Cache-Schreibvorgänge stehen auf der Preisseite.

Spezifikationen & Limits

Tokens

Kontextfenster (Herstellerangabe) 1.000.000
Max. Output (Anbieter-Spezifikation) 128.000
Knowledge Cutoff 2026-06

Prompt Caching

Modus explizit (Opt-in)
Min. Präfix 512 Anbieter-Standard: 1.024
Lebensdauer Standard 5 Min., optional 1 Std.
Schreibkosten 1.25x (5m) / 2x (1h)

Thinking

Anbieter-Parameter thinking.type
Zulässige Werte adaptive (default) · between_tools
Standardwert adaptive, effort high gilt, wenn die Anfrage nichts angibt
Abschaltbar Nein
Thinking-Verhalten Adaptive thinking is on by default. The lowest setting, between_tools, turns off up-front thinking and works at high effort or below; thinking {"type": "disabled"} and a manual {"type": "enabled", "budget_tokens": N} both return a 400 error.
Parameter reasoning_effort
Werte minimal · low · medium · high die Parameteroberfläche des Gateways - das Anbieter-Mapping oben gilt

Modell

Modalitäten Text + Bild → Text
  • Same price as Claude Sonnet 5
  • 1M context at standard pricing with no long-context tier
  • prompt-cache reads cost 0.1x input ($0.20/M)
  • minimum cacheable prompt 512 tokens
  • setting temperature, top_p or top_k to a non-default value returns a 400
  • forced tool use returns an error
  • text between tool calls comes back in thinking blocks
  • up to 300k output tokens on the Message Batches API with the output-300k-2026-03-24 beta header

laut Offizielle Anthropic-Dokumentation ↗

Ein Prompt, gemessen über das Gateway

PROMPT Explain to a CFO, in exactly three sentences, why our LLM bill doubled after we enabled reasoning. PRÜFEN genau 3 Sätze

Claude Sonnet 5.5 bestanden · 3 sentences

Enabling reasoning makes the model "think" before it answers, and that internal thinking is generated as extra text (reasoning tokens) that we are billed for at the same rate as output tokens, even though users never see it. Output tokens also cost several times more than input tokens, so a single request that used to produce a short answer can now generate many times more billable text behind the scenes. Because our request volume and prompts didn't change, the doubling comes almost entirely from this hidden thinking, and we can bring it down by limiting reasoning effort, turning it off for simple tasks, and reserving it for the workloads where the accuracy gain justifies the cost.

Ausgabe 205 tok Latenz 3.6 s

Befolgen der Anweisung (genau drei Sätze - abzählbar), Zuschnitt auf das Publikum (CFO-Register), und die Abrechnungslücke durch verstecktes Denken, die die Token-Zähler unten offenlegen.

PROMPT This function is supposed to return the longest run of consecutive equal items, but callers report it is off by one on single-element inputs and misses runs that end at the last index. Fix it and explain the bug in one paragraph. PRÜFEN Fix besteht die Tests

Claude Sonnet 5.5 bestanden · 8/8 cases

**The bug:** `cur` counts the number of *adjacent equal pairs* rather than the number of items in the run. A run of length k contains only k−1 such pairs, so every result is one too small. The counter also starts at 0 and the loop starts at index 1, so a single-element list never updates `best` and returns 0 instead of 1. (An empty list should return 0.) In addition, `cur` resets to 0 on a mismatch, when it should reset to 1, because the item at `xs[i]` begins a new run of length 1. The claim that it "misses runs ending at the last index" is really the same off-by-one: the final run is counted, but as one less than its true length. Because `best` is updated inside the loop on every extension, the last run isn't skipped. […]

Ausgabe 444 tok Latenz 6.7 s

Ob der Fix tatsächlich korrekt ist (ausführbar), die Dichte der Erklärung, und die Token-Effizienz bei einer klar begrenzten Aufgabe.

PROMPT Extract a JSON object with fields {company, quarter, revenue_usd_m, growth_pct, guidance} from this text. Use null for anything not stated; add no other fields. PRÜFEN gültiges JSON, Schema exakt

Claude Sonnet 5.5 bestanden · 5/5 fields, guidance null

Here is the extracted JSON: ```json { "company": "Meridian Robotics", "quarter": "September quarter", "revenue_usd_m": 412, "growth_pct": 11, "guidance": null } ``` Notes on the choices: - **quarter**: The text says "September quarter" without specifying a fiscal year or Q-number, so I kept the wording as stated. - **revenue_usd_m**: The text gives "412" for the topline, with no currency or unit stated explicitly. I've taken it as 412 million USD per the field definition, but the text does not confirm that, and the figure is constant-currency. - **growth_pct**: "Eleven and change" means slightly above 11, so 11 is an approximation. The exact figure isn't given. - **guidance**: Guidance is withheld pending the Q4 close, so no guidance is stated, hence null.

Ausgabe 300 tok Latenz 3.5 s

Schema-Treue (keine erfundenen Felder), Halluzinationsdruck (guidance wird ausdrücklich zurückgehalten), und Unterschiede im Pfad für strukturierte Ausgabe.

PROMPT Write a 120-word product announcement for an API gateway feature that caches prompts across providers. Forbidden words: "seamless", "unlock", "game-changer", "revolutionize", "empower". Exactly one sentence must be a question. PRÜFEN 120 Wörter, 0 verbotene Wörter

Claude Sonnet 5.5 bestanden · 122 words, 0 banned, 1 question

**Introducing Cross-Provider Prompt Caching** Your API gateway now caches prompts across every connected LLM provider, whether you route traffic to OpenAI, Anthropic, Google, or a self-hosted model. Here's how it works: the gateway fingerprints each prompt prefix, stores it centrally, and applies the right provider-specific caching mechanism automatically. When you switch providers or fail over during an outage, your cached context follows you. Early testing shows input token costs dropping by up to 70% on repeated system prompts and long documents, with latency improvements of 40% or more on cache hits. Tired of paying to send the same instructions thousands of times a day? Enable it with one configuration flag. Cache rules, TTLs, and per-team isolation are available in the dashboard today.

Ausgabe 266 tok Latenz 3.8 s

Einhaltung der Vorgaben (Wortbudget, Liste verbotener Wörter, die eine Frage), Stil-Fingerabdruck, und Längensteuerung.

Claude Sonnet 5.5 in 30 Sekunden nutzen

OpenAI-kompatibel: Tauschen Sie die base_url, behalten Sie Ihr SDK. POST /v1/chat/completions

from openai import OpenAI

client = OpenAI(
    base_url="https://synthorai.io/v1",
    api_key="sk-syn-...",
)

resp = client.chat.completions.create(
    model="claude-sonnet-5-5",
    messages=[{"role": "user", "content": "Summarize this diff"}],
    reasoning_effort="medium",
)
print(resp.choices[0].message.content)

Über Claude Sonnet 5.5

  • September 2026, und kostet genauso viel wie Claude Sonnet 5: 2 USD je Million Eingabe-Token und 10 USD je Million Ausgabe-Token, Cache-Lesezugriffe zu 10 % des Eingabepreises (0,20 USD je Million), 5-Minuten-Cache-Schreibvorgänge zu 2,50 USD und 1-Stunden-Schreibvorgänge zu 4 USD.
  • Es bietet ein Kontextfenster von 1M Token ohne Aufpreis für lange Kontexte, 128K maximale Ausgabe und nimmt Text- und Bildeingaben entgegen.
  • Adaptives Denken ist standardmäßig aktiv, die Standard-Intensität ist high; die niedrigste Stufe between_tools schaltet das vorgelagerte Denken ab und funktioniert bis einschließlich high, während ein deaktivierter thinking-Block oder eine manuelle budget_tokens-Anfrage einen 400-Fehler liefert.
  • Auch temperature, top_p oder top_k mit einem Nicht-Standardwert liefern einen 400-Fehler, und erzwungene Werkzeugnutzung wird nicht unterstützt.
  • Anthropic nennt fünf Breaking Changes für Code, der bereits auf Claude Sonnet 5 läuft, und Text zwischen Werkzeugaufrufen kommt jetzt in thinking-Blöcken zurück.
  • Der kleinste cachebare Prompt umfasst 512 Token.
  • Synthorai stellt Claude Sonnet 5.5 über dieselbe OpenAI-kompatible API bereit wie die übrige Flotte.

FAQ

Lässt sich die Claude Sonnet 5.5 API kostenlos testen?

Ja, neue Konten erhalten 10 Test-Calls und bis zu $1 kostenloses Guthaben, keine Kreditkarte erforderlich. Bei $2/M Input-Tokens deckt allein dieses Guthaben rund 62 Requests mit je ~8K Tokens gegen Claude Sonnet 5.5 ab.

Worin ist Claude Sonnet 5.5 am besten?

Beste Kombination aus Geschwindigkeit und Intelligenz im Angebot; Preis wie Sonnet 5: 2 USD Eingabe, 10 USD Ausgabe je Million Token; between_tools schaltet das vorgelagerte Denken ab. Das vollständige Bild finden Sie im Über-Abschnitt, direkt aus den offiziellen Release Notes des Anbieters.

Was kostet Claude Sonnet 5.5?

Claude Sonnet 5.5 kostet auf Synthorai $2 pro Million Input-Tokens und $10 pro Million Output-Tokens. Das ist der Listenpreis des Anbieters, ohne Plattform-Aufschlag. Gecachte Input-Tokens werden mit $0.2/M abgerechnet.

Unterstützt Claude Sonnet 5.5 Prompt-Caching?

Ja, per Opt-in: Markieren Sie stabile Präfixe mit cache_control-Breakpoints. Gecachte Input-Tokens werden mit $0.2/M statt $2/M (ungecacht) abgerechnet; Prompts brauchen ein stabiles Präfix von 512 Tokens, um gecacht zu werden (TTL Standard 5 Min., optional 1 Std.). Prompt-Caching-Guide →

Wie erhalte ich Zugang zu Claude Sonnet 5.5?

Richten Sie Ihr vorhandenes OpenAI SDK auf base_url="https://synthorai.io/v1", setzen Sie model="claude-sonnet-5-5", fertig. Ein API-Key deckt jedes Modell auf dem Gateway ab.

Wo liegt der Knowledge Cutoff von Claude Sonnet 5.5?

Der Knowledge Cutoff von Claude Sonnet 5.5 liegt bei 2026-06, laut offizieller Dokumentation des Anbieters (Stand: 2026-09-29).

Verwandte Modelle

Vergleichen

Jeder Wert auf dieser Seite ist aus der Dokumentation des Anbieters übernommen, oben verlinkt, und trägt das Datum der Prüfung. Preise werden über den gesamten Katalog verglichen; Spezifikationswerte, die Anbieter unterschiedlich definieren, werden mit benanntem Unterschied gezeigt statt grafisch verglichen. Nichts hier wird von uns gemessen, und nichts wird bewertet.

API-Schlüssel abrufen Ihre Kosten vergleichen →