Claude Haiku 4.5 vs Claude Sonnet 5
Qual usar e quando
O Sonnet 5 dobra a tabela do Haiku ($2/$10 vs $1/$5) e quintuplica o contexto (1M vs 200K). O Haiku continua sendo o piso de latência e preço da linha Claude; no momento em que os prompts ultrapassarem 200K ou precisarem de raciocínio mais profundo, o Sonnet é o destino natural.
Benchmarks
Publicado pelos provedores: Alibaba (Qwen) Anthropic ByteDance DeepSeek Google Moonshot OpenAI Tencent Z.ai
Preços
| Claude Haiku 4.5 | Claude Sonnet 5 | Δ | |
|---|---|---|---|
| Entrada / 1M tokens | $1 | $2 | 0.5× |
| Saída / 1M tokens | $5 | $10 | 0.5× |
| Leitura de cache / 1M tokens | $0.1 | $0.2 | 0.5× |
| Escrita de cache | 1.25x (5m) / 2x (1h) | 1.25x (5m) / 2x (1h) | - |
Tarifas do catálogo em tempo real, lidas no momento do build; a página de cada modelo traz a tabela de preços atual.
Onde cada um fica: preço de entrada por 1M tokens entre os 76 modelos de chat com esta unidade de cobrança (escala logarítmica)
Capacidades
| Claude Haiku 4.5 | Claude Sonnet 5 | |
|---|---|---|
| Uso de ferramentas | sim | sim |
| Controle de raciocínio | configurável | configurável |
| Saída estruturada | sim | sim |
| Cache de prompts | explícito (você marca o prefixo) | explícito (você marca o prefixo) |
| Tempo de vida do cache | 5m default, 1h option | 5m default, 1h option |
| Prefixo mínimo em cache | 1024 tokens | 1024 tokens |
Especificações
| Claude Haiku 4.5 | Claude Sonnet 5 | |
|---|---|---|
| Modalidades de entrada | texto imagem | texto imagem |
| Modalidades de saída | texto | texto |
| Lançamento | 2025-10-15 | 2026-06-30 |
| Corte de conhecimento | 2025-02 | 2026-01 |
| Janela de contexto | 200K | 1M |
| Saída máxima | 64K | 128K |
| Parâmetro de raciocínio |
|
|
| Valores aceitos | budget_tokens
"adaptive" returns 400 | thinking.type
effort
"enabled" returns 400 |
| Padrão | off, with no thinking unless the request sets it | thinking on (adaptive) effort
|
As especificações são transcritas da documentação de cada provedor; quando um provedor não publica um dado, a linha é omitida, e não inferida. Fontes completas: Claude Haiku 4.5 · Claude Sonnet 5
Um prompt, os dois modelos, medidos pelo gateway
Claude Haiku 4.5 passou · 3 sentences
# Why Your LLM Bill Doubled Reasoning models like o1 and o3 process queries through extended chains of thought before responding, consuming 5-10x more tokens internally than standard models while also charging higher per-token rates. This means both your input token usage and per-unit costs increased simultaneously—a compounding effect that easily doubles overall spend. Unless you've significantly reduced queries to other models or have very selective reasoning deployment, the math makes a doubling entirely predictable.
saída 106 tok latência 2.8 s
Claude Sonnet 5 passou · 3 sentences
Reasoning mode works by having the model generate extended internal "thinking" steps before producing its final answer, and since we're billed on total tokens, that hidden reasoning text adds significantly to what we pay for—even though you never see it in the output. This is different from simply getting longer visible responses; the model is essentially "showing its work" behind the scenes, and complex prompts can trigger thousands of extra reasoning tokens per request. The upside is improved accuracy on complex tasks, but if we want to control costs, we should selectively enable reasoning only for queries that truly need it, rather than applying it universally.
saída 194 tok latência 4.3 s
Se a instrução é seguida (exatamente três frases, dá para contar), a adequação ao público (tom para um CFO) e a diferença de cobrança causada pelo raciocínio oculto, que os medidores de tokens abaixo deixam à mostra.
Claude Haiku 4.5 passou · 8/8 cases
# Fixed Function ```python def longest_run(xs): if not xs: return 0 best = 1 cur = 1 for i in range(1, len(xs)): if xs[i] == xs[i - 1]: cur += 1 best = max(best, cur) else: cur = 1 return best ``` # Bug Explanation The original function has two related bugs stemming from incorrect initialization and loop logic. First, `best` and `cur` start at 0, which causes single-element lists to incorrectly return 0 instead of 1—a run of one identical element should have length 1. Second, `cur` is reset to 0 (not 1) when a mismatch occurs, causing the next matching pair to be undercounted. More critically, if the longest run ends at the last index, `best` is only updated *during* the loop when a mismatch is encountered; if no mismatch follows the final run, `best` never captures it. […]
saída 310 tok latência 4.0 s
Claude Sonnet 5 passou · 8/8 cases
## The Bug The function has two related off-by-one issues rooted in the same cause: `cur` and `best` are initialized to `0` instead of `1`, treating run lengths as the number of *matches* between adjacent elements rather than the number of *elements* in the run. This means a single-element input never enters the loop (since `range(1, 1)` is empty) and returns `0` instead of the correct answer of `1`. It also means that for any input, the run length is undercounted by exactly one — a run of 3 equal elements only produces 2 "match" increments, so `best` ends up as 2 instead of 3. The claim that it "misses runs that end at the last index" is actually a symptom of this same undercount: the logic does update `best` correctly on every iteration (including the last), but because it's counting matches instead of elements, the final reported value is always one less than the true longest run. […]
saída 536 tok latência 7.1 s
Se a correção está mesmo certa (dá para executar), a densidade da explicação e a eficiência em tokens em uma tarefa bem delimitada.
Claude Haiku 4.5 passou · 5/5 fields, guidance null
```json { "company": "Meridian Robotics", "quarter": "Q3", "revenue_usd_m": 412, "growth_pct": 11, "guidance": null } ```
saída 63 tok latência 1.6 s
Claude Sonnet 5 passou · 5/5 fields, guidance null
```json { "company": "Meridian Robotics", "quarter": "Q3", "revenue_usd_m": 412, "growth_pct": 11.5, "guidance": null } ```
saída 348 tok latência 5.3 s
Aderência ao schema (sem campos inventados), pressão para alucinar (o texto diz explicitamente que o guidance não foi divulgado) e diferenças no caminho de saída estruturada.
Claude Haiku 4.5 passou · 123 words, 0 banned, 1 question
# Introducing Prompt Cache Across Providers We're excited to announce our latest API gateway feature: intelligent prompt caching that works across multiple AI providers. This powerful capability stores frequently used prompts and their contexts, dramatically reducing latency and API costs for your applications. By intelligently managing cached prompts across providers like OpenAI, Anthropic, and others, you can optimize your infrastructure without changing your code. Why wait for responses when cached results can be delivered instantly? The system automatically handles cache invalidation and updates, ensuring your applications always access current information while maintaining performance gains. With support for complex multi-turn conversations and dynamic content, this feature scales seamlessly with your business needs. […]
saída 165 tok latência 3.0 s
Claude Sonnet 5 passou · 123 words, 0 banned, 1 question
## Introducing Cross-Provider Prompt Caching We're rolling out a new caching layer for our API gateway that works across every major LLM provider you connect through it. Instead of resending identical prompts and paying full price each time, the gateway recognizes repeated requests and serves cached responses in milliseconds—regardless of whether you're routing to OpenAI, Anthropic, or Google. What does this mean for your bill and your latency? Both drop, often significantly, especially for teams running high-volume, repetitive workloads like customer support bots or batch content generation. The cache is configurable per route, with adjustable TTLs and invalidation rules, so you stay in control of freshness versus cost. Available now for all Pro and Enterprise plans. Check your dashboard to enable it today.
saída 259 tok latência 4.8 s
Respeito às restrições (limite de palavras, lista de palavras proibidas, uma única pergunta), a marca de estilo de cada modelo e o controle do tamanho.
Troque de um para o outro mudando uma linha
Os dois ids estão em todas as abas abaixo; o par de linhas destacado é a única alteração. O endpoint, a chave e o formato da requisição são os mesmos.
from openai import OpenAI
client = OpenAI(
base_url="https://synthorai.io/v1",
api_key="sk-syn-...",
)
resp = client.chat.completions.create(
model="claude-haiku-4-5",
# model="claude-sonnet-5", # descomente esta linha, comente a linha acima
messages=[{"role": "user", "content": "Summarize this diff"}],
)
print(resp.choices[0].message.content)import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://synthorai.io/v1",
apiKey: "sk-syn-...",
});
const resp = await client.chat.completions.create({
model: "claude-haiku-4-5",
// model: "claude-sonnet-5", // descomente esta linha, comente a linha acima
messages: [{ role: "user", content: "Summarize this diff" }],
});
console.log(resp.choices[0].message.content);curl https://synthorai.io/v1/chat/completions \
-H "Authorization: Bearer sk-syn-..." \
-H "Content-Type: application/json" \
-d '{
"model": "claude-haiku-4-5",
# "model": "claude-sonnet-5", # descomente esta linha, comente a linha acima
"messages": [{"role": "user", "content": "Hello"}]
}'package main
import (
"context"
"fmt"
"github.com/openai/openai-go/v3"
"github.com/openai/openai-go/v3/option"
)
func main() {
client := openai.NewClient(
option.WithBaseURL("https://synthorai.io/v1"),
option.WithAPIKey("sk-syn-..."),
)
resp, _ := client.Chat.Completions.New(context.TODO(), openai.ChatCompletionNewParams{
Model: "claude-haiku-4-5",
// Model: "claude-sonnet-5", // descomente esta linha, comente a linha acima
Messages: []openai.ChatCompletionMessageParamUnion{
openai.UserMessage("Summarize this diff"),
},
})
fmt.Println(resp.Choices[0].Message.Content)
}import com.openai.client.OpenAIClient;
import com.openai.client.okhttp.OpenAIOkHttpClient;
import com.openai.models.chat.completions.*;
OpenAIClient client = OpenAIOkHttpClient.builder()
.baseUrl("https://synthorai.io/v1")
.apiKey("sk-syn-...")
.build();
ChatCompletion resp = client.chat().completions().create(
ChatCompletionCreateParams.builder()
.model("claude-haiku-4-5")
// .model("claude-sonnet-5") // descomente esta linha, comente a linha acima
.addUserMessage("Summarize this diff")
.build());
System.out.println(resp.choices().get(0).message().content().orElse(""));Perguntas frequentes
Qual é mais barato, Claude Haiku 4.5 ou Claude Sonnet 5?
Claude Haiku 4.5 sai mais barato na linha “Entrada / 1M tokens” ($1 contra $2, 2.0× de diferença). Outras linhas podem pender para o outro lado: a tabela acima mostra todos os preços, e o custo real depende do seu mix de uso.
Posso fazer um teste A/B de Claude Haiku 4.5 contra Claude Sonnet 5 sem duas integrações?
Sim. Os dois são servidos pelo mesmo endpoint compatível com OpenAI, com uma única chave de API. Para trocar, basta mudar uma linha, a string do modelo; assim você pode direcionar uma parte do tráfego para cada um e comparar as contas diretamente.
Claude Haiku 4.5 e Claude Sonnet 5 suportam cache de prompts?
Sim. Os dois cobram as leituras de cache abaixo da tarifa de entrada, então cargas de trabalho com prefixo já em cache custam menos do que os preços de tabela sugerem. Os valores exatos de leitura de cache estão na tabela de preços acima.