Novità Registrati gratis, 10 chiamate le offriamo noi. Fino a $1, senza carta.

GPT-5.4 vs GPT-5.6 Luna

vs

Quale scegliere e quando

Entrambi offrono un contesto da 1050000 token e un tetto di 128000 token in uscita, quindi decidono prezzo e strumenti. gpt-5.6-luna sta sotto su ogni tariffa: $1 contro $2.5 per milione in ingresso, $6 contro $15 in uscita, $0.1 di lettura cache contro $1.25, circa 12 volte meno sul prefisso caldo in cui si trasforma un prompt di sistema ripetuto. Scegli gpt-5.4 se ti servono computer use o chiamate di strumenti in parallelo, che gpt-5.6-luna non elenca; sulle sole tariffe non vince nulla.

Benchmark

In testaSopra la mediaNessuno miglioreGPT-5.4621 / 371 / 37GPT-5.6 Luna815 / 420 / 42

14 misurati su entrambi.

GPT-5.4 GPT-5.6 Luna altri modelli misurati media dei modelli confrontati nessun altro modello ha fatto meglio
SWE-Bench Pro
57.7%
62.7%
GeneBench Pro
N/A
10.8%
OSWorld-Verified
75%
N/A
Cybergym
79%
N/A
HealthBench
54%
55.8%
GDPval-AA v2 Elo · 642-1861
N/A
1591.8
GPQA Diamond
92.8%
92.3%
BrowseComp
82.7%
83.3%
MMMU-Pro no tools
81.2%
78.4%

Dati pubblicati dai fornitori: Alibaba (Qwen) Anthropic ByteDance DeepSeek Google MiniMax Moonshot OpenAI Tencent Z.ai

Prezzi

GPT-5.4 GPT-5.6 Luna Δ
Input / 1M token $2.5 $1 2.5×
Output / 1M token $15 $6 2.5×
Lettura cache / 1M token $1.25 $0.1 13×
Scrittura in cache nessun addebito separato nessun addebito separato -

Le tariffe provengono dal catalogo live al momento della build; la pagina di ciascun modello riporta la scheda attuale.

Dove si posizionano - prezzo di input per 1M di token rispetto a tutti gli 64 modelli chat con questa unità di fatturazione (scala logaritmica)

GPT-5.4 · $2.5 GPT-5.6 Luna · $1
$0.05 · Qwen3 VL Flash $30 · GPT-5.4 Pro

Capacità

GPT-5.4 GPT-5.6 Luna
Utilizzo dei tool
Controllo del ragionamento configurabile configurabile
Output strutturato
Prompt caching implicito (automatico) implicito (automatico)
Durata della cache 5-10m, up to 1h 5-10m, up to 1h
Prefisso minimo in cache 1024 token 1024 token

Specifiche

GPT-5.4 GPT-5.6 Luna
Modalità di input testo immagine testo immagine
Modalità di output testo testo
Rilascio 2026-03-05 2026-07-09
Cutoff di conoscenza 2025-08 2026-02
Finestra di contesto 1.1M 1.1M
Output massimo 128K 128K
Parametro thinking reasoning.effort reasoning.effort
Valori accettati
reasoning.effort
  • none
  • low
  • medium
  • high
  • xhigh
reasoning.effort
  • none
  • low
  • medium
  • high
  • xhigh
  • max
Predefinito none medium

Le specifiche sono trascritte dalla documentazione di ciascun fornitore; una riga che un fornitore non pubblica viene omessa anziché essere dedotta. Fonti complete: GPT-5.4 · GPT-5.6 Luna

Un solo prompt, entrambi i modelli - misurati attraverso il gateway

PROMPT Explain to a CFO, in exactly three sentences, why our LLM bill doubled after we enabled reasoning. VERIFICA esattamente 3 frasi

GPT-5.4 superato · 3 sentences

Enabling reasoning makes the model generate and evaluate additional internal steps before producing an answer, which increases the total tokens processed and therefore the cost per request. It also tends to increase latency and can trigger more tool calls, retries, or longer context retention, all of which compound usage-based charges. So even if user traffic stayed flat, the average cost per interaction rose materially once reasoning was turned on.

out 82 tok latenza 1.9 s

GPT-5.6 Luna superato · 3 sentences

Enabling reasoning causes the model to use substantially more computational steps and generated tokens per request, even when the visible answer is similar. Because our provider charges based largely on token consumption and compute, that increased processing translated into roughly twice the cost. We can control the expense by reserving reasoning for complex tasks, setting token and effort limits, and monitoring cost per successful outcome.

out 146 tok (+60 ragionamento) latenza 3.1 s

Rispetto dell'istruzione (esattamente tre frasi - contabili), adattamento al pubblico (registro da CFO), e il divario di fatturazione del pensiero nascosto che i contatori di token qui sotto rivelano.

PROMPT This function is supposed to return the longest run of consecutive equal items, but callers report it is off by one on single-element inputs and misses runs that end at the last index. Fix it and explain the bug in one paragraph. VERIFICA la correzione supera i test

GPT-5.4 superato · 8/8 cases

```python def longest_run(xs): if not xs: return 0 best = 1 cur = 1 for i in range(1, len(xs)): if xs[i] == xs[i - 1]: cur += 1 else: best = max(best, cur) cur = 1 best = max(best, cur) return best ``` The bug is that the original code was counting *matching transitions* between adjacent elements instead of the *length of the run itself*, so a single element started at `0` instead of `1`, making all run lengths off by one. It also only updated `best` inside the equality case, which means the final run could be lost if the loop ended before a change triggered any correction logic. […]

out 243 tok latenza 3.0 s

GPT-5.6 Luna superato · 8/8 cases

```python def longest_run(xs): if not xs: return 0 best = cur = 1 for i in range(1, len(xs)): if xs[i] == xs[i - 1]: cur += 1 else: best = max(best, cur) cur = 1 return max(best, cur) ``` The bug is that `cur` counts equal *transitions* rather than the number of items in the current run, so a single-element run returns `0` and a run of length `n` returns `n - 1`; additionally, `best` is only updated when a run is interrupted, so a run continuing through the final element is not finalized before returning. Initializing the current run to one item, resetting it to one after a change, and taking a final `max` fixes both issues.

out 356 tok (+161 ragionamento) latenza 4.7 s

Se la correzione è davvero giusta (eseguibile), la densità della spiegazione, e l'efficienza in token su un compito delimitato.

PROMPT Extract a JSON object with fields {company, quarter, revenue_usd_m, growth_pct, guidance} from this text. Use null for anything not stated; add no other fields. VERIFICA JSON valido, schema esatto

GPT-5.4 superato · 5/5 fields, guidance "withheld pending the Q4 close"

{"company":"Meridian Robotics","quarter":"September quarter","revenue_usd_m":412,"growth_pct":11,"guidance":"withheld pending the Q4 close"}

out 39 tok latenza 1.3 s

GPT-5.6 Luna superato · 5/5 fields, guidance "Withheld pending the Q4 close"

{"company":"Meridian Robotics","quarter":"September quarter","revenue_usd_m":412,"growth_pct":null,"guidance":"Withheld pending the Q4 close"}

out 227 tok (+181 ragionamento) latenza 3.6 s

Aderenza allo schema (nessun campo inventato), pressione di allucinazione (guidance è esplicitamente sospesa), e differenze nel percorso di output strutturato.

PROMPT Write a 120-word product announcement for an API gateway feature that caches prompts across providers. Forbidden words: "seamless", "unlock", "game-changer", "revolutionize", "empower". Exactly one sentence must be a question. VERIFICA 120 parole, 0 parole vietate

GPT-5.4 superato · 126 words, 0 banned, 1 question

Today we’re introducing Prompt Cache, a new API gateway feature that stores prompt results and reuses them across supported model providers. Teams can reduce latency, lower repeated compute costs, and keep application behavior more consistent when identical or near-identical prompts appear in production. Prompt Cache works at the gateway layer, so developers can adopt it without rewriting provider integrations or changing routing logic. Built-in controls let you set cache keys, expiration policies, and privacy rules for sensitive workloads. Need faster responses without duplicating optimization work for each vendor? With shared observability, you can track hit rates, savings, and cache performance in one place. […]

out 154 tok latenza 2.9 s

GPT-5.6 Luna superato · 120 words, 0 banned, 1 question

Introducing PromptCache, an API gateway feature that caches prompts across providers, helping teams reduce latency, control spend, and deliver consistent results. How much faster could your applications respond when repeated prompts are served from a shared cache instead of being sent upstream? PromptCache supports provider-aware routing, configurable time-to-live policies, encrypted storage, cache invalidation, and usage analytics through one operational layer. It works with language-model providers while preserving your authentication, observability, and fallback workflows. Developers can enable caching by endpoint, model, tenant, or prompt pattern, then monitor hit rates and savings in real time. Built for production workloads, PromptCache gives platform teams controls for performance and cost without requiring application rewrites. […]

out 948 tok (+778 ragionamento) latenza 8.0 s

Rispetto dei vincoli (budget di parole, elenco di parole vietate, l'unica domanda), impronta stilistica, e controllo della lunghezza.

Passa dall'uno all'altro con una sola riga

Entrambi gli id sono presenti in ogni scheda qui sotto - la coppia di righe evidenziata è l'unica modifica. Stesso endpoint, stessa chiave, stessa struttura della richiesta.

from openai import OpenAI

client = OpenAI(
    base_url="https://synthorai.io/v1",
    api_key="sk-syn-...",
)

resp = client.chat.completions.create(
    model="gpt-5.4",
    # model="gpt-5.6-luna",  # decommenta questa riga, commenta quella sopra
    messages=[{"role": "user", "content": "Summarize this diff"}],
    reasoning_effort="medium",
)
print(resp.choices[0].message.content)

Ottieni la tua chiave API →

FAQ

Qual è più economico, GPT-5.4 o GPT-5.6 Luna?

GPT-5.6 Luna è più economico per input / 1m token ($1 contro $2.5, 2.5× di differenza). Altre righe potrebbero indicare il contrario - la tabella sopra riporta la scheda completa, e il costo reale dipende dal tuo mix.

Posso fare un A/B test di GPT-5.4 contro GPT-5.6 Luna senza due integrazioni?

Sì. Entrambi sono serviti tramite lo stesso endpoint compatibile con OpenAI con una singola chiave API - il passaggio richiede la modifica della stringa del modello in una sola riga, quindi puoi instradare una frazione del traffico verso ciascuno e confrontare direttamente le fatture.

GPT-5.4 e GPT-5.6 Luna supportano il prompt caching?

Sì - entrambi fatturano le letture in cache a un prezzo inferiore rispetto alla loro tariffa di input, quindi i carichi di lavoro con warm-prefix costano meno di quanto suggeriscano le tariffe di listino. Le righe esatte per la lettura in cache si trovano nella tabella dei prezzi qui sopra.

Confronti correlati

Dai nostri studi misurati