GPT Realtime 2.1 vs nova-2-sonic
GPT Realtime 2.1 è disponibile su invito. Le cifre qui sotto sono le tariffe live, ma per chiamarlo il tuo workspace deve prima essere autorizzato: chiedici l'accesso prima di basare il tuo lavoro su questo confronto.
nova-2-sonic è disponibile su invito. Le cifre qui sotto sono le tariffe live, ma per chiamarlo il tuo workspace deve prima essere autorizzato: chiedici l'accesso prima di basare il tuo lavoro su questo confronto.
Quale scegliere e quando
Entrambi sono modelli audio e di testo in tempo reale fatturati per milione di token, quindi le tariffe sono confrontabili direttamente: nova-2-sonic addebita $3 per l'audio in input e $12 per l'audio in output contro $32 e $64 per gpt-realtime-2.1, e $0.06/$0.24 sul testo rispetto a $4/$24, con 1000000 di contesto e 64000 di output massimo invece di 128000 e 32000. Scegli nova-2-sonic per sessioni vocali lunghe e attente ai costi, se il suo limite di connessione di 8 minuti e l'inferenza solo in-Region si adattano alle tue esigenze; scegli il più recente gpt-realtime-2.1 (rilasciato il 2026-07-06) quando desideri letture di input in cache a $0.4 e nessun vincolo di sessione o regione del genere
Benchmark
GPT Realtime 2.1: il provider non ha pubblicato risultati di benchmark.
nova-2-sonic: 8 pubblicati, ma nessun benchmark in comune con abbastanza altri modelli per fare un confronto.
Prezzi
| GPT Realtime 2.1 | nova-2-sonic | Δ | |
|---|---|---|---|
| Input audio / 1M token | $32 | $3 | 11× |
| Output audio / 1M token | $64 | $12 | 5.3× |
| Lettura cache audio / 1M token | $0.4 | - | - |
| Input di testo / 1M token | $4 | $0.06 | 67× |
| Output di testo / 1M token | $24 | $0.24 | 100× |
| Scrittura in cache | nessun addebito separato | nessun addebito separato | - |
Tariffe lette dal catalogo live al momento della build; il listino aggiornato è sulla pagina di ciascun modello.
Dove si collocano: prezzo per 1M di token audio tra tutti i modelli di speech-to-speech in tempo reale con questa unità di fatturazione (6, scala logaritmica)
Funzionalità
| GPT Realtime 2.1 | nova-2-sonic | |
|---|---|---|
| Prompt caching | implicito (automatico) | non supportato |
| Durata della cache | 5-10m, up to 1h | non applicabile |
| Prefisso minimo in cache | 1024 token | non applicabile |
Specifiche
| GPT Realtime 2.1 | nova-2-sonic | |
|---|---|---|
| Modalità di input | testo audio | testo audio |
| Modalità di output | testo audio | testo audio |
| Rilascio | 2026-07-06 | 2025-12-02 |
| Knowledge cutoff | 2024-09 | - |
| Voci | - | Feminine- and masculine-sounding voices per locale (tiffany/matthew en-US, amy en-GB, olivia en-AU, kiara/arjun en-IN and hi-IN, ambre/florian fr-FR, beatrice/lorenzo it-IT, tina/lennart de-DE, lupe/carlos es-US, carolina/leo pt-BR) tiffany and matthew are polyglot voices that speak every supported language. |
| Funzionalità di sessione |
|
|
| Finestra di contesto | 128K | 1M |
Le specifiche sono riprese dalla documentazione di ciascun provider; se un provider non pubblica un dato, la riga viene omessa e non dedotta. Fonti complete: GPT Realtime 2.1 · nova-2-sonic
Passa dall'uno all'altro cambiando una sola riga
In ogni scheda qui sotto ci sono entrambi gli id: le due righe evidenziate sono l'unica modifica. Stesso endpoint, stessa chiave, stessa struttura della richiesta.
import asyncio, base64, json, websockets
URL = "wss://synthorai.io/v1/realtime?model=gpt-realtime-2.1"
# URL = "wss://synthorai.io/v1/realtime?model=nova-2-sonic" # decommenta questa riga, commenta quella sopra
# Send ONLY the Authorization header - the beta protocol is retired.
HEADERS = {"Authorization": "Bearer sk-syn-..."}
async def main():
async with websockets.connect(URL, additional_headers=HEADERS) as ws:
# 1) configure the speech-to-speech session
await ws.send(json.dumps({
"type": "session.update",
"session": {
"type": "realtime",
"output_modalities": ["audio"],
"audio": {"output": {"voice": "alloy"}},
},
}))
# 2) send input audio (base64 PCM16), then request a spoken reply
await ws.send(json.dumps({"type": "input_audio_buffer.append", "audio": pcm16_b64}))
await ws.send(json.dumps({"type": "input_audio_buffer.commit"}))
await ws.send(json.dumps({"type": "response.create"}))
# 3) stream the model's audio (and text) back
async for raw in ws:
ev = json.loads(raw)
if ev["type"] == "response.audio.delta":
play(base64.b64decode(ev["delta"])) # audio out
elif ev["type"] == "response.done":
break
asyncio.run(main())import WebSocket from "ws";
const ws = new WebSocket("wss://synthorai.io/v1/realtime?model=gpt-realtime-2.1", {
// const ws = new WebSocket("wss://synthorai.io/v1/realtime?model=nova-2-sonic", { // decommenta questa riga, commenta quella sopra
// Send ONLY the Authorization header — the beta protocol is retired.
headers: { Authorization: "Bearer sk-syn-..." },
});
ws.on("open", () => {
// configure the speech-to-speech session
ws.send(JSON.stringify({ type: "session.update", session: {
type: "realtime", output_modalities: ["audio"], audio: { output: { voice: "alloy" } },
} }));
// send input audio (base64 PCM16), then request a spoken reply
ws.send(JSON.stringify({ type: "input_audio_buffer.append", audio: pcm16Base64 }));
ws.send(JSON.stringify({ type: "input_audio_buffer.commit" }));
ws.send(JSON.stringify({ type: "response.create" }));
});
ws.on("message", (raw) => {
const ev = JSON.parse(raw.toString());
if (ev.type === "response.audio.delta") playAudio(Buffer.from(ev.delta, "base64")); // audio out
else if (ev.type === "response.done") ws.close();
});# Realtime is a WebSocket protocol - use a WS client such as websocat.
# Each line below is one OpenAI Realtime event (JSON) sent to the session.
websocat -H 'Authorization: Bearer sk-syn-...' \
'wss://synthorai.io/v1/realtime?model=gpt-realtime-2.1' <<'EOF'
# 'wss://synthorai.io/v1/realtime?model=nova-2-sonic' <<'EOF' # decommenta questa riga, commenta quella sopra
{"type":"session.update","session":{"type":"realtime","output_modalities":["audio"],"audio":{"output":{"voice":"alloy"}}}}
{"type":"input_audio_buffer.append","audio":"<base64-pcm16>"}
{"type":"input_audio_buffer.commit"}
{"type":"response.create"}
EOF
# Responses stream back as response.audio.delta (base64 audio out) … response.donepackage main
import (
"net/http"
"github.com/gorilla/websocket"
)
func main() {
h := http.Header{}
h.Set("Authorization", "Bearer sk-syn-...")
// Send ONLY the Authorization header — the beta protocol is retired.
c, _, err := websocket.DefaultDialer.Dial("wss://synthorai.io/v1/realtime?model=gpt-realtime-2.1", h)
// c, _, err := websocket.DefaultDialer.Dial("wss://synthorai.io/v1/realtime?model=nova-2-sonic", h) // decommenta questa riga, commenta quella sopra
if err != nil {
panic(err)
}
defer c.Close()
// configure the speech-to-speech session, send audio, request a spoken reply
c.WriteJSON(map[string]any{"type": "session.update", "session": map[string]any{
"type": "realtime", "output_modalities": []string{"audio"},
"audio": map[string]any{"output": map[string]any{"voice": "alloy"}}}})
c.WriteJSON(map[string]any{"type": "input_audio_buffer.append", "audio": pcm16B64})
c.WriteJSON(map[string]any{"type": "input_audio_buffer.commit"})
c.WriteJSON(map[string]any{"type": "response.create"})
for {
var ev struct {
Type string `json:"type"`
Delta string `json:"delta"`
}
if err := c.ReadJSON(&ev); err != nil {
return
}
if ev.Type == "response.audio.delta" {
playAudio(ev.Delta) // base64 audio out
} else if ev.Type == "response.done" {
return
}
}
}import java.net.URI;
import java.net.http.HttpClient;
import java.net.http.WebSocket;
import java.util.concurrent.CompletionStage;
// JDK built-in WebSocket — no extra dependency needed.
WebSocket ws = HttpClient.newHttpClient().newWebSocketBuilder()
.header("Authorization", "Bearer sk-syn-...")
// Send ONLY the Authorization header — the beta protocol is retired.
.buildAsync(URI.create("wss://synthorai.io/v1/realtime?model=gpt-realtime-2.1"), new WebSocket.Listener() {
// .buildAsync(URI.create("wss://synthorai.io/v1/realtime?model=nova-2-sonic"), new WebSocket.Listener() { // decommenta questa riga, commenta quella sopra
public CompletionStage<?> onText(WebSocket w, CharSequence data, boolean last) {
// handle response.audio.delta (base64 audio out) / response.done here
w.request(1);
return null;
}
}).join();
// configure the session, send input audio, then request a spoken reply
ws.sendText("{\"type\":\"session.update\",\"session\":{\"type\":\"realtime\",\"output_modalities\":[\"audio\"],\"audio\":{\"output\":{\"voice\":\"alloy\"}}}}", true);
ws.sendText("{\"type\":\"input_audio_buffer.append\",\"audio\":\"<base64-pcm16>\"}", true);
ws.sendText("{\"type\":\"input_audio_buffer.commit\"}", true);
ws.sendText("{\"type\":\"response.create\"}", true);FAQ
Qual è il più economico, GPT Realtime 2.1 o nova-2-sonic?
nova-2-sonic costa meno alla voce Input audio / 1M token ($3 contro $32, 11× di differenza). Altre voci potrebbero dire il contrario: la tabella qui sopra riporta il listino completo, e il costo reale dipende dal tuo mix di utilizzo.
Posso fare un A/B test di GPT Realtime 2.1 contro nova-2-sonic senza due integrazioni?
Sì. Si chiamano entrambi dallo stesso endpoint compatibile con OpenAI, con una sola chiave API. Per passare dall'uno all'altro basta cambiare la stringa del modello in una riga, quindi puoi mandare una parte del traffico a ciascuno e confrontare direttamente i costi.
GPT Realtime 2.1 e nova-2-sonic supportano il prompt caching?
Nel nostro feed il prezzo di lettura dalla cache è indicato solo per uno dei due; dove la tariffa manca, il provider non applica un prezzo separato alle letture dalla cache.