GPT Realtime 2.1 vs GPT Realtime 2.1 Mini
GPT Realtime 2.1 is served by invitation. Its figures below are the live rates, but calls need a workspace grant first; ask us for access before you build on this comparison.
GPT Realtime 2.1 Mini is served by invitation. Its figures below are the live rates, but calls need a workspace grant first; ask us for access before you build on this comparison.
Which one, when
The mini tier is cheaper on every line - $0.6 and $2.4 per million text tokens against $4 and $24, and $10 and $20 per million audio in and out against $32 and $64 - so roughly 3x less on audio and about 7x on text input. Both carry a 128K session context with speech-to-speech, tool use and prompt caching; the full model additionally documents configurable reasoning effort and interruption handling. Pick the mini for volume, the full model when you want to dial reasoning effort per session.
Pricing
| GPT Realtime 2.1 | GPT Realtime 2.1 Mini | Δ | |
|---|---|---|---|
| Audio input / 1M tokens | $32 | $10 | 3.2× |
| Audio output / 1M tokens | $64 | $20 | 3.2× |
| Audio cache read / 1M tokens | $0.4 | $0.3 | 1.3× |
| Text input / 1M tokens | $4 | $0.6 | 6.7× |
| Text output / 1M tokens | $24 | $2.4 | 10× |
| Cache write | no separate charge | no separate charge | - |
Rates from the live catalogue at build time; each model page carries the current rate card.
Where they sit · price per 1M audio tokens across all 6 realtime speech-to-speech models on this billing unit (log scale)
Capabilities
| GPT Realtime 2.1 | GPT Realtime 2.1 Mini | |
|---|---|---|
| Prompt caching | implicit (automatic) | implicit (automatic) |
| Cache lifetime | 5-10m, up to 1h | 5-10m, up to 1h |
| Minimum cached prefix | 1024 tokens | 1024 tokens |
Specs
| GPT Realtime 2.1 | GPT Realtime 2.1 Mini | |
|---|---|---|
| Input modalities | text audio | text audio |
| Output modalities | text audio | text audio |
| Released | 2026-07-06 | 2026-07-06 |
| Knowledge cutoff | 2024-09 | 2024-09 |
| Session capabilities |
|
|
| Context window | 128K | 128K |
Specs are transcribed from each vendor’s documentation; a row a vendor does not publish is left out rather than inferred. Full sources: GPT Realtime 2.1 · GPT Realtime 2.1 Mini
Switch between them with one line
Both ids are in every tab below; the highlighted pair of lines is the only edit. Same endpoint, same key, same request shape.
import asyncio, base64, json, websockets
URL = "wss://synthorai.io/v1/realtime?model=gpt-realtime-2.1"
# URL = "wss://synthorai.io/v1/realtime?model=gpt-realtime-2.1-mini" # uncomment this line, comment the one above
# Send ONLY the Authorization header - the beta protocol is retired.
HEADERS = {"Authorization": "Bearer sk-syn-..."}
async def main():
async with websockets.connect(URL, additional_headers=HEADERS) as ws:
# 1) configure the speech-to-speech session
await ws.send(json.dumps({
"type": "session.update",
"session": {
"type": "realtime",
"output_modalities": ["audio"],
"audio": {"output": {"voice": "alloy"}},
},
}))
# 2) send input audio (base64 PCM16), then request a spoken reply
await ws.send(json.dumps({"type": "input_audio_buffer.append", "audio": pcm16_b64}))
await ws.send(json.dumps({"type": "input_audio_buffer.commit"}))
await ws.send(json.dumps({"type": "response.create"}))
# 3) stream the model's audio (and text) back
async for raw in ws:
ev = json.loads(raw)
if ev["type"] == "response.audio.delta":
play(base64.b64decode(ev["delta"])) # audio out
elif ev["type"] == "response.done":
break
asyncio.run(main())import WebSocket from "ws";
const ws = new WebSocket("wss://synthorai.io/v1/realtime?model=gpt-realtime-2.1", {
// const ws = new WebSocket("wss://synthorai.io/v1/realtime?model=gpt-realtime-2.1-mini", { // uncomment this line, comment the one above
// Send ONLY the Authorization header — the beta protocol is retired.
headers: { Authorization: "Bearer sk-syn-..." },
});
ws.on("open", () => {
// configure the speech-to-speech session
ws.send(JSON.stringify({ type: "session.update", session: {
type: "realtime", output_modalities: ["audio"], audio: { output: { voice: "alloy" } },
} }));
// send input audio (base64 PCM16), then request a spoken reply
ws.send(JSON.stringify({ type: "input_audio_buffer.append", audio: pcm16Base64 }));
ws.send(JSON.stringify({ type: "input_audio_buffer.commit" }));
ws.send(JSON.stringify({ type: "response.create" }));
});
ws.on("message", (raw) => {
const ev = JSON.parse(raw.toString());
if (ev.type === "response.audio.delta") playAudio(Buffer.from(ev.delta, "base64")); // audio out
else if (ev.type === "response.done") ws.close();
});# Realtime is a WebSocket protocol - use a WS client such as websocat.
# Each line below is one OpenAI Realtime event (JSON) sent to the session.
websocat -H 'Authorization: Bearer sk-syn-...' \
'wss://synthorai.io/v1/realtime?model=gpt-realtime-2.1' <<'EOF'
# 'wss://synthorai.io/v1/realtime?model=gpt-realtime-2.1-mini' <<'EOF' # uncomment this line, comment the one above
{"type":"session.update","session":{"type":"realtime","output_modalities":["audio"],"audio":{"output":{"voice":"alloy"}}}}
{"type":"input_audio_buffer.append","audio":"<base64-pcm16>"}
{"type":"input_audio_buffer.commit"}
{"type":"response.create"}
EOF
# Responses stream back as response.audio.delta (base64 audio out) … response.donepackage main
import (
"net/http"
"github.com/gorilla/websocket"
)
func main() {
h := http.Header{}
h.Set("Authorization", "Bearer sk-syn-...")
// Send ONLY the Authorization header — the beta protocol is retired.
c, _, err := websocket.DefaultDialer.Dial("wss://synthorai.io/v1/realtime?model=gpt-realtime-2.1", h)
// c, _, err := websocket.DefaultDialer.Dial("wss://synthorai.io/v1/realtime?model=gpt-realtime-2.1-mini", h) // uncomment this line, comment the one above
if err != nil {
panic(err)
}
defer c.Close()
// configure the speech-to-speech session, send audio, request a spoken reply
c.WriteJSON(map[string]any{"type": "session.update", "session": map[string]any{
"type": "realtime", "output_modalities": []string{"audio"},
"audio": map[string]any{"output": map[string]any{"voice": "alloy"}}}})
c.WriteJSON(map[string]any{"type": "input_audio_buffer.append", "audio": pcm16B64})
c.WriteJSON(map[string]any{"type": "input_audio_buffer.commit"})
c.WriteJSON(map[string]any{"type": "response.create"})
for {
var ev struct {
Type string `json:"type"`
Delta string `json:"delta"`
}
if err := c.ReadJSON(&ev); err != nil {
return
}
if ev.Type == "response.audio.delta" {
playAudio(ev.Delta) // base64 audio out
} else if ev.Type == "response.done" {
return
}
}
}import java.net.URI;
import java.net.http.HttpClient;
import java.net.http.WebSocket;
import java.util.concurrent.CompletionStage;
// JDK built-in WebSocket — no extra dependency needed.
WebSocket ws = HttpClient.newHttpClient().newWebSocketBuilder()
.header("Authorization", "Bearer sk-syn-...")
// Send ONLY the Authorization header — the beta protocol is retired.
.buildAsync(URI.create("wss://synthorai.io/v1/realtime?model=gpt-realtime-2.1"), new WebSocket.Listener() {
// .buildAsync(URI.create("wss://synthorai.io/v1/realtime?model=gpt-realtime-2.1-mini"), new WebSocket.Listener() { // uncomment this line, comment the one above
public CompletionStage<?> onText(WebSocket w, CharSequence data, boolean last) {
// handle response.audio.delta (base64 audio out) / response.done here
w.request(1);
return null;
}
}).join();
// configure the session, send input audio, then request a spoken reply
ws.sendText("{\"type\":\"session.update\",\"session\":{\"type\":\"realtime\",\"output_modalities\":[\"audio\"],\"audio\":{\"output\":{\"voice\":\"alloy\"}}}}", true);
ws.sendText("{\"type\":\"input_audio_buffer.append\",\"audio\":\"<base64-pcm16>\"}", true);
ws.sendText("{\"type\":\"input_audio_buffer.commit\"}", true);
ws.sendText("{\"type\":\"response.create\"}", true);FAQ
Which is cheaper, GPT Realtime 2.1 or GPT Realtime 2.1 Mini?
GPT Realtime 2.1 Mini is cheaper on the "Audio input / 1M tokens" row ($10 vs $32, 3.2× apart). Other rows may point the other way; the table above carries the full rate card, and real cost depends on your mix.
Can I A/B test GPT Realtime 2.1 against GPT Realtime 2.1 Mini without two integrations?
Yes. Both are served through the same OpenAI-compatible endpoint with one API key. Switching is a one-line change to the model id, so you can route a fraction of traffic to each and compare bills directly.
Do GPT Realtime 2.1 and GPT Realtime 2.1 Mini support prompt caching?
Yes. Both bill cache reads below their input rate, so warm-prefix workloads cost less than the list rates suggest. The exact cache-read rows are in the pricing table above.