New Sign up free, 10 calls on us. Up to $1, no card needed.

GPT Realtime 2.1

Released 2026-07-06

invited beta

GPT Realtime 2.1 is the flagship of OpenAI's realtime speech-to-speech family, announced as an updated realtime reasoning model with improved alphanumeric recognition, silence and noise handling, and interruption behavior.

Input
audio text $4/M
Output
audio text $24/M
Audio input
$32/M
Audio output
$64/M
Cache read
$0.4/M
Knowledge cutoff
2024-09

Price in context

Where the price sits among 6 comparable models

Input$4/M
$0.06 · nova-2-sonic GPT Realtime · $4
Output$24/M
$0.24 · nova-2-sonic GPT Realtime 2.1 · $24
Cached read$0.4/M
$0.06 · GPT Realtime 2.1 Mini GPT Realtime · $0.4

The bar shows how this model’s price compares with every other model of the same kind on Synthorai. The cheapest and the most expensive are named at each end. These are base rates; batch, region and cache-write discounts are on the pricing page.

Specs & limits

Tokens

Context window (vendor spec) 128,000
Max output (vendor spec) 32,000
Knowledge cutoff 2024-09

Prompt caching

How it caches automatic
Min prefix 1,024
Lifetime 5-10m, up to 1h

Realtime

Session
  • Speech-to-speech over WebSocket
  • configurable reasoning effort
  • tool use
  • interruption handling
  • 128K context

Model

Modalities audio + text → audio + text
  • Updated realtime reasoning model with improved alphanumeric recognition, silence and noise handling, and interruption behavior
  • configurable reasoning effort and tool use for voice-agent workflows
  • 128k context

per OpenAI official docs ↗

Use GPT Realtime 2.1 in 30 seconds

OpenAI Realtime-compatible: connect over WebSocket and stream audio in, audio out. WS /v1/realtime

import asyncio, base64, json, websockets

URL = "wss://synthorai.io/v1/realtime?model=gpt-realtime-2.1"
# Send ONLY the Authorization header - the beta protocol is retired.
HEADERS = {"Authorization": "Bearer sk-syn-..."}

async def main():
    async with websockets.connect(URL, additional_headers=HEADERS) as ws:
        # 1) configure the speech-to-speech session
        await ws.send(json.dumps({
            "type": "session.update",
            "session": {
                "type": "realtime",
                "output_modalities": ["audio"],
                "audio": {"output": {"voice": "alloy"}},
            },
        }))
        # 2) send input audio (base64 PCM16), then request a spoken reply
        await ws.send(json.dumps({"type": "input_audio_buffer.append", "audio": pcm16_b64}))
        await ws.send(json.dumps({"type": "input_audio_buffer.commit"}))
        await ws.send(json.dumps({"type": "response.create"}))
        # 3) stream the model's audio (and text) back
        async for raw in ws:
            ev = json.loads(raw)
            if ev["type"] == "response.audio.delta":
                play(base64.b64decode(ev["delta"]))   # audio out
            elif ev["type"] == "response.done":
                break

asyncio.run(main())

About GPT Realtime 2.1

  • It expands the family to a 128K-token context window with 32K max output tokens, and supports configurable reasoning effort, instruction following, and tool use for complex voice-agent workflows.
  • Audio and text tokens are billed at separate per-token rates, with discounted cached input.
  • OpenAI's realtime guide names it the model to pick when building a low-latency voice agent, and recommends starting at low reasoning effort for most production agents, then raising it only where the task justifies the added latency.
  • Sessions run over WebRTC, WebSocket, or SIP, and turn-taking can use silence-based voice activity detection or a semantic detector whose eagerness setting decides how patiently it waits for a speaker to finish.
  • Interruption handling is part of the protocol: with detection enabled the server truncates unplayed audio, and WebSocket clients are expected to stop playback and report how much audio actually reached the listener.
  • Tools can be declared per session or per response, and responses can be generated outside the main conversation when a side task should not pollute the transcript.
  • On Synthorai it is served through the OpenAI-compatible /v1/realtime WebSocket endpoint, so applications built on the official Realtime SDKs work unchanged.

FAQ

Is the GPT Realtime 2.1 API free to try?

GPT Realtime 2.1 is currently in invited beta: access is application-based rather than open signup. Apply from the Synthorai console; once approved, standard pay-as-you-go pricing applies with no subscription.

What is GPT Realtime 2.1 best at?

Flagship realtime tier for production voice agents; improved interruption, noise, and alphanumeric handling; configurable reasoning effort with tool use. See the About section for the full picture from the vendor's own release notes.

How much does GPT Realtime 2.1 cost?

GPT Realtime 2.1 costs $4 per million input tokens and $24 per million output tokens on Synthorai. That is the provider's list price, with no platform markup. Cached input tokens bill at $0.4/M.

How do I connect to the GPT Realtime 2.1 API?

GPT Realtime 2.1 is a speech-to-speech model: connect over WebSocket to wss://synthorai.io/v1/realtime?model=gpt-realtime-2.1 with the OpenAI Realtime SDK (or a raw WebSocket) and stream audio in, audio out. It is not a POST /v1/audio/transcriptions file upload. Authenticate with your sk-syn key, sending only the Authorization header (the beta protocol is retired).

How do I get access to GPT Realtime 2.1?

GPT Realtime 2.1 is in invited beta: request access from the Synthorai console. Once approved it works like every other model: point your OpenAI SDK at base_url="https://synthorai.io/v1" and set model="gpt-realtime-2.1".

Related models

Compare

Every value on this page is transcribed from the vendor's own documentation, linked above, and carries the date it was checked. Prices are compared across the catalogue; specification values that vendors define differently are shown with the difference stated rather than charted. Nothing here is measured by us, and nothing is scored.

Get your API key Compare your cost →