🎁 New Sign up free, 10 calls on us. Up to $1, no card needed.

GPT Realtime 2.1 Mini Released 2026-07-06

OpenAI invited beta realtime audio
Input$0.6/M
Output$2.4/M
Audio input$10/M
Audio output$20/M
Cache read$0.06/M
Knowledge cutoff2024-09

Provider list prices: no platform markup, pay-as-you-go. These are official list prices. Logged-in customers may see effective prices including workspace discounts on /console/pricing. Effective input at a 70% cache-hit rate:$0.222/M. Automatic prefix caching once a prompt passes the minimum length, with no code changes; cached reads are discounted (up to 90% on current models) and there is no write fee.

Use GPT Realtime 2.1 Mini in 30 seconds

OpenAI Realtime-compatible: connect over WebSocket and stream audio in, audio out. WS /v1/realtime

import asyncio, base64, json, websockets

URL = "wss://synthorai.io/v1/realtime?model=gpt-realtime-2.1-mini"
HEADERS = {"Authorization": "Bearer sk-syn-...", "OpenAI-Beta": "realtime=v1"}

async def main():
    async with websockets.connect(URL, additional_headers=HEADERS) as ws:
        # 1) configure the speech-to-speech session
        await ws.send(json.dumps({
            "type": "session.update",
            "session": {"modalities": ["audio", "text"], "voice": "alloy"},
        }))
        # 2) send input audio (base64 PCM16), then request a spoken reply
        await ws.send(json.dumps({"type": "input_audio_buffer.append", "audio": pcm16_b64}))
        await ws.send(json.dumps({"type": "input_audio_buffer.commit"}))
        await ws.send(json.dumps({"type": "response.create"}))
        # 3) stream the model's audio (and text) back
        async for raw in ws:
            ev = json.loads(raw)
            if ev["type"] == "response.audio.delta":
                play(base64.b64decode(ev["delta"]))   # audio out
            elif ev["type"] == "response.done":
                break

asyncio.run(main())

About GPT Realtime 2.1 Mini

Distilled reasoning for realtime voice
Audio input under a third of the flagship's rate
Keeps 128K context, tool use, and prompt caching

GPT Realtime 2.1 Mini is the cost-efficient member of OpenAI's realtime family: a faster, lower-cost distilled reasoning model for realtime voice applications that brings reasoning and tool use to the mini tier.

  • It keeps the 2.1 generation's 128K-token context window and 32K max output tokens, supports function calling and prompt caching, and prices audio input at under a third of the flagship's rate.
  • That makes it the natural pick for always-on voice assistants and high-volume customer-facing voice agents.
  • Synthorai exposes it via the same OpenAI-compatible /v1/realtime WebSocket endpoint as the rest of the realtime lineup.

Specs & limits

Max output (vendor spec)32,000
Modalitiesaudio + text → audio + text
Featuresrealtime · speech-to-speech
NotableFaster, lower-cost distilled realtime reasoning model bringing reasoning and tool use to the mini tier; prompt caching supported; 128k context.
Prompt cachingautomatic · min 1,024-token prefix · TTL 5–10m, up to 1h

per OpenAI official docs ↗

FAQ

Is the GPT Realtime 2.1 Mini API free to try?

GPT Realtime 2.1 Mini is currently in invited beta: access is application-based rather than open signup. Apply from the Synthorai console; once approved, standard pay-as-you-go pricing applies with no subscription.

What is GPT Realtime 2.1 Mini best at?

Distilled reasoning for realtime voice, plus audio input under a third of the flagship's rate and keeps 128K context, tool use, and prompt caching. See the About section for the full picture from the vendor's own release notes.

How much does GPT Realtime 2.1 Mini cost?

GPT Realtime 2.1 Mini costs $0.6 per million input tokens and $2.4 per million output tokens on Synthorai. That is the provider's list price, with no platform markup. Cached input tokens bill at $0.06/M.

How do I connect to the GPT Realtime 2.1 Mini API?

GPT Realtime 2.1 Mini is a speech-to-speech model: connect over WebSocket to wss://synthorai.io/v1/realtime?model=gpt-realtime-2.1-mini with the OpenAI Realtime SDK (or a raw WebSocket) and stream audio in, audio out. It is not a POST /v1/audio/transcriptions file upload. Authenticate with your sk-syn key and mirror the OpenAI-Beta header.

How do I get access to GPT Realtime 2.1 Mini?

GPT Realtime 2.1 Mini is in invited beta: request access from the Synthorai console. Once approved it works like every other model: point your OpenAI SDK at base_url="https://synthorai.io/v1" and set model="gpt-realtime-2.1-mini".

Related models

Get your free API key Compare your cost →