Gemini 3.1 Flash Live is Google's low-latency, audio-to-audio model for real-time dialogue and voice-first applications, and its model page credits it with acoustic nuance detection, numeric precision, and multimodal awareness.
- Input
- text image audio video $0.75/M
- Output
- text audio $4.5/M
- Audio input
- $3/M
- Audio output
- $12/M
Price in context
Where the price sits among 6 comparable models
The bar shows how this model’s price compares with every other model of the same kind on Synthorai. The cheapest and the most expensive are named at each end. These are base rates; batch, region and cache-write discounts are on the pricing page.
Specs & limits
Tokens
| Context window (vendor spec) | 131,072 |
|---|---|
| Max output (vendor spec) | 65,536 |
Thinking
| Vendor control | thinkingLevel |
|---|---|
| Accepted values | minimal · low · medium · high |
| Default | minimal applied when the request sets nothing |
| Can be turned off | No |
| Thinking behaviour | Defaults to minimal to optimize for lowest latency; the Gemini 3 Live models replaced the 2.5 Live thinkingBudget with a level. |
| Parameter | reasoning_effort |
| Values | minimal · low · medium · high the gateway's parameter surface - the vendor mapping above applies |
Audio
| Languages | More than 90 languages for real-time multimodal conversation |
|---|---|
| Audio limits |
|
| Streaming transcription | Yes |
Realtime
| Voices | Any voice from the Gemini text-to-speech voice set; native-audio output models switch languages naturally during a conversation. |
|---|---|
| Session |
|
Model
| Modalities | text + image + audio + video → text + audio |
|---|
Low-latency audio-to-audio model for voice-first agents, documented for acoustic nuance detection, numeric precision and multimodal awareness. Uses thinkingLevel (minimal / low / medium / high) instead of thinkingBudget, defaulting to minimal for lowest latency.
Use gemini-3.1-flash-live in 30 seconds
OpenAI Realtime-compatible: connect over WebSocket and stream audio in, audio out. WS /v1/realtime
import asyncio, base64, json, websockets
URL = "wss://synthorai.io/v1/realtime?model=gemini-3.1-flash-live"
# Send ONLY the Authorization header - the beta protocol is retired.
HEADERS = {"Authorization": "Bearer sk-syn-..."}
async def main():
async with websockets.connect(URL, additional_headers=HEADERS) as ws:
# 1) configure the speech-to-speech session
await ws.send(json.dumps({
"type": "session.update",
"session": {
"type": "realtime",
"output_modalities": ["audio"],
"audio": {"output": {"voice": "alloy"}},
},
}))
# 2) send input audio (base64 PCM16), then request a spoken reply
await ws.send(json.dumps({"type": "input_audio_buffer.append", "audio": pcm16_b64}))
await ws.send(json.dumps({"type": "input_audio_buffer.commit"}))
await ws.send(json.dumps({"type": "response.create"}))
# 3) stream the model's audio (and text) back
async for raw in ws:
ev = json.loads(raw)
if ev["type"] == "response.audio.delta":
play(base64.b64decode(ev["delta"])) # audio out
elif ev["type"] == "response.done":
break
asyncio.run(main())import WebSocket from "ws";
const ws = new WebSocket("wss://synthorai.io/v1/realtime?model=gemini-3.1-flash-live", {
// Send ONLY the Authorization header — the beta protocol is retired.
headers: { Authorization: "Bearer sk-syn-..." },
});
ws.on("open", () => {
// configure the speech-to-speech session
ws.send(JSON.stringify({ type: "session.update", session: {
type: "realtime", output_modalities: ["audio"], audio: { output: { voice: "alloy" } },
} }));
// send input audio (base64 PCM16), then request a spoken reply
ws.send(JSON.stringify({ type: "input_audio_buffer.append", audio: pcm16Base64 }));
ws.send(JSON.stringify({ type: "input_audio_buffer.commit" }));
ws.send(JSON.stringify({ type: "response.create" }));
});
ws.on("message", (raw) => {
const ev = JSON.parse(raw.toString());
if (ev.type === "response.audio.delta") playAudio(Buffer.from(ev.delta, "base64")); // audio out
else if (ev.type === "response.done") ws.close();
});# Realtime is a WebSocket protocol - use a WS client such as websocat.
# Each line below is one OpenAI Realtime event (JSON) sent to the session.
websocat -H 'Authorization: Bearer sk-syn-...' \
'wss://synthorai.io/v1/realtime?model=gemini-3.1-flash-live' <<'EOF'
{"type":"session.update","session":{"type":"realtime","output_modalities":["audio"],"audio":{"output":{"voice":"alloy"}}}}
{"type":"input_audio_buffer.append","audio":"<base64-pcm16>"}
{"type":"input_audio_buffer.commit"}
{"type":"response.create"}
EOF
# Responses stream back as response.audio.delta (base64 audio out) … response.donepackage main
import (
"net/http"
"github.com/gorilla/websocket"
)
func main() {
h := http.Header{}
h.Set("Authorization", "Bearer sk-syn-...")
// Send ONLY the Authorization header — the beta protocol is retired.
c, _, err := websocket.DefaultDialer.Dial("wss://synthorai.io/v1/realtime?model=gemini-3.1-flash-live", h)
if err != nil {
panic(err)
}
defer c.Close()
// configure the speech-to-speech session, send audio, request a spoken reply
c.WriteJSON(map[string]any{"type": "session.update", "session": map[string]any{
"type": "realtime", "output_modalities": []string{"audio"},
"audio": map[string]any{"output": map[string]any{"voice": "alloy"}}}})
c.WriteJSON(map[string]any{"type": "input_audio_buffer.append", "audio": pcm16B64})
c.WriteJSON(map[string]any{"type": "input_audio_buffer.commit"})
c.WriteJSON(map[string]any{"type": "response.create"})
for {
var ev struct {
Type string `json:"type"`
Delta string `json:"delta"`
}
if err := c.ReadJSON(&ev); err != nil {
return
}
if ev.Type == "response.audio.delta" {
playAudio(ev.Delta) // base64 audio out
} else if ev.Type == "response.done" {
return
}
}
}import java.net.URI;
import java.net.http.HttpClient;
import java.net.http.WebSocket;
import java.util.concurrent.CompletionStage;
// JDK built-in WebSocket — no extra dependency needed.
WebSocket ws = HttpClient.newHttpClient().newWebSocketBuilder()
.header("Authorization", "Bearer sk-syn-...")
// Send ONLY the Authorization header — the beta protocol is retired.
.buildAsync(URI.create("wss://synthorai.io/v1/realtime?model=gemini-3.1-flash-live"), new WebSocket.Listener() {
public CompletionStage<?> onText(WebSocket w, CharSequence data, boolean last) {
// handle response.audio.delta (base64 audio out) / response.done here
w.request(1);
return null;
}
}).join();
// configure the session, send input audio, then request a spoken reply
ws.sendText("{\"type\":\"session.update\",\"session\":{\"type\":\"realtime\",\"output_modalities\":[\"audio\"],\"audio\":{\"output\":{\"voice\":\"alloy\"}}}}", true);
ws.sendText("{\"type\":\"input_audio_buffer.append\",\"audio\":\"<base64-pcm16>\"}", true);
ws.sendText("{\"type\":\"input_audio_buffer.commit\"}", true);
ws.sendText("{\"type\":\"response.create\"}", true);About gemini-3.1-flash-live
- It accepts text, images, audio, and video and returns text and audio over the Live API's stateful WebSocket, with a 131,072-token input limit and 65,536-token output limit, eight times the output ceiling of the 2.5 native-audio model it replaces.
- Function calling, search grounding, and thinking are supported, while caching, structured outputs, code execution, and the Batch API are not, and thinking depth is set with thinkingLevel at minimal, low, medium, or high, defaulting to minimal for the lowest latency instead of the numeric thinkingBudget the 2.5 generation used.
- The migration notes are the useful part, because this upgrade removes as well as adds.
- Asynchronous function calling is not yet supported, so calls are sequential and the model will not start responding until the tool response arrives; proactive audio and affective dialog are gone and their configuration has to be stripped out.
- A single server event can now carry several content parts at once, such as an audio chunk and a transcript together, so clients that read only the first part will silently drop content.
- Turn coverage now defaults to including detected audio activity and all video frames, which can raise costs for applications streaming video continuously, and client content may only seed initial history.
- Synthorai serves it over the OpenAI-compatible /v1/realtime WebSocket endpoint.
FAQ
Is the gemini-3.1-flash-live API free to try?
gemini-3.1-flash-live is currently in invited beta: access is application-based rather than open signup. Apply from the Synthorai console; once approved, standard pay-as-you-go pricing applies with no subscription.
What is gemini-3.1-flash-live best at?
Audio-to-audio model for real-time dialogue and voice-first applications; eight times the output ceiling of the 2.5 native-audio model; no asynchronous function calling, so tool calls are sequential. See the About section for the full picture from the vendor's own release notes.
How much does gemini-3.1-flash-live cost?
gemini-3.1-flash-live costs $0.75 per million input tokens and $4.5 per million output tokens on Synthorai. That is the provider's list price, with no platform markup.
How do I connect to the gemini-3.1-flash-live API?
gemini-3.1-flash-live is a speech-to-speech model: connect over WebSocket to wss://synthorai.io/v1/realtime?model=gemini-3.1-flash-live with the OpenAI Realtime SDK (or a raw WebSocket) and stream audio in, audio out. It is not a POST /v1/audio/transcriptions file upload. Authenticate with your sk-syn key, sending only the Authorization header (the beta protocol is retired).
How do I get access to gemini-3.1-flash-live?
gemini-3.1-flash-live is in invited beta: request access from the Synthorai console. Once approved it works like every other model: point your OpenAI SDK at base_url="https://synthorai.io/v1" and set model="gemini-3.1-flash-live".
Related models
Compare
Every value on this page is transcribed from the vendor's own documentation, linked above, and carries the date it was checked. Prices are compared across the catalogue; specification values that vendors define differently are shown with the difference stated rather than charted. Nothing here is measured by us, and nothing is scored.