新規 無料登録で、呼び出し 10 回分(最大 $1)をお試しいただけます。カード登録は不要です。

gemini-2.5-flash-native-audio vs gemini-3.1-flash-live

gemini-2.5-flash-native-audio は招待制で提供しています。以下の数値は現在適用されている料金ですが、呼び出すにはワークスペースへの利用許可が必要です。この比較をもとに開発を始める前に、当社まで利用をお申し込みください。

gemini-3.1-flash-live は招待制で提供しています。以下の数値は現在適用されている料金ですが、呼び出すにはワークスペースへの利用許可が必要です。この比較をもとに開発を始める前に、当社まで利用をお申し込みください。

vs

どんなときに、どちらを選ぶか

音声の料金は同一(100万あたり入力 $3、出力 $12)で、セッションの作りも同じ — 16 kHz 入力/24 kHz 出力の PCM、128k コンテキスト、音声のみ 15 分の上限 — 古い gemini-2.5-flash-native-audio はテキストがわずかに安く、100万あたり $0.5 と $2 対 $0.75 と $4.5 です。機能上の分かれ目はツールの挙動で、gemini-2.5-flash-native-audio は NON_BLOCKING の関数宣言に対応し、ツール実行中もモデルが話し続けられます。gemini-3.1-flash-live は逐次で、ツールの結果が返るまで応答しません — ただしテキスト・音声・映像に加えて画像入力を備えます。ツールを多用する会話なら 2.5、画像入力なら 3.1 を選んでください。

ベンチマーク

gemini-3.1-flash-live:ベンダーがベンチマークスコアを公表していません。

gemini-2.5-flash-native-audio:公表値は 8 件ありますが、比較できるだけの数のモデルが測定している項目がありません。

gemini-2.5-flash-native-audio gemini-3.1-flash-live 測定された他モデル 他モデルの平均 ★ これを上回るモデルなし
AMI near+far field (WER)
63.8%
N/A
Big Bench Audio
71%
N/A
BFCL v3 (single-turn, python only)
69.4%
N/A

ベンダー公表値: Amazon OpenAI

料金

gemini-2.5-flash-native-audio gemini-3.1-flash-live Δ
音声入力 / 100 万トークン $3 $3 =
音声出力 / 100 万トークン $12 $12 =
テキスト入力 / 100 万トークン $0.5 $0.75 0.67×
テキスト出力 / 100 万トークン $2 $4.5 0.44×

ビルド時点の現行カタログの料金です。最新の料金表は各モデルのページに掲載しています。

両モデルの位置:音声トークン 100 万あたりの料金(この課金単位のリアルタイム音声対話モデル全 6 件、対数スケール)

対応機能

gemini-2.5-flash-native-audio gemini-3.1-flash-live
思考の制御 設定可能 常時オン
プロンプトキャッシュ 自動 + 明示 自動 + 明示
キャッシュ保持時間 公表なし 公表なし
キャッシュ対象の最小プレフィックス 4096 トークン 4096 トークン

仕様

gemini-2.5-flash-native-audio gemini-3.1-flash-live
入力モダリティ テキスト 音声 動画 テキスト 画像 音声 動画
出力モダリティ テキスト 音声 テキスト 音声
リリース - 2026-03-26
知識カットオフ 2025-01 -
ボイス

Any voice from the Gemini text-to-speech voice set

native-audio output models switch languages naturally during a conversation.

Any voice from the Gemini text-to-speech voice set

native-audio output models switch languages naturally during a conversation.

セッション機能
  • Function calling supported, with NON_BLOCKING function declarations letting the model keep talking while a tool runs
  • VAD interruption cancels and discards the in-flight generation
  • search grounding and thinking supported
  • no caching, structured outputs, code execution or Batch API
  • Function calling is sequential only - the model will not start responding until the tool response is sent
  • search grounding and thinking supported
  • VAD interruption cancels the in-flight generation
  • no caching, structured outputs, code execution or Batch API
コンテキストウィンドウ 131K 131K

仕様は各ベンダーのドキュメントから転記しています。ベンダーが公表していない項目は、推測で埋めずに省いています。 出典の一覧: gemini-2.5-flash-native-audio · gemini-3.1-flash-live

1 行の変更で切り替え

以下のどのタブにも両方のモデル ID が入っています。書き換えるのはハイライトされた 2 行だけで、エンドポイント、キー、リクエスト形式はすべて同じです。

import asyncio, base64, json, websockets

URL = "wss://synthorai.io/v1/realtime?model=gemini-2.5-flash-native-audio"
# URL = "wss://synthorai.io/v1/realtime?model=gemini-3.1-flash-live"  # この行のコメントを外し、上の行をコメントアウト
# Send ONLY the Authorization header - the beta protocol is retired.
HEADERS = {"Authorization": "Bearer sk-syn-..."}

async def main():
    async with websockets.connect(URL, additional_headers=HEADERS) as ws:
        # 1) configure the speech-to-speech session
        await ws.send(json.dumps({
            "type": "session.update",
            "session": {
                "type": "realtime",
                "output_modalities": ["audio"],
                "audio": {"output": {"voice": "alloy"}},
            },
        }))
        # 2) send input audio (base64 PCM16), then request a spoken reply
        await ws.send(json.dumps({"type": "input_audio_buffer.append", "audio": pcm16_b64}))
        await ws.send(json.dumps({"type": "input_audio_buffer.commit"}))
        await ws.send(json.dumps({"type": "response.create"}))
        # 3) stream the model's audio (and text) back
        async for raw in ws:
            ev = json.loads(raw)
            if ev["type"] == "response.audio.delta":
                play(base64.b64decode(ev["delta"]))   # audio out
            elif ev["type"] == "response.done":
                break

asyncio.run(main())

API キーを取得 →

よくある質問

gemini-2.5-flash-native-audio と gemini-3.1-flash-live ではどちらが安いですか?

音声入力 / 100 万トークンはどちらも同じ($3)なので、価格は決め手になりません。以下の仕様と対応機能を見比べてください。

gemini-2.5-flash-native-audio と gemini-3.1-flash-live の A/B テストは、実装を 2 つ用意せずにできますか?

はい。どちらも同じ OpenAI 互換エンドポイントから 1 つの API キーで利用できます。切り替えはモデル名の文字列を 1 行変えるだけなので、トラフィックの一部をそれぞれに振り分けて、請求額を直接比較できます。

gemini-2.5-flash-native-audio と gemini-3.1-flash-live はプロンプトキャッシュに対応していますか?

当社の価格データでキャッシュ読み取りの料金が設定されているのは、2 つのうち 1 つだけです。料金の記載がないモデルでは、プロバイダーがキャッシュ読み取りに個別の価格を設定していません。

関連する比較