Gemini 3.7 Flash は Google の Flash 系で最も高性能なモデルで、2026 年 8 月 13 日に一般提供が始まり、エージェント的なワークフローとマルチモーダル推論に位置づけられています。
- 入力
- テキスト 画像 音声 動画 $0.75/M
- 出力
- テキスト $3.75/M
- 音声入力
- $0.75/M
- キャッシュ読み取り
- $0.075/M
- コンテキスト
- 1M
- GPT-4o 比
- 約 85% 割安
- 知識カットオフ
- 2026-03
ベンチマーク
ベンダー公表: Alibaba (Qwen) Anthropic DeepSeek Google Moonshot OpenAI Tencent Z.ai
価格の位置づけ
同種 68 モデル中の料金の位置
このバーは、Synthorai 上の同種モデルの中でこのモデルの価格がどこに位置するかを示します。両端には最も安いモデルと最も高いモデルの名前が入ります。表示は基本料金で、バッチ・リージョン・キャッシュ書き込みの割引は料金ページにあります。
スペックと制限
トークン
| コンテキストウィンドウ(ベンダー仕様) | 1,048,576 |
|---|---|
| 最大出力(ベンダー仕様) | 65,536 |
| 知識カットオフ | 2026-03 |
プロンプトキャッシュ
| キャッシュ方式 | 自動 + 明示 |
|---|---|
| 最小プレフィックス | 4,096 |
思考
| パラメータ | reasoning_effort |
|---|---|
| 値 | minimal · low · medium · high 受け付ける範囲はプロバイダー次第 |
モデル
| モダリティ | テキスト + 画像 + 音声 + 動画 → テキスト |
|---|
- Google's most capable Flash model for agentic workflows and multimodal reasoning
- output price includes thinking tokens, and audio and video input bill at the text rate
1つのプロンプト — ゲートウェイ経由で測定
Gemini 3.7 Flash パス · 3 sentences
Enabling reasoning causes the model to generate a high volume of invisible "thinking tokens" behind the scenes to deliberate and solve complex logic before delivering a final response. Because our vendor bills strictly on token volume, this additional internal computation effectively doubled our billable output per query. In short, our costs increased because we are now paying for the extra compute capacity required to produce significantly higher-accuracy answers.
出力 878 tok (+799 思考) レイテンシ 6.9 s
指示追従(ちょうど三文 — 数えられる)、読み手に合わせた調整(CFO 向けの語り口)、そして下のトークン計から見える隠れた思考の課金ギャップ。
Gemini 3.7 Flash パス · 8/8 cases
Here is the corrected function: ```python def longest_run(xs): if not xs: return 0 best = 1 cur = 1 for i in range(1, len(xs)): if xs[i] == xs[i - 1]: cur += 1 else: cur = 1 best = max(best, cur) return best ``` ### Explanation of the Bug The original function tracked equal adjacent transitions instead of the actual number of elements in a run, causing all non-empty run lengths to be undercounted by one (e.g., returning `0` instead of `1` for single-element lists because the loop never executes, and `1` instead of `2` for pairs). Initializing both `best` and `cur` to `1` (after handling empty inputs) and resetting `cur` to `1` on a mismatch ensures every individual element is counted as a valid run of at least length one.
出力 1050 tok (+827 思考) レイテンシ 6.4 s
修正が実際に正しいか(実行可能か)、説明の密度、そして範囲の限られたタスクでのトークン効率。
Gemini 3.7 Flash パス · 5/5 fields, guidance "withheld"
```json { "company": "Meridian Robotics", "quarter": "September quarter", "revenue_usd_m": 412, "growth_pct": null, "guidance": "withheld" } ```
出力 918 tok (+858 思考) レイテンシ 6.2 s
スキーマ遵守(フィールドをでっち上げない)、ハルシネーション圧(guidance は明示的に保留されている)、そして構造化出力の経路の違い。
Gemini 3.7 Flash パス · 120 words, 0 banned, 1 question
Why pay twice for identical AI queries simply because you routed them to different model vendors? Today, we introduce Universal Prompt Caching directly within our unified API gateway architecture. This capability stores repeated prompt contexts across OpenAI, Anthropic, and local models, instantly returning stored results to eliminate redundant computation fees. When your application sends an LLM request, the gateway inspects the payload, identifies semantic matches, and returns accurate cached responses in under ten milliseconds. Engineering teams can now slash inference latency by eighty percent while dramatically reducing monthly token expenditures across diverse production deployments. You retain complete privacy control, flexible cache eviction policies, and granular metrics through a single dashboard. Update your routing settings today to accelerate overall system performance.
出力 2858 tok (+2718 思考) レイテンシ 14.1 s
制約の遵守(語数の上限、禁止語リスト、唯一の疑問文)、文体の指紋、そして長さの制御。
30 秒で Gemini 3.7 Flash を使う
OpenAI 互換。base_url を差し替えるだけで、SDK はそのまま。POST /v1/chat/completions
from openai import OpenAI
client = OpenAI(
base_url="https://synthorai.io/v1",
api_key="sk-syn-...",
)
resp = client.chat.completions.create(
model="gemini-3.7-flash",
messages=[{"role": "user", "content": "Summarize this diff"}],
reasoning_effort="medium",
)
print(resp.choices[0].message.content)import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://synthorai.io/v1",
apiKey: "sk-syn-...",
});
const resp = await client.chat.completions.create({
model: "gemini-3.7-flash",
messages: [{ role: "user", content: "Summarize this diff" }],
reasoning_effort: "medium",
});
console.log(resp.choices[0].message.content);curl https://synthorai.io/v1/chat/completions \
-H "Authorization: Bearer sk-syn-..." \
-H "Content-Type: application/json" \
-d '{
"model": "gemini-3.7-flash",
"messages": [{"role": "user", "content": "Hello"}],
"reasoning_effort": "medium"
}'package main
import (
"context"
"fmt"
"github.com/openai/openai-go/v3"
"github.com/openai/openai-go/v3/option"
)
func main() {
client := openai.NewClient(
option.WithBaseURL("https://synthorai.io/v1"),
option.WithAPIKey("sk-syn-..."),
)
resp, _ := client.Chat.Completions.New(context.TODO(), openai.ChatCompletionNewParams{
Model: "gemini-3.7-flash",
Messages: []openai.ChatCompletionMessageParamUnion{
openai.UserMessage("Summarize this diff"),
},
ReasoningEffort: openai.ReasoningEffortMedium,
})
fmt.Println(resp.Choices[0].Message.Content)
}import com.openai.client.OpenAIClient;
import com.openai.client.okhttp.OpenAIOkHttpClient;
import com.openai.models.chat.completions.*;
import com.openai.models.ReasoningEffort;
OpenAIClient client = OpenAIOkHttpClient.builder()
.baseUrl("https://synthorai.io/v1")
.apiKey("sk-syn-...")
.build();
ChatCompletion resp = client.chat().completions().create(
ChatCompletionCreateParams.builder()
.model("gemini-3.7-flash")
.addUserMessage("Summarize this diff")
.reasoningEffort(ReasoningEffort.MEDIUM)
.build());
System.out.println(resp.choices().get(0).message().content().orElse(""));Gemini 3.7 Flash について
- テキスト、画像、音声、動画を受け取ってテキストを返すため、スクリーンショットと録音とプロンプトを 1 回の呼び出しでまとめて推論でき、複数モデルをパイプラインにつなぐ必要がありません。
- 料金面の特徴が二つあり、見積もりが立てやすくなっています。
- 第一に、**コンテキスト長による段階制がありません**。
- 90 万トークンのプロンプトでも 900 トークンでも単価は同じで、長文処理の途中で価格が跳ね上がりません。
- 第二に、**モダリティによる追加料金もありません**。
- 音声と動画の入力はテキストと同じ単価で課金され、これは多くのマルチモーダル料金体系とは逆で、送信前にメディアをダウンサンプリングする理由がなくなります。
- 出力料金には**思考トークンが含まれる**ため、推論の重い回答は可視部分だけでなくトレース全体で課金されます。
- 最大出力を設定する際はこの点を織り込んでください。
- コンテキストキャッシュは入力単価のおよそ 10 分の 1 で、別途時間単位の保管料がかかるため、同じシステムプロンプトを繰り返すエージェントのループでは安定したプレフィックスを固定する価値があります。
- Google 検索によるグラウンディングは Gemini 3 ファミリー内で共有される月次の無料枠があり、超過分から従量課金になります。
- Synthorai は OpenAI 互換の chat completions エンドポイント経由で提供します。
よくある質問
Gemini 3.7 Flash API は無料で試せますか?
はい。新規アカウントには 10 回のトライアル呼び出しと最大 $1 の無料クレジットが付与され、カード登録は不要です。入力 $0.75/M で計算すると、このクレジットだけで Gemini 3.7 Flash に対して約 166 回の ~8K トークンのリクエストを送れます。
Gemini 3.7 Flash は何が得意ですか?
Flash 系で最も高性能、エージェント向け、テキスト・画像・音声・動画入力、コンテキスト段階制もモダリティ追加料金もなし。全体像はベンダー公式のリリースノートに基づく「このモデルについて」セクションをご覧ください。
Gemini 3.7 Flash の料金はいくらですか?
Synthorai 上の Gemini 3.7 Flash は入力 100 万トークンあたり $0.75、出力 100 万トークンあたり $3.75 です。ベンダー定価のままで、プラットフォーム手数料はありません。キャッシュ済み入力トークンは $0.075/M で課金されます。
Gemini 3.7 Flash はプロンプトキャッシュに対応していますか?
はい。自動キャッシュがデフォルトで有効なうえ、確実な割引を得られる明示モードもあります。キャッシュ済み入力トークンは $0.075/M(未キャッシュは $0.75/M)で課金されます。なお、キャッシュには 4,096 トークン以上の安定したプレフィックスが必要です。 プロンプトキャッシュガイド →
Gemini 3.7 Flash を利用するには?
お使いの OpenAI SDK の base_url を "https://synthorai.io/v1" に向け、model="gemini-3.7-flash" を設定すれば完了です。API キー 1 本でゲートウェイ上のすべてのモデルを利用できます。
Gemini 3.7 Flash の知識カットオフはいつですか?
ベンダー公式ドキュメントによると、Gemini 3.7 Flash の知識カットオフは 2026-03 です(2026-08-15 時点)。
関連モデル
比較
このページの値はすべてベンダー自身のドキュメント(上部にリンク)から転記し、確認した日付を付しています。価格はカタログ全体で比較しますが、ベンダーごとに定義が異なる仕様値は差異を明記するにとどめ、図表で比較はしません。当社が測定した数値はなく、スコアも付けていません。