Gemini 3.7 Flash 是谷歌能力最強的 Flash 檔模型,2026 年 8 月 13 日正式發布,定位於智慧體工作流與多模態推理。
- 輸入
- 文字 影像 音訊 影片 $0.75/M
- 輸出
- 文字 $3.75/M
- 音訊輸入
- $0.75/M
- 快取讀取
- $0.075/M
- 上下文
- 1M
- 相較 GPT-4o
- 便宜約 85%
- 知識截止
- 2026-03
Benchmark 成績
廠商公布: Alibaba (Qwen) Anthropic DeepSeek Google Moonshot OpenAI Tencent Z.ai
價格在同類中的位置
價格在 68 個同類模型中的位置
這條線顯示該模型的價格,在 Synthorai 上同類模型裡處於什麼位置。兩端標出了最便宜和最貴的那個。這裡是基礎價,批次、區域與快取寫入的折扣見價格頁。
規格與限制
Token
| 上下文視窗(廠商規格) | 1,048,576 |
|---|---|
| 最大輸出(廠商規格) | 65,536 |
| 知識截止 | 2026-03 |
提示詞快取
| 快取方式 | 自動 + 顯式 |
|---|---|
| 最低前綴 | 4,096 |
思考
| 參數 | reasoning_effort |
|---|---|
| 取值 | minimal · low · medium · high 具體接受哪些由廠商決定 |
模型
| 模態 | 文字 + 影像 + 音訊 + 影片 → 文字 |
|---|
- Google's most capable Flash model for agentic workflows and multimodal reasoning
- output price includes thinking tokens, and audio and video input bill at the text rate
單一提示詞 — 經由閘道測量
Gemini 3.7 Flash 通過 · 3 sentences
Enabling reasoning causes the model to generate a high volume of invisible "thinking tokens" behind the scenes to deliberate and solve complex logic before delivering a final response. Because our vendor bills strictly on token volume, this additional internal computation effectively doubled our billable output per query. In short, our costs increased because we are now paying for the extra compute capacity required to produce significantly higher-accuracy answers.
輸出 878 tok (+799 思考) 延遲 6.9 s
指令遵循(恰好三句,可數)、受眾適配(面向 CFO 的語氣),以及下方 token 計量所暴露的隱藏思考計費缺口。
Gemini 3.7 Flash 通過 · 8/8 cases
Here is the corrected function: ```python def longest_run(xs): if not xs: return 0 best = 1 cur = 1 for i in range(1, len(xs)): if xs[i] == xs[i - 1]: cur += 1 else: cur = 1 best = max(best, cur) return best ``` ### Explanation of the Bug The original function tracked equal adjacent transitions instead of the actual number of elements in a run, causing all non-empty run lengths to be undercounted by one (e.g., returning `0` instead of `1` for single-element lists because the loop never executes, and `1` instead of `2` for pairs). Initializing both `best` and `cur` to `1` (after handling empty inputs) and resetting `cur` to `1` on a mismatch ensures every individual element is counted as a valid run of at least length one.
輸出 1050 tok (+827 思考) 延遲 6.4 s
修復是否真的正確(可執行)、解釋的資訊密度,以及在一個邊界明確的任務上的 token 效率。
Gemini 3.7 Flash 通過 · 5/5 fields, guidance "withheld"
```json { "company": "Meridian Robotics", "quarter": "September quarter", "revenue_usd_m": 412, "growth_pct": null, "guidance": "withheld" } ```
輸出 918 tok (+858 思考) 延遲 6.2 s
schema 服從度(不臆造欄位)、幻覺壓力(guidance 明確被暫緩給出),以及結構化輸出路徑的差異。
Gemini 3.7 Flash 通過 · 120 words, 0 banned, 1 question
Why pay twice for identical AI queries simply because you routed them to different model vendors? Today, we introduce Universal Prompt Caching directly within our unified API gateway architecture. This capability stores repeated prompt contexts across OpenAI, Anthropic, and local models, instantly returning stored results to eliminate redundant computation fees. When your application sends an LLM request, the gateway inspects the payload, identifies semantic matches, and returns accurate cached responses in under ten milliseconds. Engineering teams can now slash inference latency by eighty percent while dramatically reducing monthly token expenditures across diverse production deployments. You retain complete privacy control, flexible cache eviction policies, and granular metrics through a single dashboard. Update your routing settings today to accelerate overall system performance.
輸出 2858 tok (+2718 思考) 延遲 14.1 s
約束服從度(字數預算、禁用詞表、唯一的那句問句)、文風指紋,以及長度控制。
30 秒用上 Gemini 3.7 Flash
OpenAI 相容:換掉 base_url,SDK 不用改。POST /v1/chat/completions
from openai import OpenAI
client = OpenAI(
base_url="https://synthorai.io/v1",
api_key="sk-syn-...",
)
resp = client.chat.completions.create(
model="gemini-3.7-flash",
messages=[{"role": "user", "content": "Summarize this diff"}],
reasoning_effort="medium",
)
print(resp.choices[0].message.content)import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://synthorai.io/v1",
apiKey: "sk-syn-...",
});
const resp = await client.chat.completions.create({
model: "gemini-3.7-flash",
messages: [{ role: "user", content: "Summarize this diff" }],
reasoning_effort: "medium",
});
console.log(resp.choices[0].message.content);curl https://synthorai.io/v1/chat/completions \
-H "Authorization: Bearer sk-syn-..." \
-H "Content-Type: application/json" \
-d '{
"model": "gemini-3.7-flash",
"messages": [{"role": "user", "content": "Hello"}],
"reasoning_effort": "medium"
}'package main
import (
"context"
"fmt"
"github.com/openai/openai-go/v3"
"github.com/openai/openai-go/v3/option"
)
func main() {
client := openai.NewClient(
option.WithBaseURL("https://synthorai.io/v1"),
option.WithAPIKey("sk-syn-..."),
)
resp, _ := client.Chat.Completions.New(context.TODO(), openai.ChatCompletionNewParams{
Model: "gemini-3.7-flash",
Messages: []openai.ChatCompletionMessageParamUnion{
openai.UserMessage("Summarize this diff"),
},
ReasoningEffort: openai.ReasoningEffortMedium,
})
fmt.Println(resp.Choices[0].Message.Content)
}import com.openai.client.OpenAIClient;
import com.openai.client.okhttp.OpenAIOkHttpClient;
import com.openai.models.chat.completions.*;
import com.openai.models.ReasoningEffort;
OpenAIClient client = OpenAIOkHttpClient.builder()
.baseUrl("https://synthorai.io/v1")
.apiKey("sk-syn-...")
.build();
ChatCompletion resp = client.chat().completions().create(
ChatCompletionCreateParams.builder()
.model("gemini-3.7-flash")
.addUserMessage("Summarize this diff")
.reasoningEffort(ReasoningEffort.MEDIUM)
.build());
System.out.println(resp.choices().get(0).message().content().orElse(""));關於 Gemini 3.7 Flash
- 它接受文字、圖像、音訊和視訊輸入並回傳文字,因此一次呼叫就能同時對截圖、錄音和提示詞做推理,不必再把多個模型拼成流水線。
- 兩處定價細節讓它的預算特別好算。
- 其一,**沒有上下文長度分檔**:90 萬 token 的提示詞與 900 token 的單價相同,長上下文任務不會在文件讀到一半時價格跳檔。
- 其二,**沒有模態溢價**:音訊和視訊輸入按文字價計費,這與多數多模態定價的做法相反,也就免去了傳送前先降取樣的常規理由。
- 輸出價**含思考 token**,因此推理較重的回答按完整軌跡計費而非只按可見正文,設定最大輸出時要把這一點算進去。
- 上下文快取價約為輸入價的十分之一,另按小時收儲存費,所以對反覆重播同一段系統提示詞的智慧體迴圈,固定一個穩定前綴是划算的。
- Google 搜尋接地(grounding)每月有一份在 Gemini 3 家族內共享的免費額度,超出後按次計費。
- Synthorai 透過 OpenAI 相容的 chat completions 端點提供該模型。
常見問題
Gemini 3.7 Flash API 可以免費試用嗎?
可以,新帳號可獲得 10 次試用呼叫和最高 $1 的免費額度,無需信用卡。以輸入 $0.75/M 計算,光是這筆額度就足以對 Gemini 3.7 Flash 發出約 166 次 ~8K token 的請求。
Gemini 3.7 Flash 最擅長什麼?
Flash 檔中能力最強,面向智慧體、文字、圖像、音訊、視訊輸入、無上下文分檔、無模態溢價。完整能力請見「關於」一節,內容取自廠商官方發布說明。
Gemini 3.7 Flash 的價格是多少?
在 Synthorai 上,Gemini 3.7 Flash 輸入 $0.75/百萬 token、輸出 $3.75/百萬 token,即廠商牌價,無平台加價。快取命中的輸入 token 以 $0.075/M 計費。
Gemini 3.7 Flash 支援提示詞快取(prompt caching)嗎?
支援:自動快取預設開啟,另有顯式模式可獲得確定的折扣。快取命中的輸入 token 以 $0.075/M 計費(未命中 $0.75/M);提示詞需有 4,096 個 token 以上的穩定前綴才能命中快取。 提示詞快取指南 →
如何開通 Gemini 3.7 Flash?
把現有 OpenAI SDK 的 base_url 指向 "https://synthorai.io/v1",model 設為 "gemini-3.7-flash" 即可。一組 API key 通用閘道上的所有模型。
Gemini 3.7 Flash 的知識截止日期是什麼時候?
Gemini 3.7 Flash 的知識截止日期為 2026-03,依據廠商官方文件(資料核驗於 2026-08-15)。
相關模型
對比
本頁每個值都轉錄自廠商自己的文件(連結見上),並帶有核對日期。價格在全目錄範圍內比較;各廠商定義不同的規格值,只說明差異而不作圖表對比。此處沒有任何由我們測量的資料,也不做評分。