Gemini 3.7 Flash 是谷歌能力最强的 Flash 档模型,2026 年 8 月 13 日正式发布,定位于智能体工作流与多模态推理。
- 输入
- 文本 图像 音频 视频 $0.75/M
- 输出
- 文本 $3.75/M
- 音频输入
- $0.75/M
- 缓存读取
- $0.075/M
- 上下文
- 1M
- 对比 GPT-4o
- 便宜约 85%
- 知识截止
- 2026-03
Benchmark 成绩
厂商公布: Alibaba (Qwen) Anthropic DeepSeek Google Moonshot OpenAI Tencent Z.ai
价格在同类中的位置
价格在 68 个同类模型中的位置
这条线显示该模型的价格,在 Synthorai 上同类模型里处于什么位置。两端标出了最便宜和最贵的那个。这里是基础价,批量、区域和缓存写入的折扣见价格页。
规格与限制
Token
| 上下文窗口(厂商规格) | 1,048,576 |
|---|---|
| 最大输出(厂商规格) | 65,536 |
| 知识截止 | 2026-03 |
提示词缓存
| 缓存方式 | 自动 + 显式 |
|---|---|
| 最低前缀 | 4,096 |
思考
| 参数 | reasoning_effort |
|---|---|
| 取值 | minimal · low · medium · high 具体接受哪些由厂商决定 |
模型
| 模态 | 文本 + 图像 + 音频 + 视频 → 文本 |
|---|
- Google's most capable Flash model for agentic workflows and multimodal reasoning
- output price includes thinking tokens, and audio and video input bill at the text rate
单个 Prompt —— 通过网关实测
Gemini 3.7 Flash 通过 · 3 sentences
Enabling reasoning causes the model to generate a high volume of invisible "thinking tokens" behind the scenes to deliberate and solve complex logic before delivering a final response. Because our vendor bills strictly on token volume, this additional internal computation effectively doubled our billable output per query. In short, our costs increased because we are now paying for the extra compute capacity required to produce significantly higher-accuracy answers.
输出 878 tok (+799 思考) 延迟 6.9 s
指令遵循(恰好三句,可数)、受众适配(面向 CFO 的语气),以及下方 token 计量所暴露的隐藏思考计费缺口。
Gemini 3.7 Flash 通过 · 8/8 cases
Here is the corrected function: ```python def longest_run(xs): if not xs: return 0 best = 1 cur = 1 for i in range(1, len(xs)): if xs[i] == xs[i - 1]: cur += 1 else: cur = 1 best = max(best, cur) return best ``` ### Explanation of the Bug The original function tracked equal adjacent transitions instead of the actual number of elements in a run, causing all non-empty run lengths to be undercounted by one (e.g., returning `0` instead of `1` for single-element lists because the loop never executes, and `1` instead of `2` for pairs). Initializing both `best` and `cur` to `1` (after handling empty inputs) and resetting `cur` to `1` on a mismatch ensures every individual element is counted as a valid run of at least length one.
输出 1050 tok (+827 思考) 延迟 6.4 s
修复是否真的正确(可运行)、解释的信息密度,以及在一个边界明确的任务上的 token 效率。
Gemini 3.7 Flash 通过 · 5/5 fields, guidance "withheld"
```json { "company": "Meridian Robotics", "quarter": "September quarter", "revenue_usd_m": 412, "growth_pct": null, "guidance": "withheld" } ```
输出 918 tok (+858 思考) 延迟 6.2 s
schema 服从度(不臆造字段)、幻觉压力(guidance 明确被暂缓给出),以及结构化输出路径的差异。
Gemini 3.7 Flash 通过 · 120 words, 0 banned, 1 question
Why pay twice for identical AI queries simply because you routed them to different model vendors? Today, we introduce Universal Prompt Caching directly within our unified API gateway architecture. This capability stores repeated prompt contexts across OpenAI, Anthropic, and local models, instantly returning stored results to eliminate redundant computation fees. When your application sends an LLM request, the gateway inspects the payload, identifies semantic matches, and returns accurate cached responses in under ten milliseconds. Engineering teams can now slash inference latency by eighty percent while dramatically reducing monthly token expenditures across diverse production deployments. You retain complete privacy control, flexible cache eviction policies, and granular metrics through a single dashboard. Update your routing settings today to accelerate overall system performance.
输出 2858 tok (+2718 思考) 延迟 14.1 s
约束服从度(字数预算、禁用词表、唯一的那句问句)、文风指纹,以及长度控制。
30 秒用上 Gemini 3.7 Flash
OpenAI 兼容:换掉 base_url,SDK 不用改。POST /v1/chat/completions
from openai import OpenAI
client = OpenAI(
base_url="https://synthorai.io/v1",
api_key="sk-syn-...",
)
resp = client.chat.completions.create(
model="gemini-3.7-flash",
messages=[{"role": "user", "content": "Summarize this diff"}],
reasoning_effort="medium",
)
print(resp.choices[0].message.content)import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://synthorai.io/v1",
apiKey: "sk-syn-...",
});
const resp = await client.chat.completions.create({
model: "gemini-3.7-flash",
messages: [{ role: "user", content: "Summarize this diff" }],
reasoning_effort: "medium",
});
console.log(resp.choices[0].message.content);curl https://synthorai.io/v1/chat/completions \
-H "Authorization: Bearer sk-syn-..." \
-H "Content-Type: application/json" \
-d '{
"model": "gemini-3.7-flash",
"messages": [{"role": "user", "content": "Hello"}],
"reasoning_effort": "medium"
}'package main
import (
"context"
"fmt"
"github.com/openai/openai-go/v3"
"github.com/openai/openai-go/v3/option"
)
func main() {
client := openai.NewClient(
option.WithBaseURL("https://synthorai.io/v1"),
option.WithAPIKey("sk-syn-..."),
)
resp, _ := client.Chat.Completions.New(context.TODO(), openai.ChatCompletionNewParams{
Model: "gemini-3.7-flash",
Messages: []openai.ChatCompletionMessageParamUnion{
openai.UserMessage("Summarize this diff"),
},
ReasoningEffort: openai.ReasoningEffortMedium,
})
fmt.Println(resp.Choices[0].Message.Content)
}import com.openai.client.OpenAIClient;
import com.openai.client.okhttp.OpenAIOkHttpClient;
import com.openai.models.chat.completions.*;
import com.openai.models.ReasoningEffort;
OpenAIClient client = OpenAIOkHttpClient.builder()
.baseUrl("https://synthorai.io/v1")
.apiKey("sk-syn-...")
.build();
ChatCompletion resp = client.chat().completions().create(
ChatCompletionCreateParams.builder()
.model("gemini-3.7-flash")
.addUserMessage("Summarize this diff")
.reasoningEffort(ReasoningEffort.MEDIUM)
.build());
System.out.println(resp.choices().get(0).message().content().orElse(""));关于 Gemini 3.7 Flash
- 它接受文本、图像、音频和视频输入并返回文本,因此一次调用就能同时对截图、录音和提示词做推理,不必再把多个模型拼成流水线。
- 两处定价细节让它的预算特别好算。
- 其一,**没有上下文长度分档**:90 万 token 的提示词与 900 token 的单价相同,长上下文任务不会在文档读到一半时价格跳档。
- 其二,**没有模态溢价**:音频和视频输入按文本价计费,这与多数多模态定价的做法相反,也就免去了发送前先降采样的常规理由。
- 输出价**含思考 token**,因此推理较重的回答按完整轨迹计费而非只按可见正文,设置最大输出时要把这一点算进去。
- 上下文缓存价约为输入价的十分之一,另按小时收存储费,所以对反复重放同一段系统提示词的智能体循环,固定一个稳定前缀是划算的。
- Google 搜索接地(grounding)每月有一份在 Gemini 3 家族内共享的免费额度,超出后按次计费。
- Synthorai 通过 OpenAI 兼容的 chat completions 端点提供该模型,接进既有集成无需换用谷歌专用客户端。
常见问题
Gemini 3.7 Flash API 可以免费试用吗?
可以。新账号可获得 10 次试用调用和最高 $1 的免费额度,无需绑卡。按输入 $0.75/M 计算,仅这笔额度就足够对 Gemini 3.7 Flash 发起约 166 次 ~8K token 的请求。
Gemini 3.7 Flash 最擅长什么?
Flash 档中能力最强,面向智能体、文本、图像、音频、视频输入、无上下文分档、无模态溢价。完整能力请见「关于」部分,内容取自厂商官方发布说明。
Gemini 3.7 Flash 的价格是多少?
在 Synthorai 上,Gemini 3.7 Flash 输入 $0.75/百万 token、输出 $3.75/百万 token,即厂商牌价,无平台加价。缓存命中的输入 token 按 $0.075/M 计费。
Gemini 3.7 Flash 支持提示词缓存(prompt caching)吗?
支持:自动缓存默认开启,另有显式模式可获得确定性折扣。缓存命中的输入 token 按 $0.075/M 计费(未命中 $0.75/M);提示词需有 4,096 token 以上的稳定前缀才能命中缓存。 提示词缓存指南 →
如何开通 Gemini 3.7 Flash?
把现有 OpenAI SDK 的 base_url 指向 "https://synthorai.io/v1",model 设为 "gemini-3.7-flash" 即可。一把 API key 通用网关上的全部模型。
Gemini 3.7 Flash 的知识截止日期是什么时候?
Gemini 3.7 Flash 的知识截止日期为 2026-03,依据厂商官方文档(数据核验于 2026-08-15)。
相关模型
对比
本页每个值都转录自厂商自己的文档(链接见上),并带有核对日期。价格在全目录范围内比较;各厂商定义不同的规格值,只说明差异而不作图表对比。此处没有任何由我们测量的数据,也不做评分。