DeepSeek V4 Pro 0813 是 DeepSeek 旗舰 V4 Pro 系列在 2026 年 8 月的发布版本,模型卡把它定位为预览版的接替版,而不是并列的一个变体。
- 输入
- 文本 $1.32/M
- 输出
- 文本 $3.96/M
- 缓存读取
- $0.132/M
- 上下文
- 1M
- 对比 GPT-4o
- 便宜约 74%
Benchmark 成绩
供应商公布: Alibaba (Qwen) Anthropic ByteDance DeepSeek Google MiniMax Moonshot OpenAI Tencent Z.ai
价格在同类中的位置
价格在 68 个同类模型中的位置
这条线标出该模型的价格在 Synthorai 同类模型中的位置,两端分别是最便宜和最贵的模型。这里是基础价,批量、区域和缓存写入的折扣见价格页。
规格与限制
Token
| 上下文窗口(供应商规格) | 1,000,000 |
|---|---|
| 最大输出(供应商规格) | 393,216 |
提示词缓存
| 缓存方式 | 自动 |
|---|---|
| 有效期 | 无固定 TTL(长期未使用即自动清理) |
思考
| 供应商参数 | reasoning_effort |
|---|---|
| 可选值 | the model card documents low · high · max |
| 思考行为 | The trace returns in reasoning_content and on short prompts can account for the large majority of the completion tokens, so a tight max_tokens budget returns an empty answer that is still billed in full. DeepSeek documents three effort levels for its own API; platforms that host the model may accept a different set, so confirm against the platform you are routed to. |
| 参数 | reasoning_effort |
| 取值 | minimal · low · medium · high 这是网关侧的参数,实际以上方的供应商映射为准 |
模型
| 模态 | 文本 → 文本 |
|---|---|
| 许可证 | MIT |
August 2026 release of DeepSeek V4 Pro, superseding the preview version with what the model card describes as greatly enhanced agentic capabilities, and shipping with a DSpark speculative decoding module attached. The card states neither a parameter count nor an architecture, so neither is listed here.
一条提示词,经网关实测
DeepSeek V4 Pro (0813) 通过 · 3 sentences
Enabling reasoning causes the model to generate a hidden chain-of-thought with many additional tokens before producing the final answer, which sharply increases compute consumption per request. Those extra reasoning tokens are billed at the same or higher rates, so total usage doubled even though the volume of final outputs stayed similar. We can control this by adjusting reasoning effort, setting token limits, or routing only complex requests through reasoning mode.
输出 308 tok (+226 思考) 延迟 5.1 s
指令遵循(恰好三句,数得出来)、受众适配(对 CFO 说话的口吻),以及下方 token 计数暴露出的隐藏思考计费差额。
DeepSeek V4 Pro (0813) 通过 · 8/8 cases
```python def longest_run(xs): if not xs: return 0 best = 1 cur = 1 for i in range(1, len(xs)): if xs[i] == xs[i - 1]: cur += 1 else: cur = 1 best = max(best, cur) return best ``` The bug is that `cur` was initialized to `0` and only counted adjacent matching pairs, so a run of length `n` produced `n - 1` (or `0` for a single element) instead of `n`. The code also reset to `0` on a change, when the new run should start at length `1`. Initializing `cur` and `best` to `1` fixes both the single-element case and runs ending at the last index.
输出 1315 tok (+1130 思考) 延迟 16.9 s
修复是否真的正确(能运行)、解释的信息密度,以及在一个范围明确的任务上的 token 效率。
DeepSeek V4 Pro (0813) 通过 · 5/5 fields, guidance null
{"company":"Meridian Robotics","quarter":"September quarter","revenue_usd_m":412,"growth_pct":11.0,"guidance":null}
输出 2155 tok (+2121 思考) 延迟 26.0 s
是否严守 schema(不臆造字段)、能否顶住幻觉压力(原文明说暂不给出 guidance),以及结构化输出路径的差异。
DeepSeek V4 Pro (0813) 通过 · 121 words, 0 banned, 1 question
Introducing cross-provider prompt caching in our API gateway. Send a prompt once, and the gateway stores its computed prefix across Anthropic, OpenAI, and other LLM providers. Subsequent requests with the same prompt hit the cache, cutting latency and token costs while keeping outputs consistent across routing decisions and provider failovers. Teams can route identical prompts between providers without reprocessing shared context or lengthy system instructions. How much could you save on repeated prompt prefixes? The cache respects provider-specific key formats, handles TTLs automatically, and works with streaming and batch requests. Enable it with one configuration flag—no changes to your application code. Available today on all plans. Monitor cache hit rates, token savings, and provider-specific performance metrics in the live dashboard.
输出 2845 tok (+2694 思考) 延迟 25.5 s
是否守住约束(字数预算、禁用词表、只能有一个问句)、文风特征,以及长度控制。
30 秒用上 DeepSeek V4 Pro (0813)
兼容 OpenAI:换掉 base_url,SDK 不用改。POST /v1/chat/completions
from openai import OpenAI
client = OpenAI(
base_url="https://synthorai.io/v1",
api_key="sk-syn-...",
)
resp = client.chat.completions.create(
model="deepseek-v4-pro-0813",
messages=[{"role": "user", "content": "Summarize this diff"}],
reasoning_effort="medium",
)
print(resp.choices[0].message.content)import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://synthorai.io/v1",
apiKey: "sk-syn-...",
});
const resp = await client.chat.completions.create({
model: "deepseek-v4-pro-0813",
messages: [{ role: "user", content: "Summarize this diff" }],
reasoning_effort: "medium",
});
console.log(resp.choices[0].message.content);curl https://synthorai.io/v1/chat/completions \
-H "Authorization: Bearer sk-syn-..." \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek-v4-pro-0813",
"messages": [{"role": "user", "content": "Hello"}],
"reasoning_effort": "medium"
}'package main
import (
"context"
"fmt"
"github.com/openai/openai-go/v3"
"github.com/openai/openai-go/v3/option"
)
func main() {
client := openai.NewClient(
option.WithBaseURL("https://synthorai.io/v1"),
option.WithAPIKey("sk-syn-..."),
)
resp, _ := client.Chat.Completions.New(context.TODO(), openai.ChatCompletionNewParams{
Model: "deepseek-v4-pro-0813",
Messages: []openai.ChatCompletionMessageParamUnion{
openai.UserMessage("Summarize this diff"),
},
ReasoningEffort: openai.ReasoningEffortMedium,
})
fmt.Println(resp.Choices[0].Message.Content)
}import com.openai.client.OpenAIClient;
import com.openai.client.okhttp.OpenAIOkHttpClient;
import com.openai.models.chat.completions.*;
import com.openai.models.ReasoningEffort;
OpenAIClient client = OpenAIOkHttpClient.builder()
.baseUrl("https://synthorai.io/v1")
.apiKey("sk-syn-...")
.build();
ChatCompletion resp = client.chat().completions().create(
ChatCompletionCreateParams.builder()
.model("deepseek-v4-pro-0813")
.addUserMessage("Summarize this diff")
.reasoningEffort(ReasoningEffort.MEDIUM)
.build());
System.out.println(resp.choices().get(0).message().content().orElse(""));关于 DeepSeek V4 Pro (0813)
- DeepSeek 把这次变化归因于大幅增强的智能体能力,以及它描述为在生产环境中尤为明显的性能提升,本次发布还附带了一个 DSpark 推测解码模块。
- 模型卡在规格上异常简略:既没有给出参数量,也没有给出架构,因此凡是从此前 V4 Pro 沿用过来的数字,对本版本都应视为未经确认。
- 它确实给出的是运行边界——MIT 许可的权重、在较高推理力度下建议的 384K 最大输出长度,以及一个文档记载为 low、high、max 三档的 reasoning_effort 参数。
- 真正需要提前规划的是推理部分。
- 思维轨迹回传在 reasoning_content 中,短提示下它可能占到输出 token 的绝大部分,于是过紧的 max_tokens 预算会返回空回答,而思考所花的 token 照样全额计费;接入延迟敏感的链路之前,请为此留出预算,或者调低推理力度。
- 推理力度的成本还不止于输出:调高档位会加长模型据以工作的前置内容,因此同一条消息在高档位下计入的输入 token 比低档位更多。
- 该模型输入输出均为纯文本,不支持图像输入,需要视觉能力时应搭配视觉模型,而不是直接发多模态消息。
- 由于带日期的发布版与不带日期的滚动名都可调用,需要行为可复现时请固定带日期的 id,并预期不带日期的名字会随时间前移。
- Synthorai 通过 OpenAI 兼容的 chat completions 端点提供该模型,客户端无需改动。
常见问题
DeepSeek V4 Pro (0813) API 可以免费试用吗?
可以。新账号有 10 次试用调用,免费额度最高 $1,无需绑卡。按输入 $1.32/M 计算,光这笔额度就够向 DeepSeek V4 Pro (0813) 发约 94 次请求,每次 ~8K token。
DeepSeek V4 Pro (0813) 最擅长什么?
接替 V4 Pro 预览版、MIT 许可,建议最大输出 384K、reasoning_effort 文档记载 low、high、max 三档。完整介绍见「关于」部分,内容取自供应商的官方发布说明。
DeepSeek V4 Pro (0813) 的价格是多少?
在 Synthorai 上,DeepSeek V4 Pro (0813) 的输入价为 $1.32/百万 token,输出价为 $3.96/百万 token。这就是供应商官网价,平台不加价。缓存命中的输入 token 按 $0.132/M 计费。
DeepSeek V4 Pro (0813) 支持提示词缓存(prompt caching)吗?
支持,而且全自动:经 DeepSeek 处理的提示词会自动缓存,不用改代码。缓存命中的输入 token 按 $0.132/M 计费(未命中为 $1.32/M)(TTL 无固定 TTL(长期未使用即自动清理))。 提示词缓存指南 →
如何开通 DeepSeek V4 Pro (0813)?
把现有 OpenAI SDK 的 base_url 指向 "https://synthorai.io/v1",model 设为 "deepseek-v4-pro-0813" 即可。一个 API key 就能调用网关上的全部模型。
DeepSeek V4 Pro (0813) 开源吗?
开源:权重以 MIT 许可证发布(官方仓库链接见「关于」部分)。不想自己备 GPU 也行:这里的托管版本按量付费,不需要自建基础设施。 运行开源权重模型 →
相关模型
对比
本页的每个数值都照录自供应商自己的文档(链接见上),并标有核对日期。价格是在整个目录范围内比较的;各家供应商定义不一样的规格项,只说明差异,不画图对比。这里没有任何一项是我们自己测的,也不打分。