DeepSeek V4 Pro 0813 is the August 2026 release of DeepSeek's flagship V4 Pro line, and the model card presents it as superseding the preview version rather than sitting beside it as a variant.
- Input
- text $1.32/M
- Output
- text $3.96/M
- Cache read
- $0.132/M
- Context
- 1M
- vs GPT-4o
- ~74% cheaper
Benchmarks
Vendor-published: Alibaba (Qwen) Anthropic ByteDance DeepSeek Google MiniMax Moonshot OpenAI Tencent Z.ai
Price in context
Where the price sits among 68 comparable models
The bar shows how this model’s price compares with every other model of the same kind on Synthorai. The cheapest and the most expensive are named at each end. These are base rates; batch, region and cache-write discounts are on the pricing page.
Specs & limits
Tokens
| Context window (vendor spec) | 1,000,000 |
|---|---|
| Max output (vendor spec) | 393,216 |
Prompt caching
| How it caches | automatic |
|---|---|
| Lifetime | no fixed TTL (evicted when unused) |
Thinking
| Vendor control | reasoning_effort |
|---|---|
| Accepted values | the model card documents low · high · max |
| Thinking behaviour | The trace returns in reasoning_content and on short prompts can account for the large majority of the completion tokens, so a tight max_tokens budget returns an empty answer that is still billed in full. DeepSeek documents three effort levels for its own API; platforms that host the model may accept a different set, so confirm against the platform you are routed to. |
| Parameter | reasoning_effort |
| Values | minimal · low · medium · high the gateway's parameter surface; the vendor mapping above applies |
Model
| Modalities | text → text |
|---|---|
| License | MIT |
August 2026 release of DeepSeek V4 Pro, superseding the preview version with what the model card describes as greatly enhanced agentic capabilities, and shipping with a DSpark speculative decoding module attached. The card states neither a parameter count nor an architecture, so neither is listed here.
One prompt, measured through the gateway
DeepSeek V4 Pro (0813) passed · 3 sentences
Enabling reasoning causes the model to generate a hidden chain-of-thought with many additional tokens before producing the final answer, which sharply increases compute consumption per request. Those extra reasoning tokens are billed at the same or higher rates, so total usage doubled even though the volume of final outputs stayed similar. We can control this by adjusting reasoning effort, setting token limits, or routing only complex requests through reasoning mode.
out 308 tok (+226 thinking) latency 5.1 s
Instruction following (exactly three sentences, so countable), audience shaping (CFO register), and the hidden-thinking billing gap the token meters below expose.
DeepSeek V4 Pro (0813) passed · 8/8 cases
```python def longest_run(xs): if not xs: return 0 best = 1 cur = 1 for i in range(1, len(xs)): if xs[i] == xs[i - 1]: cur += 1 else: cur = 1 best = max(best, cur) return best ``` The bug is that `cur` was initialized to `0` and only counted adjacent matching pairs, so a run of length `n` produced `n - 1` (or `0` for a single element) instead of `n`. The code also reset to `0` on a change, when the new run should start at length `1`. Initializing `cur` and `best` to `1` fixes both the single-element case and runs ending at the last index.
out 1315 tok (+1130 thinking) latency 16.9 s
Whether the fix is actually correct (runnable), explanation density, and token efficiency on a bounded task.
DeepSeek V4 Pro (0813) passed · 5/5 fields, guidance null
{"company":"Meridian Robotics","quarter":"September quarter","revenue_usd_m":412,"growth_pct":11.0,"guidance":null}
out 2155 tok (+2121 thinking) latency 26.0 s
Schema obedience (no invented fields), hallucination pressure (guidance is explicitly withheld), and structured-output path differences.
DeepSeek V4 Pro (0813) passed · 121 words, 0 banned, 1 question
Introducing cross-provider prompt caching in our API gateway. Send a prompt once, and the gateway stores its computed prefix across Anthropic, OpenAI, and other LLM providers. Subsequent requests with the same prompt hit the cache, cutting latency and token costs while keeping outputs consistent across routing decisions and provider failovers. Teams can route identical prompts between providers without reprocessing shared context or lengthy system instructions. How much could you save on repeated prompt prefixes? The cache respects provider-specific key formats, handles TTLs automatically, and works with streaming and batch requests. Enable it with one configuration flag—no changes to your application code. Available today on all plans. Monitor cache hit rates, token savings, and provider-specific performance metrics in the live dashboard.
out 2845 tok (+2694 thinking) latency 25.5 s
Constraint obedience (word budget, banned-word list, the single question), style fingerprint, and length control.
Use DeepSeek V4 Pro (0813) in 30 seconds
OpenAI-compatible: swap the base_url, keep your SDK. POST /v1/chat/completions
from openai import OpenAI
client = OpenAI(
base_url="https://synthorai.io/v1",
api_key="sk-syn-...",
)
resp = client.chat.completions.create(
model="deepseek-v4-pro-0813",
messages=[{"role": "user", "content": "Summarize this diff"}],
reasoning_effort="medium",
)
print(resp.choices[0].message.content)import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://synthorai.io/v1",
apiKey: "sk-syn-...",
});
const resp = await client.chat.completions.create({
model: "deepseek-v4-pro-0813",
messages: [{ role: "user", content: "Summarize this diff" }],
reasoning_effort: "medium",
});
console.log(resp.choices[0].message.content);curl https://synthorai.io/v1/chat/completions \
-H "Authorization: Bearer sk-syn-..." \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek-v4-pro-0813",
"messages": [{"role": "user", "content": "Hello"}],
"reasoning_effort": "medium"
}'package main
import (
"context"
"fmt"
"github.com/openai/openai-go/v3"
"github.com/openai/openai-go/v3/option"
)
func main() {
client := openai.NewClient(
option.WithBaseURL("https://synthorai.io/v1"),
option.WithAPIKey("sk-syn-..."),
)
resp, _ := client.Chat.Completions.New(context.TODO(), openai.ChatCompletionNewParams{
Model: "deepseek-v4-pro-0813",
Messages: []openai.ChatCompletionMessageParamUnion{
openai.UserMessage("Summarize this diff"),
},
ReasoningEffort: openai.ReasoningEffortMedium,
})
fmt.Println(resp.Choices[0].Message.Content)
}import com.openai.client.OpenAIClient;
import com.openai.client.okhttp.OpenAIOkHttpClient;
import com.openai.models.chat.completions.*;
import com.openai.models.ReasoningEffort;
OpenAIClient client = OpenAIOkHttpClient.builder()
.baseUrl("https://synthorai.io/v1")
.apiKey("sk-syn-...")
.build();
ChatCompletion resp = client.chat().completions().create(
ChatCompletionCreateParams.builder()
.model("deepseek-v4-pro-0813")
.addUserMessage("Summarize this diff")
.reasoningEffort(ReasoningEffort.MEDIUM)
.build());
System.out.println(resp.choices().get(0).message().content().orElse(""));About DeepSeek V4 Pro (0813)
- DeepSeek attributes the change to greatly enhanced agentic capabilities and to performance improvements it describes as especially pronounced in production settings, and the release ships with a DSpark speculative decoding module attached.
- The card is unusually sparse on specifications: it states neither a parameter count nor an architecture, so treat any figure carried over from the earlier V4 Pro as unconfirmed for this release.
- What it does state is the operating envelope - MIT-licensed weights, a recommended maximum output length of 384K tokens at the higher reasoning-effort levels, and a reasoning_effort parameter documented at three levels, low, high and max.
- Reasoning is the thing to plan around.
- The trace comes back in reasoning_content, and on short prompts it can account for the large majority of the completion tokens, so a tight max_tokens budget will return an empty answer that is still billed in full for the tokens spent thinking; budget for it, or lower the effort, before wiring this into latency-sensitive paths.
- Effort also has a cost side beyond the output: raising it lengthens the preamble the model works from, so the same message bills more input tokens at a higher setting than at a lower one.
- The model is text-only in and text-only out - image input is not supported, so pair it with a vision model rather than sending multimodal messages.
- Because both the dated release and the rolling name remain callable, pin the dated id when you need reproducible behaviour and expect the undated name to move forward over time.
- Synthorai serves it through the OpenAI-compatible chat completions endpoint with no client changes required.
FAQ
Is the DeepSeek V4 Pro (0813) API free to try?
Yes: new accounts get 10 trial calls and up to $1 in free credit, no card required. At $1.32/M input tokens, that credit alone covers roughly 94 requests of ~8K tokens against DeepSeek V4 Pro (0813).
What is DeepSeek V4 Pro (0813) best at?
Supersedes the V4 Pro preview release; MIT-licensed, with 384K recommended max output; reasoning_effort documented at low, high and max. See the About section for the full picture from the vendor's own release notes.
How much does DeepSeek V4 Pro (0813) cost?
DeepSeek V4 Pro (0813) costs $1.32 per million input tokens and $3.96 per million output tokens on Synthorai. That is the provider's list price, with no platform markup. Cached input tokens bill at $0.132/M.
Does DeepSeek V4 Pro (0813) support prompt caching?
Yes, automatically: DeepSeek-served prompts cache with no code changes. Cached input tokens bill at $0.132/M vs $1.32/M uncached (TTL no fixed TTL (evicted when unused)). Prompt caching guide →
How do I get access to DeepSeek V4 Pro (0813)?
Point your existing OpenAI SDK at base_url="https://synthorai.io/v1", set model="deepseek-v4-pro-0813", and you're done. One API key covers every model on the gateway.
Is DeepSeek V4 Pro (0813) open source?
Yes: the weights are published under the MIT license (official repository linked in the About section). Or skip the GPUs: the hosted version here is pay-as-you-go with no infrastructure to run. Running open-weight models →
Related models
Compare
Every value on this page is transcribed from the vendor's own documentation, linked above, and carries the date it was checked. Prices are compared across the catalogue; specification values that vendors define differently are shown with the difference stated rather than charted. Nothing here is measured by us, and nothing is scored.