New Sign up free, 10 calls on us. Up to $1, no card needed.

GPT-6.1 Sol

Released 2026-09-29

chatVisionCodeTool callingReasoningPrompt caching

GPT-6.1 Sol is the September 29, 2026 upgrade to GPT-6 Sol, which OpenAI positions as near-Astra performance for complex work at a lower cost, aimed at coding, computer use, and professional work.

Input
text image $2/M
Output
text $10/M
Cache read
$0.1/M
Context
1.1M
vs GPT-4o
~60% cheaper
Knowledge cutoff
2026-04

Prompts over 272K tokens: the whole request bills at $4/M input · $15/M output

Price in context

Where the price sits among 68 comparable models

Input$2/M
$0.05 · Qwen3 VL Flash GPT-5.4 Pro · $30
Output$10/M
$0.275 · DeepSeek V4 Flash GPT-5.4 Pro · $180
Cached read$0.1/M
$0.0028 · DeepSeek V4 Flash GPT-5.4 Pro · $15

The bar shows how this model’s price compares with every other model of the same kind on Synthorai. The cheapest and the most expensive are named at each end. These are base rates; batch, region and cache-write discounts are on the pricing page.

Specs & limits

Tokens

Context window (vendor spec) 1,050,000
Max output (vendor spec) 128,000
Knowledge cutoff 2026-04

Prompt caching

How it caches automatic
Min prefix 1,024
Lifetime 5-10m, up to 1h

Thinking

Vendor control reasoning.effort
Accepted values low · medium · high · xhigh · max
Default medium applied when the request sets nothing
Can be turned off No
Parameter reasoning_effort
Values minimal · low · medium · high the gateway's parameter surface - the vendor mapping above applies

Model

Modalities text + image → text
  • Upgrade to GPT-6 Sol with near-Astra performance for complex coding, computer use and professional work. 1.05M context of which up to 922K can be input, 128K max output. Cached input is 0.05x the input rate, half of GPT-6 Sol's. Prompts over 272K input tokens move the whole request onto the long-context rate (2x input and cache, 1.5x output). Reasoning cannot be turned off: none and minimal are not supported. Tool calling requires the Responses API
  • Chat Completions is supported without tools

per OpenAI official docs ↗

One prompt, measured through the gateway

PROMPT Explain to a CFO, in exactly three sentences, why our LLM bill doubled after we enabled reasoning. CHECK exactly 3 sentences

GPT-6.1 Sol passed · 3 sentences

Enabling reasoning makes the model do more computational work before answering, and many providers bill the resulting reasoning tokens even though users never see them. That means the same number of requests can cost twice as much because each request consumes more billable tokens, even if the visible answers stay the same length. We should confirm the increase in our usage breakdown and reserve higher reasoning levels for tasks where the improvement justifies the added cost.

out 144 tok (+54 thinking) latency 6.6 s

Instruction following (exactly three sentences - countable), audience shaping (CFO register), and the hidden-thinking billing gap the token meters below expose.

PROMPT This function is supposed to return the longest run of consecutive equal items, but callers report it is off by one on single-element inputs and misses runs that end at the last index. Fix it and explain the bug in one paragraph. CHECK fix passes tests

GPT-6.1 Sol passed · 8/8 cases

```python def longest_run(xs): if not xs: return 0 best = cur = 1 for i in range(1, len(xs)): if xs[i] == xs[i - 1]: cur += 1 else: cur = 1 best = max(best, cur) return best ``` The original function counts equal adjacent pairs rather than items in a run: a run of length `n` has only `n - 1` equal pairs, so a single-element input returns `0`, and a run ending at the last index is also undercounted. It does process the last index, so the issue is not a missing final check. Initializing each nonempty run to `1` and updating `best` after every iteration correctly counts items, including singleton and trailing runs; empty inputs still return `0`.

out 189 tok latency 6.3 s

Whether the fix is actually correct (runnable), explanation density, and token efficiency on a bounded task.

PROMPT Extract a JSON object with fields {company, quarter, revenue_usd_m, growth_pct, guidance} from this text. Use null for anything not stated; add no other fields. CHECK valid JSON, schema exact

GPT-6.1 Sol passed · 5/5 fields, guidance "Withheld pending the Q4 close"

{ "company": "Meridian Robotics", "quarter": "September quarter", "revenue_usd_m": null, "growth_pct": "eleven and change", "guidance": "Withheld pending the Q4 close" }

out 345 tok (+285 thinking) latency 10.2 s

Schema obedience (no invented fields), hallucination pressure (guidance is explicitly withheld), and structured-output path differences.

PROMPT Write a 120-word product announcement for an API gateway feature that caches prompts across providers. Forbidden words: "seamless", "unlock", "game-changer", "revolutionize", "empower". Exactly one sentence must be a question. CHECK 120 words, 0 banned words

GPT-6.1 Sol passed · 120 words, 0 banned, 1 question

Introducing Cross-Provider Prompt Cache, a new API gateway feature that stores reusable prompts and manages caching across your supported AI providers. Why rebuild the same context every time your application switches models? With one configuration, teams can reuse shared instructions, standardize cache policies, and reduce repeated prompt processing wherever provider caching is available. The gateway handles provider-specific requirements while giving you clear visibility into cache hits, usage, and estimated savings. Set expiration windows, isolate cached content by project, and invalidate entries when prompts change. Your existing routing logic stays intact, so you can compare models without rebuilding your caching workflow. […]

out 588 tok (+435 thinking) latency 13.9 s

Constraint obedience (word budget, banned-word list, the single question), style fingerprint, and length control.

Use GPT-6.1 Sol in 30 seconds

OpenAI-compatible: swap the base_url, keep your SDK. POST /v1/chat/completions

from openai import OpenAI

client = OpenAI(
    base_url="https://synthorai.io/v1",
    api_key="sk-syn-...",
)

resp = client.chat.completions.create(
    model="gpt-6.1-sol",
    messages=[{"role": "user", "content": "Summarize this diff"}],
    reasoning_effort="medium",
)
print(resp.choices[0].message.content)

About GPT-6.1 Sol

  • It keeps the GPT-6 family's 1,050,000-token context window, of which up to 922,000 tokens can be input, with 128K max output tokens, text and image input, structured outputs, streaming, function calling, prompt caching, and an April 2026 knowledge cutoff.
  • Input and output list prices match GPT-6 Sol, but cached input costs half as much, which adds up in agent loops that resend long prefixes.
  • Prompts above 272K input tokens move the whole request onto the long-context rate at twice the input and cache rates and 1.5x the output price.
  • Reasoning is always on: effort runs from low to max with medium as the default, and OpenAI does not support none or minimal on this model, so Synthorai runs requests that ask for them at low instead of rejecting them.
  • OpenAI reserves tool calling for the Responses API; through Synthorai a Chat Completions request with tools still works, because the gateway sends it upstream in the form that supports tool calling.
  • Synthorai serves GPT-6.1 Sol through the same OpenAI-compatible API as the rest of the fleet.

FAQ

Is the GPT-6.1 Sol API free to try?

Yes: new accounts get 10 trial calls and up to $1 in free credit, no card required. At $2/M input tokens, that credit alone covers roughly 62 requests of ~8K tokens against GPT-6.1 Sol.

What is GPT-6.1 Sol best at?

GPT-6 Sol upgrade with near-Astra performance for coding, computer use and professional work; 1.05M context, 128K output, reasoning effort low to max; cached input at half GPT-6 Sol's rate, long-context pricing above 272K input tokens. See the About section for the full picture from the vendor's own release notes.

How much does GPT-6.1 Sol cost?

GPT-6.1 Sol costs $2 per million input tokens and $10 per million output tokens on Synthorai. That is the provider's list price, with no platform markup. Cached input tokens bill at $0.1/M.

Does GPT-6.1 Sol support prompt caching?

Yes, automatically: OpenAI-served prompts cache with no code changes. Cached input tokens bill at $0.1/M vs $2/M uncached; prompts need a 1,024-token stable prefix to cache (TTL 5-10m, up to 1h). Prompt caching guide →

How do I get access to GPT-6.1 Sol?

Point your existing OpenAI SDK at base_url="https://synthorai.io/v1", set model="gpt-6.1-sol", and you're done. One API key covers every model on the gateway.

What is GPT-6.1 Sol's knowledge cutoff?

GPT-6.1 Sol's knowledge cutoff is 2026-04, per the vendor's official documentation (as of 2026-09-30).

Related models

Compare

Every value on this page is transcribed from the vendor's own documentation, linked above, and carries the date it was checked. Prices are compared across the catalogue; specification values that vendors define differently are shown with the difference stated rather than charted. Nothing here is measured by us, and nothing is scored.

Get your API key Compare your cost →