Claude Fable 5.1 vs GPT-6 Luna
Claude Fable 5.1 is served by invitation. Its figures below are the live rates, but calls need a workspace grant first; ask us for access before you build on this comparison.
Which one, when
Both take text and image in and return text, cap output at 128,000 tokens, and offer about a million tokens of context (1,000,000 for claude-fable-5-1, 1,050,000 for gpt-6-luna), so the real split is price and control over thinking. gpt-6-luna costs $0.1 input and $0.5 output per million against $10 and $50, 100x cheaper on both legs, with cache reads at $0.01 versus $0.25, and its thinking can be turned off, which suits high-volume or latency-sensitive work. Pick claude-fable-5-1 when you want Anthropic's always-on reasoning mode and will pay for it.
Benchmarks
Vendor-published: Alibaba (Qwen) Anthropic ByteDance DeepSeek Google MiniMax Moonshot OpenAI Tencent Z.ai
Pricing
| Claude Fable 5.1 | GPT-6 Luna | Δ | |
|---|---|---|---|
| Input / 1M tokens | $10 | $0.1 | 100× |
| Output / 1M tokens | $50 | $0.5 | 100× |
| Cache read / 1M tokens | $0.25 | $0.01 | 25× |
| Cache write | 1.25x (5m) / 2x (1h) | no separate charge | - |
Rates from the live catalog at build time; each model page carries the current card.
Where they sit - input price per 1M tokens across all 74 chat models on this billing unit (log scale)
Capabilities
| Claude Fable 5.1 | GPT-6 Luna | |
|---|---|---|
| Tool use | yes | yes |
| Thinking control | always on | configurable |
| Structured output | yes | yes |
| Prompt caching | explicit (you mark the prefix) | implicit (automatic) |
| Cache lifetime | 5m default, 1h option | 5-10m, up to 1h |
| Minimum cached prefix | 1024 tokens | 1024 tokens |
Specs
| Claude Fable 5.1 | GPT-6 Luna | |
|---|---|---|
| Input modalities | text image | text image |
| Output modalities | text | text |
| Released | 2026-09-01 | 2026-09-22 |
| Knowledge cutoff | 2026-06 | 2026-05 |
| Context window | 1M | 1.1M |
| Max output | 128K | 128K |
| Thinking parameter | output_config.effort (thinking is adaptive-only and always on) | reasoning.effort |
| Accepted values | effort
| reasoning.effort
|
| Default | thinking always on (adaptive) effort
| medium |
Specs are transcribed from each vendor’s documentation; a row a vendor does not publish is left out rather than inferred. Full sources: Claude Fable 5.1 · GPT-6 Luna
One prompt, both models - measured through the gateway
Claude Fable 5.1 passed · 3 sentences
When we enabled reasoning, the model started generating a long chain of internal "thinking" tokens before producing each answer, and we're billed for those tokens at the same output rate even though users never see them. In practice, that hidden reasoning often runs several times longer than the visible response, so the total tokens per request roughly doubled while the number of requests stayed flat. The upside is measurably better accuracy on complex tasks, so the right conversation is whether that quality gain justifies the cost—and if not, we can restrict reasoning to the workflows that actually need it rather than turning it on globally.
out 186 tok latency 3.2 s
GPT-6 Luna passed · 3 sentences
Enabling reasoning typically makes the model process more information and generate more internal tokens before producing its answer. Because usage is billed by tokens—and reasoning may also route requests to a more expensive model—the cost per request can rise sharply. If request volume stayed similar, that higher cost per request could explain why the bill doubled.
out 117 tok (+48 thinking) latency 3.4 s
Instruction following (exactly three sentences - countable), audience shaping (CFO register), and the hidden-thinking billing gap the token meters below expose.
Claude Fable 5.1 passed · 8/8 cases
```python def longest_run(xs): if not xs: return 0 best = cur = 1 for i in range(1, len(xs)): cur = cur + 1 if xs[i] == xs[i - 1] else 1 best = max(best, cur) return best ``` The bug is that `cur` counts *matches between adjacent elements* rather than *elements in the run*: it starts at 0 and only increments when `xs[i] == xs[i-1]`, so a run of length k produces a count of k−1. That's why a single element returns 0 instead of 1, and why every other run also comes out one short — the "misses runs at the end" reports are the same off-by-one seen on inputs like `[1, 2, 2]`, where the trailing run of length 2 is reported as 1 (the original code does update `best` inside the loop, so end-of-list runs aren't actually skipped, just undercounted). […]
out 882 tok (+507 thinking) latency 3.4 s
GPT-6 Luna passed · 8/8 cases
```python def longest_run(xs): if not xs: return 0 best = cur = 1 for i in range(1, len(xs)): if xs[i] == xs[i - 1]: cur += 1 else: cur = 1 best = max(best, cur) return best ``` The original code counted matching *adjacent pairs* rather than items, so a run of length `n` was reported as `n - 1`, and a single-item input returned `0`. This version counts the current item as part of each run and updates `best` at every index, including the last; it returns `0` for an empty input.
out 514 tok (+356 thinking) latency 6.9 s
Whether the fix is actually correct (runnable), explanation density, and token efficiency on a bounded task.
Claude Fable 5.1 passed · 5/5 fields, guidance null
```json { "company": "Meridian Robotics", "quarter": "Q3 (September quarter)", "revenue_usd_m": 412, "growth_pct": 11, "guidance": null } ``` Notes on judgment calls (not part of the object): - **revenue_usd_m**: The text says "412 for the topline" but never states the currency or that the figure is in millions; it's also on a constant-currency basis and prefaced with "call it," so it's approximate. If you want strict literalism, this should be `null`. - **growth_pct**: "Eleven and change" means slightly above 11%; the exact figure isn't given. - **quarter**: "September quarter" is Q3 on a calendar basis, consistent with the reference to a pending Q4 close. - **guidance**: Explicitly withheld, so `null`.
out 633 tok (+349 thinking) latency 3.8 s
GPT-6 Luna passed · 5/5 fields, guidance "Withheld pending the Q4 close"
{"company":"Meridian Robotics","quarter":"September quarter","revenue_usd_m":412,"growth_pct":null,"guidance":"Withheld pending the Q4 close"}
out 161 tok (+119 thinking) latency 21.7 s
Schema obedience (no invented fields), hallucination pressure (guidance is explicitly withheld), and structured-output path differences.
Claude Fable 5.1 passed · 121 words, 0 banned, 1 question
**Introducing Cross-Provider Prompt Caching for the Gateway.** Starting today, your API gateway can cache prompts once and reuse them across OpenAI, Anthropic, Google, and any other provider you route to. Long system instructions, retrieval context, and few-shot examples are stored at the gateway layer, so switching models no longer means paying to resend the same tokens. Why send the same 8,000-token preamble to three different vendors when you could send it once? Our cache normalizes prompt segments, fingerprints them, and serves matching prefixes from memory before the request ever reaches an upstream API. Early customers report token costs falling by 30 to 60 percent on high-volume workloads, with median latency dropping alongside. Available today on all Team and Enterprise plans.
out 1722 tok (+1463 thinking) latency 2.9 s
GPT-6 Luna passed · 120 words, 0 banned, 1 question
Introducing Prompt Cache, a new API gateway feature that recognizes repeat prompt prefixes and reuses provider-side cached context across supported models. Teams can route requests to different AI providers while preserving eligible cache hits, reducing redundant input processing and helping lower latency and token costs. Configure cache policies in one place, monitor hit rates by provider, and keep existing client integrations unchanged. The gateway applies provider-specific rules automatically, so developers do not need to build separate caching logic for each endpoint. Which workflows could benefit from faster responses and predictable spend? Prompt Cache is available today in preview for eligible accounts, with usage details, supported providers, and setup guidance in the dashboard. Start with one route, compare results, then expand.
out 959 tok (+813 thinking) latency 14.6 s
Constraint obedience (word budget, banned-word list, the single question), style fingerprint, and length control.
Switch between them with one line
Both ids are in every tab below - the highlighted pair of lines is the only edit. Same endpoint, same key, same request shape.
from openai import OpenAI
client = OpenAI(
base_url="https://synthorai.io/v1",
api_key="sk-syn-...",
)
resp = client.chat.completions.create(
model="claude-fable-5-1",
# model="gpt-6-luna", # uncomment this line, comment the one above
messages=[{"role": "user", "content": "Summarize this diff"}],
reasoning_effort="medium",
)
print(resp.choices[0].message.content)import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://synthorai.io/v1",
apiKey: "sk-syn-...",
});
const resp = await client.chat.completions.create({
model: "claude-fable-5-1",
// model: "gpt-6-luna", // uncomment this line, comment the one above
messages: [{ role: "user", content: "Summarize this diff" }],
reasoning_effort: "medium",
});
console.log(resp.choices[0].message.content);curl https://synthorai.io/v1/chat/completions \
-H "Authorization: Bearer sk-syn-..." \
-H "Content-Type: application/json" \
-d '{
"model": "claude-fable-5-1",
# "model": "gpt-6-luna", # uncomment this line, comment the one above
"messages": [{"role": "user", "content": "Hello"}],
"reasoning_effort": "medium"
}'package main
import (
"context"
"fmt"
"github.com/openai/openai-go/v3"
"github.com/openai/openai-go/v3/option"
)
func main() {
client := openai.NewClient(
option.WithBaseURL("https://synthorai.io/v1"),
option.WithAPIKey("sk-syn-..."),
)
resp, _ := client.Chat.Completions.New(context.TODO(), openai.ChatCompletionNewParams{
Model: "claude-fable-5-1",
// Model: "gpt-6-luna", // uncomment this line, comment the one above
Messages: []openai.ChatCompletionMessageParamUnion{
openai.UserMessage("Summarize this diff"),
},
ReasoningEffort: openai.ReasoningEffortMedium,
})
fmt.Println(resp.Choices[0].Message.Content)
}import com.openai.client.OpenAIClient;
import com.openai.client.okhttp.OpenAIOkHttpClient;
import com.openai.models.chat.completions.*;
import com.openai.models.ReasoningEffort;
OpenAIClient client = OpenAIOkHttpClient.builder()
.baseUrl("https://synthorai.io/v1")
.apiKey("sk-syn-...")
.build();
ChatCompletion resp = client.chat().completions().create(
ChatCompletionCreateParams.builder()
.model("claude-fable-5-1")
// .model("gpt-6-luna") // uncomment this line, comment the one above
.addUserMessage("Summarize this diff")
.reasoningEffort(ReasoningEffort.MEDIUM)
.build());
System.out.println(resp.choices().get(0).message().content().orElse(""));FAQ
Which is cheaper, Claude Fable 5.1 or GPT-6 Luna?
GPT-6 Luna is cheaper on input / 1m tokens ($0.1 vs $10, 100× apart). Other rows may point the other way - the table above carries the full card, and real cost depends on your mix.
Can I A/B test Claude Fable 5.1 against GPT-6 Luna without two integrations?
Yes. Both are served through the same OpenAI-compatible endpoint with one API key - switching is a one-line model-string change, so you can route a fraction of traffic to each and compare bills directly.
Do Claude Fable 5.1 and GPT-6 Luna support prompt caching?
Yes - both bill cached reads below their input rate, so warm-prefix workloads cost less than the list rates suggest. The exact cache-read rows are in the pricing table above.