New Sign up free, 10 calls on us. Up to $1, no card needed.

GPT-6 Luna vs GPT-5.6 Luna: Half the Price, Half the Words

Contents
  1. How does GPT-6 Luna compare with GPT-5.6 Luna on benchmarks?
  2. What else changed besides the price?
  3. How did we test it?
  4. Does GPT-6 Luna write less than GPT-5.6 Luna?
  5. Does GPT-6 Luna leave things out?
  6. What brings the length and the details back?
  7. How much does GPT-6 Luna save over GPT-5.6 Luna?
  8. Do our results match other published tests?
  9. What did we observe, and which should you use?
  10. FAQ

GPT-6 Luna is level with GPT-5.6 Luna on public benchmarks (38 against 37 on the Artificial Analysis Intelligence Index) and costs about half as much per task. The difference is in what it writes: when a prompt only states a goal (“write the Q3 review”), GPT-6 Luna wrote 0.55x the words across 80 business documents. It did not skip parts it was asked for. With the required parts named, the two models were equally complete, and the one detail GPT-6 Luna dropped more often on goal-only prompts was the duration of an incident.

TL;DR

  • Artificial Analysis scores GPT-6 Luna 38 and GPT-5.6 Luna 37 at max effort, at $0.07 against $0.18 per task.
  • On goal-only prompts, GPT-6 Luna wrote 384 words per document against GPT-5.6 Luna’s 703.
  • GPT-6 Luna left the incident duration out of 9 of 20 goal-only incident postmortems; GPT-5.6 Luna out of 2.
  • With the required parts named, GPT-6 Luna was complete on 74 of 80 documents against 69, at $0.81 against $1.61 per 1,000.

How does GPT-6 Luna compare with GPT-5.6 Luna on benchmarks?

GPT-6 Luna matches GPT-5.6 Luna on general benchmarks at 48% to 61% lower cost per task, and the two separate only on documents scored against a rubric (a scoring checklist). Scores are quoted at a reasoning effort, the API setting for how much the model reasons before it answers (none to max, default medium). The figures were read from the publishers’ pages on 2026-10-02.

GPT-6 LunaGPT-5.6 Luna
Released2026-09-222026-07-09
List price, input / output per 1M tokens$0.10 / $0.50$0.20 / $1.20
Artificial Analysis Intelligence Index v4.3.2 at max (10 evals: knowledge, reasoning, coding, agentic and business-document tasks)3837
Cost per index task$0.07$0.18
Vals Index (coding, legal, tax, finance and other professional tasks)51.2%51.7%
Cost per Vals test$0.43$0.82
GDPval-AA v2.1 Elo at max (business documents scored against a hidden rubric; Elo is a rating from head-to-head comparisons, higher is better)14381463
GDPval-AA v2.1 Elo at medium12621133
Output speed, tokens per second131124

GDPval-AA is where the complaint about GPT-6 Luna started. In its launch analysis, Artificial Analysis put GPT-6 Luna about 75 Elo points below GPT-5.6 Luna on GDPval-AA and about 45 below on AA-Briefcase v1.1, a second document benchmark, and attributed it to “reduced presentation quality and deliverables that omit rubric elements”; reviewers named it the “format tax”. The leaderboard now shows a 25-point gap at max, and at medium GPT-6 Luna is 129 points ahead.

Against other vendors, GPT-6 Luna competes with the flash tier, their low-price, high-speed models. On the same index, DeepSeek V4.1 Flash at max scores 39 at $0.27 per task and Gemini 3.8 Flash at high scores 41 at $1.24, so GPT-6 Luna gives up 1 to 3 points for a cost per task 4 to 18 times lower. Both generate faster, at 209 and 249 tokens per second.

What else changed besides the price?

GPT-6 Luna keeps GPT-5.6 Luna’s limits and settings; only the cache price and the knowledge cutoff moved. Both have a 1,050,000-token context window and a 128,000-token output limit, and both charge 2x input and 1.5x output for a whole request once the prompt passes 272K tokens. Cached input costs $0.01 per million tokens, down from $0.02, and the knowledge cutoff moved from Feb 16 to May 18, 2026. Effort is set with reasoning.effort on the Responses API (none, low, medium, high, xhigh, max), and the reasoning is hidden and billed as output tokens.

A week after the launch, OpenAI released GPT-6.1 Sol as the successor to GPT-6 Sol in the tier above, and our GPT-6.1 Sol measurements found its answers back at the old length. Luna got no update, so we measured what GPT-6 Luna writes shorter, what it leaves out, and what brings it back.

How did we test it?

We wrote document tasks whose requirements can be checked in code, so no judge model is involved. Each task gives the model a block of synthetic data and asks for a document built from it:

DocumentDataWhat a reader expects in it
Quarterly reviewrevenue for 6 to 12 regionsevery region covered, the region with the largest decline named
Incident postmortem7 to 10 timestamped eventsevery event, the duration of the incident
Vendor comparison5 to 8 hosting quotesevery vendor covered
Customer replies6 to 14 support ticketsa reply to every ticket, addressed to the customer by name

Each task comes in two prompt styles over the same data. A goal-only prompt says what the document is and nothing about its parts (“Write the Q3 regional review for the leadership team”), and is graded only on the items in the table. A prompt with the parts named lists the sections, counts and word ranges, and is graded on all of them. A document is complete when nothing required is missing, too short or under-counted. Goal-only vendor comparisons count for length only, because the right vendor is a judgment call.

On 2026-10-02 we ran the 80 goal-only items on both models at five effort settings and on GPT-6 Luna at one verbosity setting, and the 80 named-part items on both at the default effort: 1,040 Responses API calls. Every prompt carried a unique random string so no response came from a cache, and cost uses OpenAI list prices. Both models answered the same items, so we compare them with an exact McNemar test (a paired test on the items where the two models differ) and a Holm correction, which tightens the threshold when several comparisons are tested at once. Only gaps that survive it are called differences.

Does GPT-6 Luna write less than GPT-5.6 Luna?

GPT-6 Luna wrote 0.55x the words of GPT-5.6 Luna on goal-only prompts at the default effort: 384 words per document against 703, shorter on all 80 items. It was shorter in every document type, most in postmortems (441 against 981 words) and least in customer replies (576 against 767).

With the parts named, the gap disappears: 738 words against 724.

Horizontal bar chart of words per answer for GPT-6 Luna and GPT-5.6 Luna. On goal-only prompts GPT-6 Luna writes 365 to 436 words across the effort settings and GPT-5.6 Luna 656 to 727; at medium, 384 against 703 words, complete on 51 and 58 of 60. With the parts named, 738 against 724 words, complete on 74 and 69 of 80

The shorter default is not specific to Luna. On the same items, GPT-6 Sol wrote 325 words against 592 for GPT-5.6 Sol, and GPT-6.1 Sol is back at 565.

Does GPT-6 Luna leave things out?

On goal-only prompts GPT-6 Luna left out one thing more often than GPT-5.6 Luna: the duration of the incident in a postmortem. It usually gave the time window (“Incident window: 18:00-18:39 UTC”) and left the subtraction to the reader. Everything else was there at every effort setting: every region and the worst one, every event, and a reply addressed by name for every ticket, with a single exception at none.

Goal-only promptsGPT-6 Luna completeGPT-6 Luna postmortems without the durationGPT-5.6 Luna completeGPT-5.6 Luna postmortems without the duration
none44 / 6015 / 2057 / 603 / 20
low47 / 6013 / 2054 / 606 / 20
medium51 / 609 / 2058 / 602 / 20
high53 / 607 / 2057 / 603 / 20
max59 / 601 / 2060 / 600 / 20

The gap at none (44 against 57) survives the correction. At medium (51 against 58) it is significant only before correction, so read it as a direction, and at max the two are tied. All of it comes from one document type, so it shows GPT-6 Luna skipping a derived figure in a postmortem and no general habit of dropping content.

With the parts named, GPT-6 Luna was complete on 74 of 80 documents and GPT-5.6 Luna on 69, a tie. Every miss on both models was a memo or impact paragraph below its requested word range: 6 for GPT-6 Luna and 11 for GPT-5.6 Luna.

What brings the length and the details back?

Naming the parts in the prompt does, and the two request parameters do not. The same GPT-6 Luna at the same effort went from 384 words to 738 and stated the duration in every postmortem once the prompt asked for it. A request like this is enough:

Write the postmortem with these sections: Timeline (a table with every event),
Impact (120-200 words, state the total duration in minutes), Root Cause,
Action Items (5, each ending "Owner: <team>"), Lessons Learned (at least 3 bullets).

text.verbosity is a Responses API parameter (low, medium, high) for how long and detailed the visible answer is. Set to high, it made GPT-6 Luna’s goal-only documents 7% longer (410 words) and left completeness where it was, 52 of 60 against 51.

Raising reasoning.effort changes how much the model thinks more than how much it writes. From none to max, GPT-6 Luna’s documents grew from 370 to 436 words while reasoning went from 0 to 4,192 tokens per answer. max did recover the duration (missing from 1 of 20), at 4.5 times the cost of medium.

Both parameters with the OpenAI Python SDK:

from openai import OpenAI

client = OpenAI()
prompt = "Write the postmortem with these sections: ..."  # data and required parts
resp = client.responses.create(
    model="gpt-6-luna",
    reasoning={"effort": "medium"},   # none, low, medium (default), high, xhigh, max
    text={"verbosity": "high"},       # 7% longer in our runs; did not change completeness
    input=prompt,
)
print(resp.output_text)
usage = resp.usage
print(usage.output_tokens, usage.output_tokens_details.reasoning_tokens)

output_tokens includes the reasoning tokens, so the visible answer is the difference between the two numbers.

How much does GPT-6 Luna save over GPT-5.6 Luna?

For the same document, GPT-6 Luna cost 49% less than GPT-5.6 Luna: $0.81 against $1.61 per 1,000 documents with the parts named. The price cut alone would have given more. GPT-5.6 Luna’s tokens repriced at GPT-6 Luna’s rates come to $0.68, 58% less, and GPT-6 Luna’s answers cost 20% more than that, because it reasoned twice as much (537 reasoning tokens per answer against 271) and its total output grew from 1,284 to 1,558 tokens.

On goal-only prompts the bill looks better, $0.54 against $1.66 per 1,000 documents (68% less), but most of the extra saving is GPT-6 Luna writing half as much, which is a saving only if the other half was not needed.

Speed is about the same. GPT-6 Luna generated 115 output tokens per second of total request time against 125, and a goal-only document took a median of 8.1 seconds against 11.1 because there was less to write.

Do our results match other published tests?

Our results agree with the public benchmarks where they overlap, and add what is shorter, what is missing and what fixes it.

QuestionPublished elsewhereOur measurementVerdict
General capability, GPT-6 Luna against GPT-5.6 LunaArtificial Analysis 38 against 37; Vals 51.2% against 51.7%complete on 74 against 69 of 80 with the parts named, a tieAgrees: level
Document qualityGDPval-AA v2.1 Elo: 25 lower at max, 129 higher at mediumat medium, 0.55x the words on goal-only prompts and the incident duration missing from 9 of 20 postmortems against 2Mixed: neither shows a clear loss at the default effort; ours shows the shorter default and the one item it costs
Output tokens51K against 41K per index task at max, 24% more1,558 against 1,284 per document at medium, 21% moreAgrees
Cost61% less per index task; 48% less per Vals test49% less with the parts named, 68% less on goal-only promptsAgrees

What did we observe, and which should you use?

Use GPT-6 Luna wherever you control the prompt, and name the parts you need; keep GPT-5.6 Luna only where prompts cannot be changed and readers expect long documents. Three observations lead there:

  1. GPT-6 Luna writes to the request. A goal gets a short document, and a list of parts gets a full one that matches GPT-5.6 Luna’s.
  2. What it drops is narrow. Across four document types, the only expected item missing more often was the incident duration.
  3. The saving is the price cut. It uses 21% more output tokens for the same document, so the 49% saving is below the 58% that the list prices alone would give.
WorkloadPickNumbers
Machine-read output: classification, extraction, routingGPT-6 Luna at the lowest effort your accuracy allowslist prices 50% (input) and 58% (output) lower; answer length does not matter when a program reads it
Documents from a template or a prompt you writeGPT-6 Luna with the sections, counts and lengths in the promptcomplete on 74 of 80 against 69; $0.81 against $1.61 per 1,000
Documents from end-user prompts you cannot editGPT-6 Luna if you can add the expected parts to the request; otherwise GPT-5.6 Lunagoal-only: 384 against 703 words; duration missing from 9 of 20 postmortems against 2
Rubric-scored deliverablesPut the rubric in the prompt; do not rely on text.verbosityverbosity high: 52 against 51 of 60, 7% more words

For documents that go to customers, check the required elements in code before sending, whichever model wrote them.

FAQ

Is GPT-6 Luna worse than GPT-5.6 Luna? GPT-6 Luna was as complete as GPT-5.6 Luna when the prompt named the required parts (74 against 69 of 80 documents, a tie) at half the cost. It writes about half as much when a prompt only states a goal, and on those prompts it left the incident duration out of postmortems more often.

Does text.verbosity: "high" make GPT-6 Luna write as much as GPT-5.6 Luna? text.verbosity: "high" added 7% more words on GPT-6 Luna in our runs, which leaves it at about 58% of GPT-5.6 Luna’s length, and completeness did not change. Naming the parts in the prompt brought it to the same length.

Did GPT-6.1 fix GPT-6 Luna’s short answers? OpenAI’s 6.1 update covers only the Sol tier: GPT-6.1 Sol writes at GPT-5.6 Sol’s length again, and GPT-6 Luna is unchanged since its 2026-09-22 release.

Related measurements: GPT-6.1 Sol vs Claude Sonnet 5.5, flash-tier LLMs compared, and GPT-5.6 cost levers.

← Back to blog