Best Flash LLM API 2026: GLM Best Value, Gemini 3.8 Strongest
Contents
- How do the nine compare, and how did we test them?
- Which vendors’ models count as flash, and should you move to the newest?
- What do the nine cost?
- What can each one take in, and can you turn thinking off?
- What does a whole task cost, not a token?
- Which flash model should you pick?
- Which flash models fail at their default?
- Is a flash model enough, or do you need the bigger one?
- What breaks when you put a flash model in production?
- How Synthorai handles it
- FAQ
The best flash-tier LLM API in September 2026 depends on what you are buying: GLM 5.3 Flash gives the most benchmark score per dollar, Gemini 3.8 Flash is the strongest all-rounder and the one for audio and video, Claude Sonnet 5 invents the fewest answers at the highest price, and Seed 2.0 Mini is the cheapest model that passed every check we ran. Flash tier means each vendor’s small, fast model. We evaluated sixteen models, found that every vendor’s current model beats its previous small one, and compare the nine current ones on list price, cache pricing, input types, thinking controls, independent benchmarks, cost per task and what breaks in production.
TL;DR
- Move to each vendor’s current model: Claude Sonnet 5, GPT-5.6 Terra and Gemini 3.8 Flash cost more per token than Haiku 4.5, GPT-5.4 mini and Gemini 3.5 Flash-Lite, and 2.5 to 3.9 times less per correct answer.
- Best value: GLM 5.3 Flash, an Artificial Analysis index of 41.9 at $0.25 per task.
- Strongest and best multimodal: Gemini 3.8 Flash.
- Fewest invented answers: Claude Sonnet 5, at the highest cost per task.
- Cheapest: Seed 2.0 Mini at $0.10 / $0.40.
How do the nine compare, and how did we test them?
Prices and features come from vendor pages, benchmarks from independent leaderboards that run every model on one harness, and our own check is an admission test, not a ranking. The leaderboards are the Artificial Analysis (AA) Intelligence Index v4.3 and its non-hallucination rate (how often a model declines instead of guessing), LiveBench, the Vals Index, and Arena’s text and WebDev votes. Our check is 11 tasks with a checkable number (digit counts, prime sums, grid paths, coin change, a modular power, a knapsack and similar), 3 runs each, plus five questions about a company, an institute, a town, an alloy and a prize that do not exist, read in full to see whether the model declines or invents. Thinking means the reasoning tokens a model writes before its answer, billed as output and set per request with a field such as reasoning_effort; every model ran at its default and at every thinking setting it accepts.
| Model | AA index | AA non-hallucination | LiveBench | Vals Index | Arena text / WebDev | Our check: default / best | Invented at default |
|---|---|---|---|---|---|---|---|
| GPT-5.6 Terra | 42.3 at max | 12.1 | 77.9 | 59.6 | 1466 / 1521 | 31 / 33 | 2 of 5 |
| GLM 5.3 Flash | 41.9 | 72.4 | 71.6 | 47.2 | 1475 / 1607 | 31 / 31 | 0 of 5 |
| Gemini 3.8 Flash | 41.2 at high | 44.8 | 75.8 | 62.3 | 1493 / 1568 | 32 / 33 | 0 of 5 |
| Qwen3.8 Flash | 39.9 | 54.7 | 76.2 | not listed | not listed / 1635 | 33 / 33 | 2 of 5 |
| DeepSeek V4.1 Flash | 39.5 at max | 3.5 | 81.1 | 57.9 | not listed / 1614 | 33 / 33 | 4 of 5 |
| Claude Sonnet 5 | 38.4 at max | 60.6 | 76.0 | 59.6 | 1461 / 1537 | 33 / 33 | 0 of 5 |
| GPT-5.6 Luna | 37.5 at max | 7.4 | 73.6 | 59.9 | 1453 / 1519 | 31 / 32 | 4 of 5 |
| Hy3 | 25.8 | 25.9 | not listed | not listed | 1456 / 1513 | 33 / 33 | 0 of 5 |
| Seed 2.0 Mini | not listed | not listed | not listed | not listed | not listed | 33 / 33 | 0 of 5 |
Sources, read 2026-09-17: Artificial Analysis, LiveBench, Vals Index, Arena text and Arena WebDev; Arena scores are Elo-style ratings where a 20-point gap is small. Qwen3.8 Flash is listed there as Qwen3.8-Flash-Next, the open-weight model behind the API. Measurements ran on 2026-09-16 and 17, with two models rerun on the second day to confirm the days compare.
No model leads every board: Terra tops the index, DeepSeek tops LiveBench, Gemini 3.8 Flash tops Vals and Arena text, Qwen3.8 Flash tops Arena WebDev. The AA non-hallucination rate and our five invented questions agree on who guesses: DeepSeek V4.1 Flash, Luna and Terra rarely decline, GLM 5.3 Flash and Sonnet 5 usually do. Seed 2.0 Mini has no independent score anywhere, so our check is the only evidence for it.
Which vendors’ models count as flash, and should you move to the newest?
Every vendor’s newest small model, and yes. Google, DeepSeek, Alibaba, Z.ai, ByteDance and Tencent label theirs Flash, Mini or a single cheap model. Anthropic’s current generation is Fable, Opus and Sonnet, and Haiku 4.5 is a generation behind (October 2025, retirement not before October 15, 2026), so we treat Claude Sonnet 5 as Claude’s flash tier. OpenAI’s docs map GPT-5.6 Terra to its old mini tier and GPT-5.6 Luna to nano. Sonnet 5 ($2 / $10) and Terra ($2 / $12) cost several times more per token than the other seven and sit in their own price band below.
We also evaluated the previous small model from four vendors, scoring each on cost per correct answer: everything spent on the 33 runs divided by correct answers, so a model that is right half the time pays twice per answer. The current model won every pair, and in the three pairs from the TL;DR it cost more per token but less per correct answer:
- Claude Sonnet 5 over Claude Haiku 4.5. Haiku reached 32 of 33 only at maximum thinking; Sonnet 5 got 33 of 33 at every thinking-on setting and cost 2.9 times less per correct answer at
max, its cheapest. Both declined all five invented questions. - GPT-5.6 Terra over GPT-5.4 mini. Mini scored 9 of 33 at its default and needed
highfor 33; Terra atlowgot 33 for 2.5 times less. Terra invented two answers where mini declined all five. - Gemini 3.8 Flash over Gemini 3.5 Flash-Lite and Gemini 3.7 Flash. Against Flash-Lite, 3.9 times less per correct answer, and it declined the three invented questions Flash-Lite answered; Flash-Lite keeps a role only for single-step extraction. Against 3.7 Flash the list price is the same and 3.8 costs about 10% more per call, so that upgrade buys capability, not savings.
- GPT-5.6 Luna over GPT-5.4 nano at about a third of the cost per correct answer for the same 31 of 33, and Qwen3.8 Flash over Qwen3.6 Flash and Qwen3.5 Flash, which think about ten times as much and cost 6 to 27 times more per correct answer.
The rest of this guide covers the nine current models only.
What do the nine cost?
Three bands by output list price, with the bar showing our measured cost per correct answer at each model’s best setting, so a cheap list price and an expensive thinking habit show up together. The label carries the list price and the cache read price, what a repeated prompt prefix costs as a share of the input price.
The band is a weak predictor of the bill: GPT-5.6 Terra, in the most expensive band, cost less per correct answer than Gemini 3.8 Flash and only about a quarter more than Seed 2.0 Mini, because the thinking a model spends moves cost as much as the list price. Prices also move often and vary by time and place: Google’s 3.x Flash prices double on January 1, 2027 (Gemini 3.8 Flash to $1.50 / $7.50), DeepSeek charges half outside 01:00 to 04:00 and 06:00 to 10:00 UTC on weekdays, OpenAI cut GPT-5.6 Luna by 80% in July, and Qwen3.8 Flash costs 19 to 25% less outside Alibaba’s Singapore region. Three current small models were not available on Synthorai on the test days and are not measured: Mistral Small 4 ($0.15 / $0.60), StepFun Step 3.7 Flash ($0.20 / $1.15) and Amazon Nova 2 Lite ($0.30 / $2.50).
What can each one take in, and can you turn thinking off?
Eight of the nine accept images, four take video, two take audio, and Hy3, Tencent’s Hunyuan 3, reads only text. The last column is whether JSON extraction against a strict schema came back with correct values in 8 runs of an invoice extraction.
| Model | Input | Context / max output | Open weights | Thinking off switch | Thinking default | Schema JSON |
|---|---|---|---|---|---|---|
| Claude Sonnet 5 | text, image | 1M / 128K | no | yes | adaptive, high | 8/8 via tool call |
| GPT-5.6 Terra | text, image | 1.05M / 128K | no | yes | medium | 8/8 |
| Gemini 3.8 Flash | text, image, audio, video, PDF | 1M / 64K | no | no | medium | 8/8 |
| GPT-5.6 Luna | text, image | 1.05M / 128K | no | yes | medium | 8/8 |
| DeepSeek V4.1 Flash | text, image | 1M / 384K | MIT | yes | high | json_object only, 8/8 |
| Seed 2.0 Mini | text, image, audio, video | 256K / 128K | no | yes | medium | 8/8 |
| Qwen3.8 Flash | text, image, video | 1M / 128K | Qwen license | yes | xhigh | 3/8, wrong integers |
| GLM 5.3 Flash | text, image, video, file | 1M / 128K | MIT | no | max | json_object only, 8/8 |
| Hy3 | text | 256K / 128K | Apache 2.0 | yes | high | json_object only, 7/8 |
“Via tool call” means the schema went in as a tool’s input and Claude was forced to call it. “json_object only” means a strict schema did not come back right on our request path, and {"type": "json_object"} with the keys named in the prompt did. Audio input means Gemini 3.8 Flash or Seed 2.0 Mini; output above 128K means DeepSeek V4.1 Flash; weights you can host yourself means DeepSeek, GLM, Hy3 or Qwen3.8 Flash.
What does a whole task cost, not a token?
A call costs the price times the tokens the model chooses to spend, and in this tier the second factor varies more than the first. The fairest public measure is Artificial Analysis’s cost to run one task of its index, with thinking tokens counted; it exists for eight of the nine.
GLM 5.3 Flash scores within half a point of GPT-5.6 Terra for a fifth of the cost per task. Qwen3.8 Flash has almost the same list price as GLM but costs half as much again per task, because it wrote 240 million output tokens to run the index against GLM’s 180 million. Claude Sonnet 5 at max costs $5.09 per task against $1.79 at its default high, for six more index points. Our single calls show the same effect: Seed 2.0 Mini has a lower list price than Qwen3.8 Flash but cost 3.4 times more per correct answer, because it thought a median of 2,810 tokens per task against 530.
Cache reads cost 2% of input at DeepSeek, about 10% at Google, OpenAI, Anthropic and Alibaba, 20% at Z.ai and ByteDance and 25% at Tencent. For agent and RAG traffic (retrieval-augmented generation, where retrieved documents are pasted into every prompt), the read price is the bill: a cached 100K-token prefix costs $0.0006 per call at DeepSeek V4.1 Flash against $0.02 on Sonnet 5 or Terra, and OpenRouter reported that 90% of the tokens V4.1 Flash served in one day after launch were cache reads (post). Batch tiers, answered within hours, halve the Gemini 3.8 Flash, GPT-5.6 Luna, Terra and Sonnet 5 prices; Qwen3.8 Flash, GLM 5.3 Flash and Hy3 have none.
Which flash model should you pick?
GLM 5.3 Flash for value, Gemini 3.8 Flash for strength and multimodal input, Claude Sonnet 5 for the fewest invented answers, Seed 2.0 Mini for price.
| Question | Pick | Why |
|---|---|---|
| Best value | GLM 5.3 Flash | index 41.9, second only to Terra, at $0.25 per task; highest non-hallucination rate of the nine (72.4); $0.15 / $0.50; image, video and file input; MIT weights |
| Strongest all-rounder | Gemini 3.8 Flash | tops Vals and Arena text of the nine, 1.1 index points behind Terra for 11% less per task; 33 of 33 at low, declined every invented question |
| Best multimodal | Gemini 3.8 Flash | audio, video and PDF input, audio priced as text; schema JSON 8 of 8 |
| Fewest invented answers, fast | Claude Sonnet 5 | 33 of 33 at every thinking-on setting, declined all five invented questions with thinking on and off, first answer token in 2.2 seconds; the most expensive per task, so use it where a made-up answer costs more than the call |
| Cheapest | Seed 2.0 Mini | $0.10 / $0.40; 33 of 33 at default, schema JSON 8 of 8, declined all five invented questions, audio and video input |
Three models miss a verdict for one reason each. GPT-5.6 Terra has the top index and 33 of 33 at low for $0.00199 per correct answer, but invented two of five answers at $12 per million output tokens. DeepSeek V4.1 Flash tops LiveBench and has the cheapest cache, but reads only text and images and invented four of five answers with thinking on. Qwen3.8 Flash had the cheapest perfect score in our check, $0.00046 per correct answer, and tops Arena WebDev, but invented two answers, wrote wrong integers under a strict schema, and took almost three minutes on a 400-word explanation.
Route by request rather than picking one model: closed tasks with a checkable answer to the cheapest model that passes (GLM 5.3 Flash, Qwen3.8 Flash, Hy3), open factual questions to a model that declines (GLM 5.3 Flash, Sonnet 5, Gemini 3.8 Flash), long agent runs to the strongest model you can afford.
Which flash models fail at their default?
Two of the nine, and both pass once open questions go to a setting that declines. The admission floors, applied at the setting a model ships with: at least 30 of 33 on the tasks and no more than two of five invented answers.
| Model | At its default | Change | After the change |
|---|---|---|---|
| DeepSeek V4.1 Flash | invented 4 of 5 (“50 employees”, “1790 °C”, a Simpsons character as prize winner) | thinking off for open questions | declined 5 of 5, but 21 of 33 on the tasks, so send only open questions there |
| GPT-5.6 Luna | invented 4 of 5 (“approximately 20 employees”, “1974”, “1,400 °C”) | thinking off for open questions | declined 4 of 5, but 7 of 33 on the tasks, so the same routing applies |
GPT-5.6 Terra and Qwen3.8 Flash sit on the floor with two invented answers each (“0 employees” and “about 1,455 °C” for Terra), so treat them the same way for open questions.
Is a flash model enough, or do you need the bigger one?
For bounded work, yes. On the shorter agent benchmarks, where the model works in a shell against real repositories, DeepSeek V4.1 Flash beats DeepSeek V4 Pro on Terminal-Bench 2.1 (90.6 against 87.9) and DeepSWE (74.2 against 62.7), and Gemini 3.8 Flash scores 89.4 on Terminal-Bench 2.1 against 89.1 for Claude Opus 5. All nine completed our two-turn tool-calling loop 3 of 3 times.
The gap reopens on long tasks and long contexts. On Terminal-Bench 4.0, Google’s own table has Gemini 3.8 Flash at 19.1 against 51.8 for Opus 5. On OpenAI’s multi-needle retrieval test between 512K and 1M tokens, GPT-5.6 Luna scores 41.3 against 72.5 for GPT-5.6 Terra, so a 1M window does not mean 1M tokens of reliable recall on the smallest models. Keep flash models on extraction, classification, short tool steps and code changes of a known shape, and a larger model on planning long agent runs and questions over very long documents.
What breaks when you put a flash model in production?
Defaults and request shapes, more than model ability.
- Two models cannot turn thinking off. Gemini 3.8 Flash and GLM 5.3 Flash returned a 400 for every off setting we tried.
- Defaults sit at both extremes. GLM 5.3 Flash defaults to
maxand Qwen3.8 Flash toxhigh(Alibaba); on our five invented questions Qwen3.8 Flash spent a median of 11,680 reasoning tokens. Google counts thinking towardmax_output_tokens, so a tight cap can return a truncated or empty answer that is still billed (thinking guide). - Claude uses its own controls. Claude Sonnet 5 sets thinking depth with
output_config.effortand rejects the olderbudget_tokens(effort);reasoning_effortis not one of its parameters. - Tool loops must send reasoning back. DeepSeek’s thinking mode expects the
reasoning_contentof previous turns (thinking mode), and Gemini attaches thought signatures to function-call parts. Frameworks that rebuild the message history drop both. - Strict JSON is not uniform. Qwen3.8 Flash enforced the schema’s types but wrote wrong integers in 5 of 8 runs (
-561994line items). Usejson_objectwith the keys named in the prompt on Qwen, GLM, Hy3 and DeepSeek, and a tool call on Claude. - Thinking changes honesty. DeepSeek V4.1 Flash invented four of five answers with thinking on and none with it off.
- Model ids move.
deepseek-v4-flashhas served V4.1 Flash since September 10 (pricing) and Claude Haiku 4.5 may retire from October 15 (deprecations). Pin a dated snapshot where one exists, such asseed-2-0-mini-260428. - Third-party hosts of open weights vary. DeepSeek, GLM, Qwen and Hy3 are served by many providers with different quantization and cache behaviour, and in September one reseller was found routing paid requests to a cheaper model (Hacker News). Compare your provider’s answers against the vendor’s own API.
Speed follows the default thinking setting more than model size. A streamed 400-word explanation at each default, 3 runs each, on 2026-09-17:
| Model, default setting | First answer token | Total time | Output tokens per second |
|---|---|---|---|
| Claude Sonnet 5 | 2.2 s | 10.5 s | 70 |
| Gemini 3.8 Flash | 10.7 s | 14.3 s | 114 |
| GPT-5.6 Luna | 24.9 s | 26.1 s | 89 |
| GPT-5.6 Terra | 33.3 s | 34.2 s | 63 |
| GLM 5.3 Flash | 36.7 s | 39.3 s | 128 |
| DeepSeek V4.1 Flash | 40.4 s | 40.6 s | 158 |
| Seed 2.0 Mini | 61.3 s | 61.5 s | 81 |
| Hy3 | 123.0 s | 134.5 s | 35 |
| Qwen3.8 Flash | 159.9 s | 165.8 s | 72 |
Tokens per second count reasoning and answer over the whole call. Sonnet 5’s adaptive thinking spent nothing on a writing task and started the answer at once; Qwen3.8 Flash spent a median of 10,697 reasoning tokens on it, so an open-ended prompt can take minutes on a model that answers closed tasks cheaply.
The setting is one field on an OpenAI-compatible request, and the thinking spend comes back in usage:
from openai import OpenAI
client = OpenAI(base_url="https://synthorai.io/v1", api_key="sk-syn-...")
r = client.chat.completions.create(
model="gemini-3.8-flash", # or glm-5.3-flash, gpt-5.6-terra, ...
reasoning_effort="low",
max_tokens=32768,
messages=[{"role": "user", "content": "Compute 7 raised to the power 222, modulo 1000. Reply with a single integer."}],
)
u = r.usage
print(r.choices[0].message.content, u.completion_tokens, u.completion_tokens_details.reasoning_tokens)
For Claude Sonnet 5, send the same request to the Messages API with output_config: {"effort": "low"}.
How Synthorai handles it
Every model above is one model id on the same endpoint, OpenAI-compatible or Anthropic-compatible, and the model compare page puts any two side by side on price, cache, input types, context and vendor benchmarks. Current prices for all nine are on the model price comparison.
FAQ
Is Claude Sonnet 5 a flash model? Anthropic does not call it one, but it is the smallest model in Claude’s current generation, since Claude Haiku 4.5 dates from October 2025 and may retire from October 15, 2026. Claude Sonnet 5 costs $2 / $10 against Haiku’s $1 / $5, yet cost 2.9 times less per correct answer on our check.
Is the Gemini Flash API free? Gemini 3.8 Flash has a free tier in Google AI Studio, but Google uses free-tier content to improve its products (pricing) and allows only the paid service for apps serving the EEA, Switzerland or the UK (terms). The paid tier is $0.75 / $3.75 until December 31, 2026, then $1.50 / $7.50.
Should I use Gemini Flash or Flash-Lite? Gemini 3.8 Flash unless every request is a single-step extraction: it scores 41.2 on the Artificial Analysis index against 23 for Gemini 3.5 Flash-Lite, cost 3.9 times less per correct answer on our check, and declined all five invented questions where Flash-Lite answered three. Flash-Lite is $0.30 / $2.50 and answers in about 4 seconds.
GPT-5.6 Luna or GPT-5.6 Terra?
GPT-5.6 Luna for volume, GPT-5.6 Terra for harder reasoning. Luna is $0.20 / $1.20 and scored 31 of 33 at its default for $0.00030 per correct answer; Terra is $2 / $12, scored 33 of 33 at low for $0.00199, and has the top index score of the nine. Both invented answers at their default, Luna four of five and Terra two.
Which flash model is best for coding? Qwen3.8 Flash, DeepSeek V4.1 Flash and GLM 5.3 Flash lead Arena WebDev (1635, 1614 and 1607), DeepSeek leads LiveBench’s agentic coding category (77.3), and DeepSeek and Claude Sonnet 5 lead Vals’s Vibe Code Bench (84.7 and 81.3).
Related: DeepSeek V4.1 Flash API cost, GLM 5.3 API cost, Gemini 3.7 Flash API cost, prompt caching across providers, best LLM by use case.