Posts tagged pricing
17 posts about pricing.
-
Voice Agent API Cost: a 10-Minute Call Runs $0.04 to $0.57
Every leg measured: STT at $0.002-0.006/min, TTS at $0.004-0.034 per audio-minute, LLM turns, and GPT Realtime. Cascade vs speech-to-speech, priced per call.
-
DeepSeek V4 Pro GA vs Preview, Measured: 18-62% Less Thinking
deepseek-v4-pro-0813 vs the preview build on identical tasks: 18-62% fewer reasoning tokens, a fixed 8,192-token runaway, and a lost 'I don't know'.
-
Gemini 3.7 Flash API Cost, Measured: Tasks Bill 2.5-8x Less Than 3.6
Day-one probes of gemini-3.7-flash: half-price intro until Dec 31, thinking down 26-77% vs 3.6, off-switch gone, prefill 400s. Task bills drop 2.5-8x.
-
How Many Tokens Is an Image? 15 APIs Measured, Same Icon 6 to 1,298
The same 64px icon bills 6 tokens on GPT-5.6 and 1,298 on ByteDance Seed. Three billing schemes, three broken doc rules; dollars follow the rate, not the image.
-
LLM Thinking Controls: What 13 Models Accept, Ignore, or Enforce
reasoning_effort and thinking_budget across 13 LLM APIs: three vendors enforce budgets to the token, others silently ignore them, a cap of 16 can burn 97.
-
Web Search API and Web Fetch API: How They Work and What $0.01 Buys
How server-side web search and web fetch bill: a $0.01 fee, 1,500-3,100 injected result tokens at your model's rate, and a turn two that can re-search.
-
Qwen-Image 3.0 API, Measured: 9 Scenes for One $0.03 Image
Day-one probes of qwen-image-3.0: $0.03 per image with prompts unbilled even at 5,000 tokens, a real nine-in-one trick, and small print that misspells itself.
-
DeepSeek V4 Flash API Cost: Thinking Mode Corrupts Strict JSON
Day-one 0731 measurements: $0.14/$0.28 list, a 1,024-token cache page, and a defect: thinking plus strict json_schema corrupted integers in 8 of 13 runs.
-
Qwen 3.8 Max API Pricing: 16 Thinking Tokens Beat the Off Switch
Day-one measurements of qwen3.8-max at $2/$6: an effort dial that is really a budget cap, a 1/6 accuracy collapse with thinking off, and a 16-token fix.
-
MCP Tool Overhead, Measured: 26 Tools, $0.03 Every Call
Every MCP tool is re-billed as input tokens on every call: one small tool is 401 tokens on Claude, the GitHub server $0.03/call, and families differ 2.7x.
-
Prompt Cache Write Cost: When Does the 1.25x Premium Pay?
One measured agent suite, one premium: a 6% loss on RAG, an 83% saving on batch. The read:write ratio decides, and you can measure yours before the invoice.
-
Claude Opus 5 vs Opus 4.8, Measured: Same Price, 3x Apart
Opus 5 and Opus 4.8 share a $5/$25 rate card, yet Opus 5 billed 3.1x more on identical tasks out of the box. Where it goes, and the switch that closes it.
-
Speech-to-Text APIs: 14 Models, $0.002 to $0.016 a Minute
14 speech-to-text APIs behind one gateway: per-minute rates, who streams, what token-billed gpt-4o converts to, and the China ASR tier most comparisons skip.
-
Gemini 3.6 Flash: the Thinking Dial That Moves Cost 30x (Measured)
Gemini 3.6 Flash bills reasoning you never see, and one request setting swings the same task's cost up to 30x. Measured across five task types, with the catch.
-
Seedance API Pricing, Measured: the Video-Token Formula, Solved
Seedance bills W×H×(24s+1)/1024 video tokens; we solved it to the token. 720p bills as 1248×704, and 4k's cheaper rate costs 2.1x more per second. Measured.
-
Kimi K3 API Pricing, Measured: Turn Off the 'Always-On' Reasoning
Kimi K3's docs say reasoning can't be disabled. reasoning_effort:'none' works and cuts simple queries 6x. Measured: effort dial, cache floor, 9-language rates.
-
GPT Live API Pricing: gpt-realtime Speaking Costs 4x
GPT Live is the ChatGPT feature, not an API; behind it is gpt-realtime-2.1: $0.019/min listening, $0.077/min speaking, silence free, cached replay 1/80th.