Posts tagged llm-cost
3 posts about llm-cost.
-
LLM Token Usage: Why a 4-Token Answer Bills 217 Tokens
Measured across GPT-5.6, Claude Fable 5, Qwen3.7-max and five more families: reasoning dominates the output bill. How to read every usage field, and cap them.
-
Prompt Cache Minimums: The Docs Under-State by 1.4-2.4x
Vendors publish a prompt-cache token minimum. Measured across LLM families, auto-cache needs 1.4-2.4x more than the docs say; Claude's explicit cache is exact.
-
Which LLM Is Cheapest for Your Language? Tokenizer Costs Measured
GPT-5.5 bills fewest tokens for European languages, Kimi for Chinese, DeepSeek for Japanese; Claude Fable 5, Opus 4.8 and Sonnet 5 run 1.2-2.3x. Measured.