Token Cost Radar

Token Cost Radar

September 29, 2026

Today's token-cost story is finally separating price per token from cost per task. Anthropic kept Claude Sonnet 5.5 at the same $2 input and $10 output rate as its predecessor but says it needs fewer tokens to finish the same work. At the same time, a Futurum report says agentic workloads can consume 10 to 100 times more tokens than simple inference, open-weight models are winning enterprise workloads on price-performance, and Chinese models processed more than four times the OpenRouter token volume of U.S. models last week. The vocabulary arc from tokenmaxxing through tokenminimizing to token yield is landing on a useful principle: the cheapest token is irrelevant if the workload needs far more of them to finish.

Top Developments (Last 24 Hours)

1What if the token price stays flat but the task gets 30% cheaper?

Reuters reports that Anthropic released Claude Sonnet 5.5 on September 28 at the same $2 per million input tokens and $10 per million output tokens as Sonnet 5. Anthropic says the new model needs fewer tokens to complete the same work. That distinction matters for AI budgeting: a model can lower cost per completed task without changing its nominal token rate.

Reuters ↗

2Agentic AI puts the per-token meter under pressure

A September 28 SiliconANGLE guest column examines a new Futurum Research report, sponsored by QumulusAI, which estimates that agentic AI can consume 10 to 100 times more tokens per task than simple inference. The report forecasts agent and reasoning inference growing 219% this year and argues that sustained production workloads can make purely variable per-token pricing difficult to predict. The figures are report estimates, not independent measurements, but the underlying FinOps problem is increasingly visible.

SiliconANGLE ↗

3Chinese models process 62.22 trillion tokens in one week

China Daily reports September 29, citing OpenRouter data, that Chinese models processed 62.22 trillion tokens from September 21 through September 27 versus 14.2 trillion for U.S. models. DeepSeek V4.1 Flash led with 19.6 trillion tokens, up 24% week over week, while GLM 5.3 Flash processed 16.3 trillion. The report attributes the shift partly to competitive pricing, open weights, rapid iteration, and price-performance for agent and tool-calling workloads.

China Daily ↗

4Corporate America routes more work toward cheaper open-weight models

The Financial Times reports September 28 that U.S. companies are increasingly using lower-cost open-weight models as AI infrastructure bills rise. The report says open-weight models have exceeded half of observed usage on some AI platforms, with DeepSeek, Zhipu, Mistral, and Nvidia among the model makers benefiting. The trend is modelmaxxing in practice: reserve premium closed models for workloads where their additional capability justifies the premium.

Financial Times ↗

From Tokenmaxxing to Task Efficiency

Tokenmaxxing is the practice of maximizing AI consumption as a proxy for productivity. Tokenminimizing is the practice of reducing avoidable token consumption while preserving the required result. Modelmaxxing means matching each workload to the least expensive model capable of performing it reliably. Token yield measures useful work relative to token consumption. Today's model releases and agent-cost coverage sharpen the arc: token efficiency increasingly means measuring how much inference a model needs to complete the job, not merely what one million tokens cost.

IFX

The IFX Inference Index closed September 28 at 83.87, down 0.62% from its previous reading. Its capability-adjusted board currently puts the cheapest blended rate meeting its frontier threshold at $3.375 per million tokens, its capable threshold at $1.6875, and its budget threshold at $0.4625. The structure captures modelmaxxing neatly: price the capability floor the workload actually requires.

IFX ↗

BenchLM

BenchLM's September 28 pricing index puts frontier token prices 81.3% below its March 2023 baseline, with a median flagship blended rate of $7 per million tokens across 22 active constituents. Its latest snapshot includes an 80% blended-price cut for GPT-5.6 Luna and a 50% cut for Gemini 3.6 Flash. Token deflation continues, but the index itself warns that model-mix changes can make the frontier cohort more expensive month to month.

BenchLM ↗

Surplus Intelligence

The September 29 marketplace snapshot records 2.574 trillion fresh input tokens, 2.021 trillion cache tokens, and 47.73 billion output tokens across 52.3 million requests over 28 days. Its latest seven full days of eligible traffic show an 82.8% mean realized discount versus direct-provider pricing. Cache volume equal to roughly 79% of fresh input again shows why raw token counts are a poor proxy for the actual bill.

Surplus Intelligence ↗

Flexera

Flexera's September 29 enterprise AI governance guide explicitly assigns AI spend, budgets, and cost allocation to finance and FinOps while recommending an inventory of employee AI tools, agents, SaaS features, and accounts. Its research says only 31% of organizations have visibility into AI software. Per-seat and per-tool cost control starts with knowing which seats and tools exist.

Flexera ↗

Anthropic tool-use research

Anthropic's tool-search work remains a useful benchmark for the tool-surface problem. Its published example found 58 tool definitions consuming roughly 55,000 tokens before the conversation began, with larger setups reaching 134,000. Retrieving tools on demand reduced token usage by 85% in its evaluation. As MCP catalogs grow, loading every available schema into every request increasingly looks like paying rent on tools the agent never touches.

Anthropic ↗

Research Watch

Total Cost of Agency: Exact Attribution of Memory Injection Cost in Multi-Agent LLM Workflows

This September study decomposes multi-agent cost into base prompt, inference, memory injection, miss penalty, and context accumulation. Across a 200-task enterprise benchmark using real model APIs, memory injection represented about 12% of full billed cost and reached 27.6% of controllable variable cost at workflow depth six. Reducing retrieval capacity from 32 entries to 2 cut injected tokens by 28.7% with accuracy movement within seed-level variation.

Why it matters: Tokenminimizing gets more actionable when the bill is attributed to architectural causes. Memory tokens look like ordinary input to the provider, but the application controls how much memory gets injected.

arXiv ↗

An Empirical Cost Attribution of Context-Compression Gateways in Multi-Turn Coding Agents

This empirical study separates tool-schema filtering, content compression, and history summarization as independent token-saving levers. Tool-schema filtering removed roughly 21,000 to 57,000 tokens from a typical turn and produced an immediate linear saving, while compressed file reads accumulated savings across later turns. The authors caution that single-shot compression quality does not by itself prove lower multi-turn agent cost.

Why it matters: The tool-surface tax compounds because schemas are repeatedly resident. Removing unused definitions can be more predictable than trying to squeeze every file read or conversation turn.

arXiv ↗

Toollery: Scaling LLM Agents to Thousands of Skills and Tools

Toollery treats large capability catalogs as a retrieval problem rather than placing every skill and tool specification into the model context. Evaluations span a roughly 79,000-capability benchmark, BFCL-V4 with more than 440 tools, and 3,396 proprietary requests across 220 tools. The framework retrieves a compact candidate set before final model selection and reports improved recall over ordinary specification retrieval in its evaluated settings.

Why it matters: On-demand tool loading attacks two costs at once: fewer schema tokens and fewer distractors. The model sees the small set of capabilities relevant to the current request instead of paying attention to the entire warehouse.

arXiv ↗

Token-Budget-Aware Pool Routing for Cost-Efficient LLM Inference

This serving study estimates total token demand before dispatch and routes short and long workloads into differently configured inference pools. On Azure and LMSYS traces, the authors report 17% to 39% fewer required GPU instances, corresponding to roughly $1.2 million to $2 million in annual savings at 1,000 requests per second. A Qwen3-235B case study projects $15.4 million in annual savings at 10,000 requests per second.

Why it matters: Token budgets can control infrastructure as well as prompts. Expected context footprint determines KV-cache demand and concurrency, making routing by token budget a direct lever on cost-to-serve.

arXiv ↗

Phrase of the Day

“Task efficiency”

Task efficiency is the amount of token consumption, inference cost, model calls, and other compute required to complete a useful unit of work at an acceptable quality level.

  1. AI adoption
  2. Tokenmaxxing
  3. Token shock
  4. Tokenminimizing
  5. Modelmaxxing
  6. Token discipline
  7. Token yield
  8. Task efficiency

The likely winners are systems that measure the complete execution path, then reduce the tokens, model calls, tool schemas, memory, retries, and premium-model steps required to finish the work without lowering acceptable quality.

Price per token tells you what the fuel costs. Task efficiency tells you whether the thing gets twelve miles to the gallon.

Reuters ↗

The jCodeMunch read

Today's task-efficiency and tool-surface stories point to the same discipline: do not make the model process information merely because it is available. jCodeMunch reduces code-reading tokens through tree-sitter symbol retrieval and byte-precise context, keeping relevant code evidence available without making repository-scale context the default.

See how the 95%+ cut is measured →

← All editions