Token Cost Radar

Token Cost Radar

October 7, 2026

Today's token-cost story is splitting into two layers. At the market layer, Reuters notes that token prices have fallen roughly 40% while questioning whether usage is growing fast enough to offset the decline for AI labs. At the product layer, Cohere is putting token quotas, rate limits, and organization-wide caps directly into its enterprise agent control plane, while Mistral has launched a trillion-parameter mixture-of-experts model at $1.36 per million input tokens and $4.18 per million output tokens, with a two-week launch discount cutting those rates in half. Meanwhile, fresh MCP measurements show that tool-schema cost depends heavily on the harness rather than the protocol alone. The arc from tokenmaxxing through tokenminimizing to token yield is becoming allocation economics: which model, context, tool schema, and unit of compute deserves to be resident for this particular piece of work?

Top Developments (Last 24 Hours)

1What happens if token prices keep falling but usage does not rise fast enough?

Reuters reports October 7 that AI labs are betting lower token prices will stimulate enough additional usage to compensate for declining unit prices, a version of Jevons' paradox. Reuters cites recent data showing token prices down roughly 40%, while questioning whether demand is currently accelerating fast enough to offset that decline. The issue moves token economics beyond customer savings: falling inference prices also change the revenue model of the companies supplying those tokens.

Reuters ↗

2Mistral launches a 1T model at $1.36 input and $4.18 output

Mistral opened Mistral Large 4 to public API preview October 6. The mixture-of-experts model has 1.05 trillion total parameters, 49 billion active parameters per token, and a 1 million-token context window. Standard pricing is $1.36 per million input tokens, $0.14 cached input, and $4.18 output. Mistral's changelog says launch pricing is 50% off for two weeks, bringing those rates to $0.68, $0.07, and $2.09 respectively. The company says downloadable weights will follow later this month.

Mistral ↗

3Enterprise agent budgets move into the control plane

Cohere's North 2 adds granular cost controls, rate limits, user quotas, consumption tiers based on requests and token rates, and organization-wide caps. Administrators can monitor usage down to individual users and agents and control which models teams can use. The launch is a useful AI FinOps milestone because token governance is becoming an execution-time policy rather than a dashboard teams inspect after the invoice arrives.

Cohere ↗

4MCP schema overhead gets measured instead of merely complained about

ImportStatic published measurements October 6 across four real MCP servers and found that GitHub's 46 default tools consume 11,207 tokens of tool definitions under its test tokenizer. One issue-writing tool alone costs 807 tokens to describe. More importantly, the study argues that MCP overhead depends on the harness: clients that defer or selectively load tool definitions avoid paying the entire schema cost on every interaction. The protocol exposes capabilities, but the harness decides how much of that capability surface becomes resident context.

ImportStatic ↗

From Tokenmaxxing to Allocation Economics

Tokenmaxxing is the practice of maximizing AI consumption as a proxy for productivity. Tokenminimizing is the practice of reducing avoidable token consumption while preserving the required result, with tokenminning and tokenmining also appearing as variant spellings in the efficiency conversation. Modelmaxxing means matching work to the least expensive model capable of performing it reliably. Token yield measures useful work relative to token consumption. Today's model, governance, and MCP stories push the arc toward allocation economics: deciding what deserves expensive inference before optimizing the price of that inference.

Mistral pricing

Mistral Large 4's model card makes caching and batch processing explicit pricing dimensions. Standard rates are $1.36 input, $0.14 cached input, and $4.18 output per million tokens, while the current two-week launch tier is $0.68, $0.07, and $2.09. A 10-fold gap between fresh and cached input again makes context reuse an economic primitive rather than a minor API feature.

Mistral Docs ↗

Current model pricing

BenchLM's October 6 board tracks 173 paid models across 32 providers. It lists a decision model at $0.04 per million input tokens with free output as the cheapest API, MiMo-V2.6-Pro at $0.43 input and $0.87 output as the cheapest model clearing its production-grade threshold, and Claude Sonnet 5.5 at $2 and $10 as the cheapest model clearing its frontier threshold. Cheapest token, cheapest adequate model, and cheapest frontier model remain three different questions.

BenchLM ↗

Agentic coding economics

A pricing analysis updated October 6 models an illustrative coding-agent session with 10 million input tokens, 90% cache hits, and 300,000 output tokens. Its estimated bill ranges from $0.34 on GPT-6 Luna and $0.71 on DeepSeek V4.1 Flash to $5.90 on GPT-6.1 Sol and $34 on GPT-6 Astra. The assumptions are illustrative rather than benchmark results, but they show why cache-read pricing and model routing can dominate the economics of context-heavy agents.

The Vibelog ↗

DeepSeek

DeepSeek V4.1 Flash remains an unusually clear example of schedule-aware and cache-aware inference pricing. Current independently verified rates put off-peak fresh input at $0.15 per million tokens, cache hits at $0.003, and output at $0.60, with peak rates twice as high. At off-peak rates a cached input token costs one-fiftieth as much as a fresh one, turning workload timing and prefix reuse into routing variables.

DeepSeek Pricing Guide ↗

MCP tool surfaces

A September measurement of a large routed MCP catalog found that exposing 47 hosted definitions consumed 9,286 tokens, while representing the same catalog as one MCP tool per skill would have consumed an estimated 257,047 tokens. The two routing tools alone used 1,070 tokens. The specific implementation is one vendor's measurement, but the architectural lesson is broader: on-demand capability discovery can make the accessible tool surface much larger than the resident tool surface.

ToolRouter ↗

Research Watch

Harnessing LLMs as Agents: What Does It Cost?

This October 1 paper introduces the Language Model Agent Machine, a resource model that charges agent harnesses for bounded context, persistent memory, tool use, verification, communication, and repeated execution instead of treating the underlying model call as the whole computational cost. The authors derive bounds for context-memory traffic, recomputation, and verification and test several predictions on GPT-6 Astra.

Why it matters: Agent FinOps needs a cost model for the harness, not merely the model. Context movement, memory access, repeated calls, verification, and checkpointing can all consume resources even when the nominal model price is unchanged.

arXiv ↗

What Does a Token Cost? A Mixture-of-Agents Measurement of Sufficient Per-Token Compute

A panel of fifteen models found that a 0.5B model could reproduce 92% to 95% of reference tokens across three benchmarks, while the hardest 10% of tokens accounted for 64% to 80% of estimated FLOPs. Using the resulting difficulty map for routing reduced projected MATH-500 latency from 7.59 to 5.12 seconds while slightly improving accuracy, and reduced draft-token use by 32.6% at similar accuracy.

Why it matters: Modelmaxxing can become finer than routing whole requests. If token difficulty is highly uneven, sufficient compute suggests allocating expensive capacity only where individual pieces of generation actually require it.

arXiv ↗

Total Cost of Agency: Exact Attribution of Memory Injection Cost in Multi-Agent LLM Workflows

On a 200-task enterprise benchmark using real model APIs, the authors attribute roughly 12% of total billed cost to memory-injection tokens and find the share rising to 27.6% at workflow depth six. Reducing retrieval-window capacity from 32 entries to 2 lowered injected tokens by 28.7% with accuracy change within seed-level variation in their experiment.

Why it matters: Memory is not free merely because the model did not generate it. Retrieved state becomes billable input, so memory-window sizing belongs in token budgeting alongside model selection and output limits.

arXiv ↗

Token-Efficient Data Reasoning Agents via Adaptive Structuring of Unstructured Data

This work argues that repeatedly reopening unstructured documents forces agents to repurchase the same evidence in context. On FanOutQA, reasoning over an ideal pre-structured store was 28 times cheaper than repeatedly reasoning over source documents. Its adaptive data-cracking method structures useful information as a byproduct of earlier reasoning and cut cost by 53% in an extended evaluation while preserving accuracy.

Why it matters: Token yield improves when knowledge that inference already paid to discover becomes reusable structure. The cheapest future context may be information the agent never needs to reread as prose.

arXiv ↗

Phrase of the Day

“Allocation economics”

Allocation economics is the practice of deciding where expensive model capacity, context, memory, and tool definitions actually improve an AI workload before optimizing the unit price of those resources.

  1. AI adoption
  2. Tokenmaxxing
  3. Token shock
  4. Tokenminimizing
  5. Modelmaxxing
  6. Token discipline
  7. Token yield
  8. Allocation economics

The likely winners are systems that allocate scarce model attention rather than merely purchasing cheaper attention, combining capability-aware routing with cache reuse, bounded memory, on-demand tools, and execution-time budget controls.

The cheapest token is still expensive when it had no reason to be in the room.

Cohere ↗

The jCodeMunch read

Today's allocation-economics and tool-surface stories land on the same principle: availability does not require residency. jCodeMunch reduces code-reading tokens through tree-sitter symbol retrieval and byte-precise context, keeping relevant code evidence available without making repository-scale context the default.

See how the 95%+ cut is measured →

← All editions