Token Cost Radar

Token Cost Radar

September 28, 2026

Today's token-cost story is about the denominator. SaaStr reports a median engineer now generates $213 a week in AI coding spend, while EY's new agentic-AI ROI framework argues that token bills alone cannot tell an enterprise whether the money produced value. Fresh pricing data also shows why the accounting is getting harder: Grok 4.7 doubles its token rates once prompts cross 200,000 tokens, DeepSeek remains dramatically cheaper, and cached traffic can dominate real agent workloads. The vocabulary arc from tokenmaxxing through tokenminimizing to token yield is converging on a more useful question: what did the tokens buy?

Top Developments (Last 24 Hours)

1What did your $213-a-week coding-agent bill actually buy?

SaaStr reports September 28 that the median engineer in the dataset it cites now generates $213 per week in AI coding spend. The article argues that companies increasingly need to connect token and seat costs to engineering output rather than treating adoption or raw consumption as success metrics. That moves AI cost governance directly toward return-on-token measurement.

SaaStr ↗

2Grok 4.7 makes context inflation visible on the rate card

Pricing verified September 28 lists Grok 4.7 at $2 input, $0.50 cached input, and $6 output per million tokens below 200,000 prompt tokens. Crossing that threshold doubles the standard rates to $4, $1, and $12 for the entire request. The faster serving tier costs more again. Long context is therefore no longer merely a capacity feature. It can change the marginal economics of every token in the request.

OrcaRouter ↗

3The current price board spans from ten-cent frontier input to fifty-dollar output

LLM Cost Hub's September 27 official-pricing refresh tracks 124 model rows. GPT-6 Luna is listed at $0.10 input, $0.01 cached input, and $0.50 output per million tokens, while GPT-6 Astra reaches $10 input and $50 output. GPT-6 Sol sits between them at $2 and $10. A 100-fold input-price spread inside one provider's current lineup makes workload-to-model matching an increasingly obvious budget control.

LLM Cost Hub ↗

4Kimi's agent subscriptions and API now expose two different budget models

A pricing guide updated September 28 lists Kimi consumer memberships from 49 to 699 yuan per month with agent-task quotas, while programmatic API use remains token-metered. Kimi K3 is listed at $3 input and $15 output per million tokens. The split illustrates an increasingly important FinOps distinction: agent capacity may be sold as quotas or seats at one surface and as variable token consumption at another.

Cabina.AI ↗

From Tokenmaxxing to Return on Token

Tokenmaxxing is the practice of maximizing AI consumption as a proxy for productivity. Tokenminimizing is the practice of reducing avoidable token consumption while preserving the required result. Tokenminning and tokenmining are variant spellings now appearing in the same efficiency conversation. Modelmaxxing means matching each workload to the least expensive model capable of performing it reliably. Token yield measures useful work relative to token consumption. Today's enterprise coverage pushes the arc toward return on token: not how many tokens were saved, but how much useful output the remaining spend purchased.

EY

EY's September 21 agentic-AI ROI framework argues that organizations need to measure the full cost of agents and the value they create rather than stopping at token consumption. It separates infrastructure, model usage, integration, governance, and human oversight from benefits such as productivity, revenue, quality, and risk reduction. That is the finance-side version of token yield: the denominator needs an outcome.

EY ↗

Weave Index

Weave's Return on Token Spend metric explicitly divides observed engineering output by attributed AI spend. Its August 2026 snapshot reports 0.091625 estimated expert hours of engineering output per dollar across 7,020 engineers with both spend and merged-code telemetry. The methodology is imperfect by design, but the vocabulary is important: AI spend is being measured against work rather than usage.

Weave Index ↗

Tokenminning

The tokenminning vocabulary has expanded from a slogan into a proposed engineering discipline covering model routing, context hygiene, spend attribution, agent caps, and deployment enforcement. Its current definition is deliberately outcome-preserving: reduce token consumption where spend does not convert to useful output. That distinction separates tokenminimizing from simply starving a workload.

Tokenminning ↗

BenchLM

BenchLM's September 27 inference guide lists Claude Sonnet 5 at $2 input and $10 output per million tokens, GPT-5.6 Terra at $2 and $12, and DeepSeek V4.1 Flash at $0.30 and $1.20. Even before cache discounts, latency, or quality are considered, the output side spans a tenfold price range. Modelmaxxing increasingly means routing by required capability rather than letting one default model inherit every workload.

BenchLM ↗

Data Driven Investor

A recent measurement of MCP tool schemas argues that eagerly loaded tool definitions can consume 6.6 times more context than a more selective representation under the newer specification. The broader tool-surface lesson remains independent of the exact implementation: an agent should be able to discover thousands of capabilities without paying to place thousands of full schemas into every model request.

Data Driven Investor ↗

Research Watch

Grow the Harness, Not the Context: From Strategy-Free Scaffolds to Reusable Specialist Agents

Growing Harness moves recurring agent control logic out of repeatedly reconstructed model context and into reusable executable code learned from task failures. Across BrowseComp-Plus and WebArena-Verified with deployment models from 4B to 120B parameters, the authors report 76.0% to 91.8% fewer LLM calls and 74.4% to 98.6% lower deployed-agent inference cost relative to a tool-calling agent.

Why it matters: This is tokenminimizing by architectural substitution. If recurring reasoning becomes reusable program logic, the system avoids paying the model to rediscover the same procedure on every run.

arXiv ↗

AgenticCache: Cache-Driven Asynchronous Planning for Embodied AI Agents

AgenticCache reuses frequent plan transitions rather than making an LLM call at every agent step, while a background updater validates and refreshes cached plans. Across four multi-agent embodied benchmarks and three models, the authors report 22% higher average task success, 65% lower simulation latency, and 50% lower token usage.

Why it matters: Caching can improve token yield by eliminating complete inference calls rather than merely discounting repeated input. The useful unit being reused is the decision itself.

arXiv ↗

Cost-Optimal LLM Routing with Limited User Feedback under User Satisfaction Guarantees

SLARouter learns an online cost-aware routing policy from sparse production feedback while enforcing a user-satisfaction constraint. Across its benchmark evaluations, the authors report satisfying quality constraints while reducing operating cost by as much as 2.2 times compared with existing routing baselines, without per-benchmark tuning.

Why it matters: Modelmaxxing needs a quality floor. Routing to the cheapest model only improves token economics when the resulting answer remains acceptable to the user.

arXiv ↗

Agent-as-a-Router: Agentic Model Routing for Coding Tasks

Agent-as-a-Router treats model selection as a continuing context-action-feedback loop rather than a one-time classifier. The authors report that simply adding task-dimension performance statistics to a vanilla LLM router produced a 15.3% relative gain, and their full ACRouter accumulated execution-grounded experience while routing among eight frontier models across roughly 10,000 coding tasks.

Why it matters: The best-value model for a workload is not necessarily static. A router that learns from actual outcomes can improve return on token as prices, models, and task mixes change.

arXiv ↗

Phrase of the Day

“Return on token”

Return on token is the useful work or business value produced relative to the money spent on AI token consumption, shifting the optimization target from raw usage or raw savings to measurable outcomes per dollar.

  1. AI adoption
  2. Tokenmaxxing
  3. Token shock
  4. Tokenminimizing
  5. Modelmaxxing
  6. Token discipline
  7. Token yield
  8. Return on token

The likely winners are organizations that can connect AI consumption to completed work, then improve that ratio through model routing, selective context, cache reuse, smaller tool surfaces, bounded agent execution, and avoided inference.

A smaller token bill is nice. A larger pile of useful work for every dollar on that bill is the actual point.

Weave Index ↗

The jCodeMunch read

Today's return-on-token and tool-surface stories reward the same basic discipline: make useful information available without making all available information resident. jCodeMunch reduces code-reading tokens through tree-sitter symbol retrieval and byte-precise context, keeping relevant code evidence available without making repository-scale context the default.

See how the 95%+ cut is measured →

← All editions