Token Cost Radar

Token Cost Radar

October 1, 2026

Today's token-cost story has moved beyond counting tokens to deciding which tokens deserve to exist. Google's Gemini 4 Argon arrives with introductory pricing of $2 per million input tokens and $10 per million output tokens, plus a 95% cached-input discount, while its million-token output ceiling makes runaway generation a newly visible budget variable. Enterprise AI cost management is attracting fresh capital, with Ascerta raising $18 million around connecting AI spend to measurable business value. Meanwhile, new research shows a decision-only evaluator can handle routine judgments for a tiny fraction of a frontier judge's fee and escalate uncertain cases. The arc from tokenmaxxing through tokenminimizing to token yield is becoming selective inference: spend expensive intelligence only where expensive intelligence changes the result.

Top Developments (Last 24 Hours)

1How do you budget for a model that can generate a million tokens in one run?

Google announced Gemini 4 Argon on September 30 with a 1 million-token output limit and introductory API pricing of $2 per million input tokens and $10 per million output tokens. Cached input receives a 95% discount. Google says the introductory rates later rise to $4 input and $20 output, although it has not announced when that change occurs. Argon is initially restricted to trusted cyber defenders, with paid API customers and Google AI Ultra users next in line. The unusually large output ceiling makes generation limits and agent budgets especially important because output remains the expensive side of the meter.

Google ↗

2AI cost management attracts $18 million around a new question: which spend actually pays off?

Ascerta, formerly Pay-i, announced September 30 that it raised an $18 million Series A led by Dell Technologies Capital, bringing total funding to $22.9 million. The company is expanding from AI cost management toward measuring adoption, costs, resource use, and business outcomes together. The funding is a useful market signal: enterprise buyers increasingly want to connect model calls, coding-agent usage, licenses, and infrastructure spend to measurable ROI rather than merely observe the bill.

SiliconANGLE ↗

3CIOs are being told token consumption is not a success metric

CIO's September 30 AI ROI analysis argues that enterprises need workload-level cost controls and business-value measurement rather than treating adoption or token consumption as evidence of success. The framing reflects a broader FinOps shift from asking how much AI was used to asking what each workload cost and what outcome it produced.

CIO ↗

4A software company says it deliberately does not count employee tokens

In a September 30 interview, Railsware told dev.ua that it does not centrally count individual token consumption and instead chooses AI tools and subscription structures where granular token accounting is unnecessary for ordinary employee use. The company still tracks overall AI spending and economics. The approach highlights a competing governance model: cap or bundle consumption at the seat and tool level instead of forcing every employee to become a miniature FinOps analyst.

dev.ua ↗

From Tokenmaxxing to Selective Inference

Tokenmaxxing is the practice of maximizing AI consumption as a proxy for productivity. Tokenminimizing is the practice of reducing avoidable token consumption while preserving the required result, with tokenminning and tokenmining also appearing as variants in the efficiency conversation. Modelmaxxing means matching work to the least expensive model capable of performing it reliably. Token yield measures useful work relative to token consumption. Today's emerging pattern is selective inference: reserve expensive generation, reasoning, and long context for the parts of a workload where they materially improve the outcome.

Flexera

Flexera's current FinOps for AI guide says 98% of FinOps teams now manage AI spend, up from 31% two years earlier. Its framework treats cost as an architectural metric alongside latency, throughput, and accuracy, with model rightsizing, caching, workload ownership, and spend allocation built into the operating discipline. AI FinOps is moving from invoice archaeology into the execution path.

Flexera ↗

ReqKey

Pricing verified October 1 shows the current Claude ladder spanning Haiku 4.5 at $1 input and $5 output per million tokens, Sonnet 5 at $2 and $10, Opus 5.5 at $4 and $20, and Fable 5.1 at $10 and $50. Cached input ranges from $0.10 to $0.25. A 10-fold input-price range inside one model family makes modelmaxxing a straightforward budget lever before more elaborate routing begins.

ReqKey ↗

DeepSeek and GLM pricing

A current comparison puts DeepSeek V4.1 Flash at $0.15 per million fresh input tokens off-peak, $0.003 for cache hits, and $0.60 for output, versus GLM 5.3 Flash at $0.15 fresh input, $0.03 cached input, and $0.50 output. DeepSeek's cache-hit input is therefore 50 times cheaper than its own fresh input off-peak. Cheap tokens are no longer one number because reuse patterns can dominate the effective rate.

Intelligent Living ↗

AI gateways

September 30 coverage of AI gateways frames them increasingly as cost-control infrastructure, centralizing model routing, budgets, caching, rate limits, retries, and spend attribution rather than functioning only as API proxies. The architectural appeal is simple: a shared gateway can enforce token discipline before an inefficient request or retry loop reaches the provider bill.

Technology.org ↗

Anthropic tool-use research

Anthropic's tool-search measurements remain an important reference for the tool-surface lane. Its published example found 58 tool definitions consuming roughly 55,000 tokens before the conversation began, with larger configurations reaching 134,000. Retrieving tools on demand reduced token usage by 85% in its evaluation. The durable lesson is that capability can remain discoverable without every schema becoming permanent context.

Anthropic ↗

Research Watch

JEV-as-a-Judge: Accept When Confident, Escalate When Unsure

This recently updated study compares a decision-only evaluator with sixteen generative and reward-model judges. On ordinary preference and evidence-grounded factuality, JEV came within three percentage points of the strongest LLM judge at 0.36% of its estimated fee. A confidence-gated cascade accepted inexpensive judgments when confidence was high and escalated uncertain cases, retaining nearly all of the stronger judge's performance at lower cost.

Why it matters: Selective inference can be more powerful than simply selecting a cheaper generative model. Some bounded decisions do not need generated reasoning at all, while confidence can determine when paying for the stronger model is justified.

arXiv ↗

Grow the Harness, Not the Context: From Strategy-Free Scaffolds to Reusable Specialist Agents

Growing Harness moves recurring agent control decisions out of repeatedly reconstructed model context and into reusable executable code learned from task feedback. Across BrowseComp-Plus and WebArena-Verified with deployment models from 4B to 120B parameters, the authors report 76.0% to 91.8% fewer LLM calls and 74.4% to 98.6% lower deployed-agent inference cost relative to a tool-calling agent.

Why it matters: This is tokenminimizing by substitution rather than compression. Once recurring reasoning can safely become executable logic, there is little economic reason to buy the same reasoning tokens again on every run.

arXiv ↗

AgentRouter: Heterogeneous Model Routing for Cost-Optimal Multi-Step Agentic Workflows

AgentRouter assigns individual steps inside an agent trajectory to one of four model tiers instead of routing the entire workflow to a frontier model. Trained on 50,000 annotated trajectory steps, the authors report a 72% cost reduction relative to frontier-only execution while retaining 97.3% of frontier-only quality, with less than 5 milliseconds of routing overhead per step on an A100.

Why it matters: Modelmaxxing becomes more precise when the routing unit is the individual step. Planning may justify frontier reasoning while extraction, formatting, or routine tool selection may not.

arXiv ↗

Dual-Pool Token-Budget Routing for Cost-Efficient and Reliable LLM Serving

This serving study finds that 80% to 95% of production requests in its analyzed traces are short even though inference instances are commonly provisioned for worst-case context length. Routing requests by estimated token budget into short-context and long-context pools reduced GPU-hours by 31% to 42% on Azure and LMSYS traces, while lowering preemption rates 5.4 times.

Why it matters: Token budgets can route infrastructure as well as models. Expected context size changes KV-cache demand and concurrency, so token discipline can reduce the physical cost of inference before provider pricing enters the equation.

arXiv ↗

Phrase of the Day

“Selective inference”

Selective inference is the practice of reserving expensive model calls, reasoning, context, or generation for the parts of a workload where they materially improve the result, while routing simpler decisions to cheaper models, cached results, retrieval, or deterministic logic.

  1. AI adoption
  2. Tokenmaxxing
  3. Token shock
  4. Tokenminimizing
  5. Modelmaxxing
  6. Token discipline
  7. Token yield
  8. Selective inference

The likely winners are systems that decide whether inference is necessary before deciding which model should perform it, then spend frontier tokens only on the steps where frontier capability changes the outcome.

Tokenminimizing trims the grocery bill. Selective inference asks whether every item needed to go in the cart.

arXiv ↗

The jCodeMunch read

Today's selective-inference and tool-surface stories point toward the same architecture: make information available without making the model process all of it by default. jCodeMunch reduces code-reading tokens through tree-sitter symbol retrieval and byte-precise context, keeping relevant code evidence available without making repository-scale context the default.

See how the 95%+ cut is measured →

← All editions