Token Cost Radar

Token Cost Radar

September 30, 2026

Today's token-cost story is moving from cheap tokens to productive tokens. OpenAI's GPT-6.1 Sol launch puts a concrete number on the distinction, with the company reporting near-Astra capability at one-fifth of Astra's standard input and output token prices and substantially lower cost per task on its scientific benchmark. DevDay also introduced an Ultrafast tier that charges more for latency, while fresh market data shows identical open-weight models can vary nearly 20-fold in price depending on the host. Meanwhile, DeepSeek is working with Huawei on infrastructure aimed at extracting more performance from non-NVIDIA hardware. The arc from tokenmaxxing through tokenminimizing to token yield is becoming workload economics: price the model, the host, the speed, the cache behavior, and finally the completed task.

Top Developments (Last 24 Hours)

1What if the better model costs less per task without costing less per token?

OpenAI released GPT-6.1 Sol on September 29 at $2 per million input tokens, $0.10 per million cached input tokens, and $10 per million output tokens. On Terminal-Bench Science 0.1 at maximum reasoning effort, OpenAI reports an average cost of $5.47 per task versus $23.80 for GPT-6 Astra, while GPT-6.1 Sol more than doubled the score of GPT-6 Sol. The important unit is no longer just dollars per million tokens. It is dollars required to reach the desired result.

OpenAI ↗

2OpenAI puts an explicit price on speed

OpenAI's new Ultrafast API service tier offers GPT-6 Astra at up to eight times Standard speed and recommends persistent WebSocket connections for tool-heavy agent workloads. GPT-6.1 Sol's API documentation separately lists Fast mode at twice Standard token prices, while Batch and Flex are 50% below Standard. The same model can therefore occupy multiple price points depending on whether the workload values latency or thrift.

OpenAI API ↗

3OpenAI turns heavy AI usage into a $500 monthly seat

Business Insider reports September 30 that OpenAI's new Pro 500 subscription costs $500 per month and bundles its highest usage allowances with access to Ultrafast compute in products including ChatGPT Work and Codex. The move is another example of AI spend splitting into two meters: variable API token billing for developers and increasingly expensive capacity or seat tiers for heavy interactive users.

Business Insider ↗

4DeepSeek and Huawei push inference economics down into the chip stack

Reuters reports September 30 that DeepSeek and Huawei are collaborating on open-source programming infrastructure for Huawei Ascend AI chips, including compute and communication libraries and the TileLang programming language. The work accompanies a supernode design using 128 Ascend 950 chips. The project is strategically about reducing dependence on NVIDIA, but economically it is another attempt to improve the software-to-hardware efficiency underneath every generated token.

Reuters ↗

From Tokenmaxxing to Workload Economics

Tokenmaxxing is the practice of maximizing AI consumption as a proxy for productivity. Tokenminimizing is the practice of reducing avoidable token consumption while preserving the required result. Tokenminning and tokenmining remain variant spellings in the same efficiency conversation. Modelmaxxing means matching each workload to the least expensive model capable of performing it reliably. Token yield measures useful work relative to token consumption. Today's pricing data extends the arc again: the economical choice can depend as much on host, latency tier, cache reuse, and task-level token burn as on the model name.

BenchLM

BenchLM's September 29 price board tracks 168 paid models across 30 providers. It lists Qwen3.7 Flash as the cheapest API at $0.03 input and $0.13 output per million tokens, MiMo-V2.6-Pro as the cheapest model clearing its production-grade capability threshold, and Claude Sonnet 5.5 as the cheapest model clearing its frontier threshold. That separation captures modelmaxxing neatly: cheapest token, cheapest adequate model, and cheapest frontier model are three different questions.

BenchLM ↗

LLM Latency

A September 29 host comparison finds that identical model weights can carry dramatically different API prices. DeepSeek V4 Flash 0731 spans a reported 19.5-fold input-price range across eight hosts, while DeepSeek V4.1 Flash spans fivefold. Host selection is therefore becoming another routing dimension alongside model capability: the model can stay fixed while both latency and effective token price move.

LLM Latency ↗

AI Pricing Guru

DeepSeek V4.1 Flash remains a useful extreme in cache-aware pricing. Current official off-peak rates are $0.15 per million fresh input tokens, $0.003 for cache hits, and $0.60 for output, with peak rates exactly double. A cache-hit input token is therefore 50 times cheaper than fresh input off-peak, making prefix reuse and workload scheduling independent cost controls.

AI Pricing Guru ↗

HowManyTokens

The September 29 cost leaderboard prices a fixed workload of 1,000-token input and 200-token output calls across current models. GPT-5 Nano leads the tracked non-deprecated set at $130 per million such calls, followed by Qwen3.8 Flash at $146, with GPT-6 Luna and DeepSeek V4 Flash also in the low-cost group. Converting rate cards into a fixed workload makes the output-token component visible and moves comparison closer to actual task economics.

HowManyTokens ↗

Anthropic

Anthropic's September 29 Microsoft Foundry session highlighted tool search alongside MCP connectors and other agent tools. The architecture is relevant to the tool-surface lane because tool search allows an agent to select capabilities dynamically instead of treating every connected tool as permanently resident context. That matters as MCP catalogs grow and schema tokens compete with task context for the same window.

Anthropic ↗

Research Watch

An Empirical Cost Attribution of Context-Compression Gateways in Multi-Turn Coding Agents

This study separates tool-schema filtering, content compression, and history summarization as distinct token-saving levers. Tool-schema filtering removed roughly 21,000 to 57,000 tokens on a typical turn and produced a reproducible linear saving, while compressed file reads accumulated across history and produced approximately quadratic savings until the context window capped the effect.

Why it matters: Tool-surface bloat is unusually attractive to attack because unused schemas impose a predictable recurring tax. Tokenminimizing starts with removing context the agent did not need in the first place.

arXiv ↗

AgentRouter: Heterogeneous Model Routing for Cost-Optimal Multi-Step Agentic Workflows

AgentRouter assigns individual steps inside an agent trajectory to one of four model tiers rather than routing the entire workflow to a frontier model. Trained on 50,000 annotated trajectory steps, the authors report 72% lower cost than frontier-only execution while retaining 97.3% of frontier-only quality, with less than 5 milliseconds of routing overhead per step on an A100.

Why it matters: Modelmaxxing becomes more precise when the routing unit shrinks from the application to the task and then to the individual step. Planning may deserve frontier inference while formatting does not.

arXiv ↗

What Does Context Compression Cost an Agent? Interaction Costs Unrevealed by Task-Completion Metrics

This study finds that aggressive context compression can leave task-completion rates apparently unchanged while forcing agents to spend additional tool calls reacquiring information that was removed. At one fivefold compression point, GPT-5.5 completion moved from 80% to 85% without statistical significance while retrieval calls rose from 21.0 to 63.9.

Why it matters: A smaller context is not automatically higher token yield. If an agent must repeatedly fetch discarded state, apparent token savings can migrate into tool calls, latency, and later context.

arXiv ↗

How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks

Across eight frontier models on SWE-bench Verified, researchers found agentic coding tasks consuming roughly 1,000 times more tokens than simpler code reasoning and chat tasks. Repeated runs on the same task varied by as much as 30 times in token consumption, higher spending did not reliably produce higher accuracy, and models systematically underestimated their own eventual token use.

Why it matters: Agent budgets need external controls because token consumption is both stochastic and poorly self-predicted. More inference is not automatically more useful inference.

arXiv ↗

Phrase of the Day

“Cost per task”

Cost per task is the total model and inference spend required to complete one defined unit of useful work at an acceptable quality level, incorporating both token prices and the number and type of tokens the workload actually consumes.

  1. AI adoption
  2. Tokenmaxxing
  3. Token shock
  4. Tokenminimizing
  5. Modelmaxxing
  6. Token discipline
  7. Token yield
  8. Cost per task

The likely winners are organizations that benchmark complete workloads instead of rate cards, then improve task economics through model routing, cache reuse, selective context, smaller tool surfaces, latency-tier selection, and bounded agent execution.

A cheap token is a fine thing to buy. Needing fewer of them to finish the job is better.

OpenAI ↗

The jCodeMunch read

Today's cost-per-task and tool-surface stories reward selective context rather than context starvation. jCodeMunch reduces code-reading tokens through tree-sitter symbol retrieval and byte-precise context, keeping relevant code evidence available without making repository-scale context the default.

See how the 95%+ cut is measured →

← All editions