Today's token-cost story is about the price of keeping capability permanently visible. Uber has published the architecture behind an MCP estate that grew to roughly 800 servers and 5,000 tools, forcing the company to build a gateway for discovery, access control, and tool selection rather than expose the entire catalog indiscriminately. That lands beside fresh pricing data showing open-weight APIs at an 81% median discount to proprietary models and frontier token prices 84% below their 2023 baseline. Cheap inference is abundant. Cheap context is not. The vocabulary arc from tokenmaxxing through tokenminimizing to token yield is increasingly about context residency: what must the model see now, what can remain discoverable, and what should never enter the prompt at all.
Top Developments (Last 24 Hours)
1What happens when your agent has 5,000 tools?
Uber published October 2 details of its MCP Gateway after internal adoption expanded to roughly 800 MCP servers exposing about 5,000 tools. Rather than make clients manage that estate directly, Uber built a centralized gateway for server discovery, authentication, authorization, configuration, and access. The architecture is a concrete example of the tool-surface problem reaching enterprise scale: capability catalogs can grow far beyond what should be treated as permanently resident agent context.
Uber Engineering ↗2Open-weight inference carries an 81% median price discount
BenchLM's October 2 pricing statistics track 181 models and put the median API rate at $0.95 per million input tokens and $3.75 per million output tokens. At a 3:1 input-to-output mix, open-weight models have a median blended price of $0.50 per million tokens versus $2.63 for proprietary models, an 81% discount. The data reinforces the modelmaxxing case: workload routing has a wide enough price surface to matter even before caching and batch discounts enter the calculation.
BenchLM ↗3Frontier token prices hold at $2.13 per million blended tokens
The Token Price Index updated October 2 at $2.13 per million blended tokens, down 9.1% month over month but up 80.8% year over year. Claude Sonnet 5.5 replaced Sonnet 5 at unchanged $2 input and $10 output rates, while GPT-6.1 Sol replaced GPT-6 Sol after only ten days. The combination illustrates why AI budgets cannot assume a simple downward price curve: capability, model composition, and product turnover can move faster than procurement cycles.
Token Price Index ↗4AI gateways move from routing traffic to governing the bill
CRN reports October 2 that CData's Connect AI Gateway is being positioned as a control point between enterprise agents and underlying IT systems as companies confront high AI costs, uncertain ROI, and limited governance. The product combines access and action controls with centralized agent connectivity. The larger trend is increasingly clear: the gateway layer is becoming where enterprises expect to govern both what agents can reach and how much those interactions cost.
CRN ↗From Tokenmaxxing to Context Residency
Tokenmaxxing is the practice of maximizing AI consumption as a proxy for productivity. Tokenminimizing is the practice of reducing avoidable token consumption while preserving the required result, with tokenminning and tokenmining also appearing as variants in the efficiency conversation. Modelmaxxing means matching work to the least expensive model capable of performing it reliably. Token yield measures useful work relative to token consumption. Today's MCP and pricing signals add another control point: context residency asks which information and capabilities must occupy the model's prompt now, versus remaining cheaply discoverable until needed.
Claude ecosystem
An October 2 developer briefing reports that Claude Code 2.1.287 can defer an MCP server's tools behind tool search with alwaysLoad set to false, rather than listing the server's tools in the prompt up front. The same briefing notes a similar search-first pattern in Pi. The implementation detail matters because tool discovery is becoming an explicit alternative to permanent schema residency.
MindPattern ↗IFX
The IFX Inference Index closed October 2 unchanged at 85.32 across a 29-model basket. Its capability-adjusted board lists a blended $3.375 per million tokens for the cheapest model clearing its frontier threshold, $1.6875 for its capable threshold, and $0.4625 for its budget threshold. Modelmaxxing turns that spread into a routing question: what is the cheapest capability tier that can reliably finish this particular job?
IFX ↗Current API pricing
An October 2 pricing snapshot shows how many independent levers now sit behind one nominal model bill. Claude Sonnet 5.5 is listed at $2 input, $0.20 cache read, $2.50 cache write, and $10 output per million tokens, while GPT-6.1 Sol is $2 input, $0.10 cached input, $2.50 cache write, and $10 output. Batch and Flex tiers cut the latter to $1 input and $5 output. Fresh input, reused input, output, and processing tier are now separate budget variables.
Prompt20 ↗DeepSeek self-hosting economics
Tom's Hardware reports on a consultancy that rented four NVIDIA H200 GPUs to test claims that self-hosting DeepSeek V4.1 Flash would dramatically undercut Claude. The rented hardware cost about $440.88 per day, or roughly $13,200 per month, versus an estimated $184 to $223 per day for DeepSeek's hosted API and about $5,500 per month for the consultancy's Claude subscription. The experiment is a useful warning that cheap open weights do not automatically mean cheap self-hosted inference.
Tom's Hardware ↗MCP spending controls
Runtime's MCP documentation, updated October 2, explicitly places spending checks in the same shared control plane as authentication, ownership, and idempotency. That pairing is a small but telling signal for agent FinOps: tool access and budget enforcement are beginning to live at the same execution boundary instead of being reconciled after the agent finishes spending.
Runtime ↗Research Watch
Tool-Schema Compression Enables Agentic RAG Under Constrained Context Budgets
This study evaluates 28 tool definitions across 6,566 controlled API calls and finds that conservative tool-schema compression saves 44% to 50% of schema tokens. At an 8K context budget, ordinary JSON schemas overflowed the available window and produced 2.6% average exact match, while compressed schemas restored operation with a 20.5 percentage-point average lift. Scaling tests found ordinary schemas overflowing at roughly 494 tools while compressed representations remained operational beyond 800.
Why it matters: Tool-surface bloat is not merely an invoice problem. Schemas compete directly with task evidence for finite context, so reducing resident tool metadata can preserve both token budget and usable capability.
arXiv ↗Risk-Constrained Freshness-Aware Semantic Caching for Open-Web Retrieval-Augmented LLMs
FreshCache treats semantic-cache reuse as a freshness-risk decision rather than relying on similarity alone. On an 8,072-query benchmark expanded to 31,201 paraphrased queries, its learned policy achieved 97% search API savings at a 24-hour evaluation window with 0.1% hash-based stale error. The authors estimate answer-affecting stale error at roughly 0.034% in a separate judged sample.
Why it matters: Semantic caching improves token yield only when reuse remains trustworthy. A risk gate lets applications avoid retrieval and inference work when cached knowledge is likely still valid while escalating cases where freshness matters.
arXiv ↗C2KV: Compressed and Composable KV Cache Reuse for Efficient LLM Inference
C2KV combines KV-cache compression with reusable non-prefix context representations designed to be concatenated across long-context workloads. Across multiple model families and long-context benchmarks, the authors report up to 17 times faster inference while preserving generation quality, with the design targeting both redundant prefill computation and the storage and transfer cost of large caches.
Why it matters: Repeated context has two prices: recomputing it and keeping its inference state around. Efficient reusable caches attack both, moving token economics below the visible prompt.
arXiv ↗Dual-Pool Token-Budget Routing for Cost-Efficient and Reliable LLM Serving
This serving study finds that 80% to 95% of requests in its production traces are short despite fleets commonly being configured for worst-case context. Routing requests by estimated token budget into separate short-context and long-context pools reduced GPU-hours by 31% to 42% on Azure and LMSYS traces and lowered preemption rates 5.4 times.
Why it matters: Token budgeting is becoming an infrastructure primitive. The predicted token footprint can determine which serving pool should handle a request before the first inference token is produced.
arXiv ↗Phrase of the Day
“Context residency”
Context residency is the amount of information, instructions, history, and tool metadata kept permanently visible to a model rather than retrieved or loaded only when the current task requires it.
- AI adoption
- Tokenmaxxing
- Token shock
- Tokenminimizing
- Modelmaxxing
- Token discipline
- Token yield
- Context residency
The likely winners are systems that separate availability from residency, keeping large capability and knowledge surfaces discoverable while placing only the small task-relevant subset into expensive model context.
- on-demand tool discovery
- retrieval-based schema loading
- semantic caching
- context-budget enforcement
- cost-aware model routing
- token-budget-aware serving
- cost-per-completed-task measurement
Owning 5,000 tools is capability. Making the model read all 5,000 menus before ordering lunch is overhead.
Uber Engineering ↗