What is the State of Inference Spend 2026?
Original o10 research quantifying compliant price spread across inference venues, workload savings models, and enterprise routing benchmarks. June 2026.
An illustrative framework for comparing inference costs at a defined quality threshold. Its examples explain the calculation method; they are not measured customer results.
Methodology preview · Illustrative workload models · Worked examples, not measured customer results
Start with the core questions, then examine the examples and tradeoffs below.
Original o10 research quantifying compliant price spread across inference venues, workload savings models, and enterprise routing benchmarks. June 2026.
Traffic stays on default model routes across fragmented gateways. Finance sees blended invoices; platform teams lack a control point to enforce envelopes when prompts or retries change.
o10 replayed representative enterprise workloads against candidate models on unified inference gateway, OpenRouter, Amazon Bedrock committed capacity, and owned open-weight. At identical eval floors per use case.
Benchmarks use per-use-case quality floors, not global averages.
June 2026 venue survey across gateway, aggregator, and committed capacity pricing.
Open-weight 8B-class models clear many batch workloads at $0.05/1M on committed capacity versus $0.12 on gateways.
Mini-class tiers show 20–30% spread between gateway and committed routes. Material at billions of tokens per month.
Frontier tiers remain expensive; most enterprise use cases clear below sonnet-class when evals are workload-specific.
| Tier | Gateway | Aggregator | Committed |
|---|---|---|---|
| Open-weight 8B | $0.12 | $0.08 | $0.05 |
| Haiku-class | $0.65 | $0.48 | $0.42 |
| GPT-mini-class | $2.40 | $2.10 | $1.85 |
| Sonnet-class | $9.40 | $8.10 | $7.20 |
| Frontier | $31.90 | $28.00 | $24.50 |
Savings depend on use case, quality floor, and venue mix, not a single percentage.
RAG summarization at balanced floor: up to 80% versus default sonnet-class routing.
Support assistants at strict QA floor: 40–60% with mini-class compliant routes.
Batch classification at lean floor: up to 94% routing to open-weight when evals permit.
Inference spend is controllable when routing sits in the path with eval-gated selection.
Dashboards report last month's tokens. A control plane changes next month's routes, with proof in shadow first.
KYI adds board-grade governance above raw savings: performance, economics, integration, strategy, and risk.
Support, RAG, code, batch. Each has different volume and floor.
Define the cheapest compliant tier, not the default frontier model.
Build verified savings baseline per use case.
Hold budget and policy on every subsequent call.
o10 State of Inference Spend 2026. Venue survey June 2026. Workload models use per-use-case eval floors. o10 KYI framework.
The spread is not uniform. It varies by workload, eval floor, and venue mix, but it demonstrates that default model routes routinely overshoot cheapest compliant supply. Shadow mode proves your organization's spread against your traffic; this report documents the benchmark methodology and venue price tables behind the headline number.
The gap exists because traffic stays on default models across fragmented gateways while cheaper tiers would clear the same eval floor. Dashboards reveal the overspend after invoices arrive; a control plane captures it on the next call. Segmenting by use case is essential. RAG and batch often show the largest absolute delta.
The June 2026 venue survey covers per-token API gateways (unified inference gateway), OpenRouter (multi-provider aggregator), Amazon Bedrock (per-token and committed capacity), and owned or open-weight infrastructure. Pricing tables in the report show five model tiers from open-weight 8B ($0.05/1M on committed) through frontier ($31.90/1M on gateway). Production stacks commonly combine multiple venues; o10 routes above all of them under one policy and ledger.
Check current prices before a procurement or routing decision. Keep the date and endpoint with each reference rate, and compare it with your own usage records before forecasting costs.
A quality floor is the minimum eval score a model must achieve for a specific use case before o10 routes production traffic to it. Floors are per workload. Support, RAG, code, and batch clear at different bars, and measured by replaying representative traffic through eval suites, not assumed from vendor benchmarks. Once a cheaper candidate passes the floor, o10 can route to it in shadow (proof) or enforce (live). Floors without evals are hopes; evals without floors are expensive defaults. Benchmark comparisons use identical floors across venues. Otherwise price spread measurements compare unlike routes and overstate savings.
Yes. Routing compliant workloads through Amazon Bedrock committed capacity draws down sunk cloud commitments and lowers marginal $/1M versus pure per-token API rates. Many enterprises underutilize signed Bedrock spend while live traffic bills marginal gateway rates. o10 models capex/opex crossover per use case and routes through committed venues when evals permit. Turning reserved capacity into inference value.
High-volume RAG summarization and batch classification typically show the largest absolute monthly savings because token volume multiplies small per-million price deltas. Support assistants save materially at strict QA floors when mini-class models clear evals. Code and agents benefit from per-step routing to prevent frontier defaults on every hop. Savings percentages range from 40–94% depending on floor and venue mix. Shadow mode verifies your workloads.
Shadow mode mirrors live inference traffic through o10 without changing production routes. For every request, o10 evaluates candidate models against your per-use-case quality floors and records which route would have been cheapest and compliant. Along with the cost delta, while the original provider still serves the response. Engineering sees proof without production risk; finance gets a verified savings figure tied to your traffic, not industry averages. Most teams run shadow for 7–14 days segmented by use case (support, RAG, code, batch) before flipping enforce mode. The report provides industry benchmark context; shadow provides your organization's proof against that methodology.
The State of Inference Spend 2026 report is published by o10 with benchmark methodology and venue survey data from June 2026. Framework context draws on Know Your Inference (KYI) by o10, which governs how boards evaluate inference supply chains beyond per-token cost. Cite the report URL and methodology section when reproducing spread or pricing statistics.
The KYI whitepaper explains the framework’s scoring pillars and how to use them in a review. Follow the related link to read its methodology and limitations.
Paste a week of traffic. Get the number that books the audit.
See what you're overpaying →