o10Inference spend control plane

Frontier for frontier.
Cheapest compliant for everything else.

o10 holds your quality floor in the request path — and routes each call to the cheapest model that clears it. You pay actual cost, proven per call.

in the path · not after the fact
In path above
3 venues
Spread observed
638×
Routing
Frontier when required
The Book · cost per 1M tokens
Saved / month · ILLUSTRATIVE
$0
vs. current routing · estimate only
envelope
cheapest compliantmost expensive
Routed to
Quality floor
Routed spend$0
drag the envelope ← → to set the budget
FOR ENGINEERING

One base URL. Same OpenAI-compatible API.

Point at https://app.o10.io/v1. Set o10/auto — or frontier, squad, pinned.

How to use → Modes →
FOR FINANCE

Set the envelope. o10 holds it in the path.

Shadow proves savings on your traffic. Enforce holds the budget. Ledger per use case.

Four questions →
FOR BUSINESS TEAMS

DeepShell: an agent on your computer, on your budget.

Approval-gated desktop agent with a tamper-evident audit log — powered by o10 routing. Beta.

DeepShell → o10 Tune →
01The problem

Your AI spend is a surprise,
not a plan.

fx

Frontier pricing on non-frontier work

Support, RAG, and batch run on default frontier routes. Productivity improves — but every increment costs frontier tokens.

4+

Spend is fragmented

Calls scatter across gateways and cloud accounts. No single ledger, no single owner, no single number.

T+30

The bill lands after the fact

The dashboards you have report what was spent — a month late, when the money is already gone.

You can see the leak. You can't pull a lever on it.

02The shift

Marginal cost of tokens Marginal cost of productivity

Frontier models for frontier problems. o10 routes everything else to the cheapest compliant model that clears your quality floor — in the path, not on a dashboard.

03How it works

One plane, every venue.

The same model is often priced differently across venues. o10 routes every call to the cheapest compliant supply — and starts in shadow mode before it ever changes a route.

Works with Unified inference gateway · OpenRouter · Amazon Bedrock · owned / open-weight
routing live
enforce · in the path
YOUR APP live traffic o10 CONTROL PLANE quality floor · evals eval & budget route → cheapest IN THE PATH CHEAPEST COMPLIANT SUPPLY Gateway PER-TOKEN API $9.40 /1M TOKENS Aggregator MULTI-PROVIDER $8.10 /1M TOKENS Your commitment YOUR BEDROCK / BYOK $1.85 /1M TOKENS ← o10 would route here venues: per-token APIs · aggregators · your Bedrock commitment · BYOK / open-weight
in the path · routing live
calls routed: 0
saving · illustrative$312K/mo

The same model is priced differently across venues — $9.40 on a per-token API, $1.85 when drawn from your committed Bedrock spend. o10 routes every call to the cheapest venue you already have that clears your floor. Start in shadow — o10 mirrors traffic and shows what would route and save, with evals that prove the cheaper model clears your quality floor — then flip to enforce. That ramp is the trust story.

Start freeHow to use o10

Free route key. One base URL.
Same OpenAI-compatible API.

Start free — no card. Signing up at app.o10.io/signup creates an org and issues a free o10 route key (o10_sk_…), shown once. Free is $0 BYOK: you bring your own provider keys; o10 can run shadow receipts on live traffic. Paid plans add routed credits.

01
Start free

Sign up — get your free route key. No card.

02
Point the base URL

OpenAI-compatible client → https://app.o10.io/v1 (only a base-URL change).

03
Set the model

Use o10/auto (or another mode). Paste the free route key as the API key.

Shadow mode shows what you’d save before enforce. Same OpenAI-compatible API.

curl https://app.o10.io/v1/chat/completions \
  -H "Authorization: Bearer $O10_ROUTE_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "o10/auto",
    "messages": [{"role": "user", "content": "Hello"}]
  }'
# Copy your free key after signup — do not use a shared demo key.
Start free (no card) See modes
ModesVirtual model IDs

Four modes. One control plane.

Set model to a virtual alias — the developer surface of eval-gated routing. Base URL https://app.o10.io/v1. Concrete model slugs are also allowed; pin with ! to skip routing. Or use o10/pinned for an org-level saved model.

o10/autoDefault — routes each prompt to the cheapest model that still clears the quality bar
o10/frontierValue-first across frontier-class models — open-weight and proprietary; orgs can set a preferred frontier model
o10/squadMulti-agent: planner → workers → judge. For hard multi-step work (typically 1–4 minutes)
o10/pinnedOrg-saved model override — skips auto-downgrade; falls back to auto with a visible notice if unavailable

New — o10 Tune: free prompt optimization. Same quality, cheaper position, with a before/after receipt. Tune a prompt →

Frontier Tokens

o10/frontier

Value-first across frontier-class models — including open-weight (e.g. DeepSeek R1, GLM, Kimi) and proprietary (e.g. GPT-5.5, Claude Opus/Sonnet). Beta / live for entitled accounts.

Learn more Open in app →
Squad

o10/squad

Multi-agent: planner → workers → judge. For hard multi-step work. Latency typically 1–4 minutes — not chat-speed. Beta / live.

Learn more Open in app →
Pinned + Tag

o10/pinned · Slack

o10/pinned always uses your org-saved model. Tag: mention @o10 in Slack with receipts. Tune and DeepShell sit alongside.

All modes → o10 Tag →
04The four questions

Answer the four CFO questions
in under a minute. Then act on the answers.

Q1The fully loaded cost of each AI use case in production.→ live ledger per use case
Q2The cost per business outcome for each one.→ efficiency ratio, not token totals
Q3Which use cases can't defend their unit economics.→ flagged against the floor
Q4The forecast for next quarter, tied to a volume driver.→ bound to a business metric
THE o10 DIFFERENCE

Every answer comes with a lever. The use case that can't defend its unit economics gets auto-rightsized or capped at the control plane — not added to a slide for next quarter.

05Capex / opex

Route across the venues you already have.

Cost vs. monthly volume
per-token API (opex) your committed / open-weight

The make-vs-buy decision, modelled.

o10 routes across venues you already have — per-token APIs, aggregators, your AWS/Bedrock committed spend, and self-hosted open-weight models — and shows the crossover where moving a workload pays off. That committed capacity is yours, not o10's.

Routing through Bedrock can draw down committed cloud spend the company has already signed — turning a sunk commitment into realized inference value.

A capital-allocation decision a CFO owns — the thing no FinOps dashboard touches.
06Governance & sovereignty

Control which model sees which data.

/ POLICY

Policy-ready routing

Route by quality floor, budget, and model approval. Jurisdiction and data-residency controls are on the roadmap — not a live guarantee today.

Roadmap
/ LEDGER

Immutable audit trail

Every call leaves an immutable record — model, venue, and cost. A defensible ledger, per request.

per-call record
07Value-add · Know Your Inference

o10 enforces the spend.
KYI governs the chain.

Routing controls the bill. Know Your Inference is the framework above it — scoring every use case across performance, economics, integration, strategy, and risk, so your AI supply chain is sustainable, governable, and defensible to a board.

Public research: the Quality-Cost Frontier Index publishes what it costs to clear a quality bar by task type — with observation and org counts.

Explore the KYI framework Frontier Index
76/100 Recommended · sample
Performance 25%81
Economics 25%86
Integration 20%74
Strategy 20%58
Risk 10%79
08The spread

The same answer, 638× the price.

Pick a workload and a quality floor. o10 estimates what you'd have saved by routing to the cheapest compliant model. The embarrassing number that books the meeting. Same quality floor. Different token tier. That's how productivity improvement stops costing frontier prices.

Workload
Quality floor
Estimate only · paste a week of real
traffic in the audit for your number.
Estimated monthly saving
$0
0% of current spend, same quality floor
Current spend
$0
o10 routed
$0
Price spread
Cuts to spend
09Pricing & adoption

Start in shadow. Pay from what you save.

o10 is the only layer in the stack that earns more when your bill goes down — we profit from the spread we prove, not from your volume.

/ CONTROL

Governance fee

flat platform subscription

The control plane, policy engine, audit ledger, and routing across every venue. Priced like infrastructure, not a percentage of your bill.

/ GAINSHARE

Share of verified savings

measured against shadow baseline

A share of savings o10 proves against your own shadow baseline. You win only when the customer wins — no savings, no gainshare.

01
Shadow

o10 mirrors traffic and shows what would have routed and saved. Nothing changes.

02
Prove

A verified savings figure against your own baseline, per use case.

03
Enforce

Flip the switch. o10 holds the envelope in the path, on Monday.

Start free at app.o10.io/signup — no card. Free is $0 BYOK (bring your own provider keys) with shadow receipts; paid plans add routed credits. See how to use.

10Built for the compact

Finance sets the envelope.
Engineering keeps the keys.

The CFO / board

Owns the number.

  • The budget envelope, set and enforced.
  • The unit-economic floor each use case must clear.
  • The kill criteria when it can't.
  • A forecast tied to a business volume driver.
The CIO / platform

Keeps the keys.

  • Eval floors, latency, and reliability.
  • Data-residency controls (roadmap).
  • Architecture — no six-week ticket per target change.
  • Full visibility into every routed call.
o10 enforces the line
o10Set the envelope. o10 holds it.

See what you're overpaying.

Paste a week of traffic. Get the number that books the audit.

See what you're overpaying Start free (no card)
the same class of answer is priced up to 638× apart across venues — o10 routes to the cheapest one that clears your quality floor
FAQFrequently asked questions

Common questions

What is o10?

o10 is the control plane for inference spend — an OpenAI-compatible LLM routing gateway. It holds your quality floor in the request path and routes every AI inference call to the cheapest model that clears it across per-token APIs, aggregators, AWS Bedrock, and venues you already have (including your own cloud commitments and BYOK provider keys). Shadow mode proves savings without changing production; enforce mode holds the bar and budget envelopes in the path. Evals define per-use-case quality floors; KYI governs the supply chain for board reporting; an immutable ledger records model, venue, and fully loaded cost on every call.

What is the difference between shadow and enforce mode?

Shadow mode mirrors live traffic and shows what would have routed and saved — without changing production responses. Enforce mode places o10 in the request path and actively routes each call to the cheapest eval-passing model within your budget envelope. Teams always start in shadow to build a verified per-use-case baseline; finance signs off before enforce flips. Both modes write to the immutable ledger; only enforce changes spend.

How much can inference routing save?

o10 has observed up to 638× compliant price spread for the same quality floor across venues in June 2026 benchmarks. Workload-specific monthly savings typically range from 40–94% depending on use case, eval floor, and venue mix — RAG and batch at lean floors show the largest percentages. These are not guarantees: shadow mode proves your organization's number against your traffic. Teams without routing in the path leave an estimated 40–70% of compliant savings uncaptured.

What is Know Your Inference (KYI)?

Know Your Inference (KYI) is a governance framework by Shen Pandi that scores inference systems across five weighted pillars: Performance (25%), Economics (25%), Integration (20%), Strategy (20%), and Risk (10%). Each pillar scores 0–100; the composite rolls into a confidence level and board-signable recommendation. KYI runs continuously in the o10 control plane — not as a one-off audit — so every routed call and eval updates the score. A composite floor of 65 triggers enforcement levers: cap, rightsizing, or sunset per policy.

Does o10 replace my AI gateway?

No. o10 does not replace your AI gateway or developer-facing APIs. It sits above gateways and clouds, adding spend enforcement, eval-gated routing, policy, and CFO-grade ledger — not proxy compatibility. Teams keep their per-token API gateway, OpenRouter, or LiteLLM for access; o10 changes which model and venue serve each request based on cost, eval floor, and governance rules. The split is intentional: gateways provide doors; control planes enforce economics.

What is a quality floor?

A quality floor is the minimum eval score a model must achieve for a specific use case before o10 routes production traffic to it. Floors are per workload — support, RAG, code, and batch clear at different bars — and measured by replaying representative traffic through eval suites, not assumed from vendor benchmarks. Once a cheaper candidate passes the floor, o10 can route to it in shadow (proof) or enforce (live). Floors without evals are hopes; evals without floors are expensive defaults.

Which providers does o10 support?

o10 fulfills routed traffic across venues including per-token API gateways, OpenRouter, Amazon Bedrock (including your committed Bedrock spend as a venue), and BYOK provider keys. o10 does not own datacenters or hold committed/reserved capacity of its own — committed spend you already have can be a venue o10 routes through. A single control plane sits above all venues.

How are savings verified?

Savings are verified against your own shadow baseline per use case — not industry averages or vendor marketing claims. o10 mirrors a week or more of production traffic, segments by workload, and compares what you actually spent versus what you would have spent on the cheapest eval-passing route at the same quality floor. Finance signs off on the delta before enforce mode flips. Gainshare pricing ties o10 fees to this verified number, so savings must be real and auditable.

How do I integrate with a free route key?

Start free at https://app.o10.io/signup — no card. Signup creates an org and issues a free o10 route key (o10_sk_…), shown once. Point your OpenAI-compatible client at https://app.o10.io/v1, set model to o10/auto (or another mode), and paste the free route key as the API key. Use a placeholder like $O10_ROUTE_KEY in docs — copy your own key after signup; do not use a shared demo key. Shadow mode shows what you'd save before enforce.

What is Free / BYOK?

Free is $0 with no card required. Free is bring-your-own provider keys (BYOK): you route on your own provider keys; o10 can run shadow receipts on live traffic. Paid plans add routed credits. Free does not include prepaid o10-hosted frontier credits.

What are the o10/* modes?

Four modes. o10/auto — default: cheapest model that clears the quality bar. o10/frontier — value-first across frontier-class models (open-weight and proprietary); orgs can set a preferred frontier model. o10/squad — multi-agent planner → workers → judge for hard multi-step work (typically 1–4 minutes). o10/pinned — org-saved model override; if unavailable, falls back to auto with a visible notice. Base URL https://app.o10.io/v1. Concrete slugs and ! pin syntax also work per-request.

What are Frontier Tokens?

Frontier Tokens (model o10/frontier): value-first across frontier-class models — including open-weight (e.g. DeepSeek R1, GLM, Kimi) and proprietary (e.g. GPT-5.5, Claude Opus/Sonnet). Orgs can set a preferred frontier model in settings. Status: beta / live for entitled accounts. Open https://app.o10.io/frontier after signup.

What is o10/pinned?

o10/pinned always routes to the specific model your org saves in settings — an explicit override that skips auto-downgrade. If the pinned model is unavailable, o10 falls back to auto with a visible notice. Concrete slugs and ! pin syntax also still work per-request.

What is Squad and how long does it take?

Squad (model o10/squad): multi-agent planner → workers → judge. For hard multi-step work. Latency is typically 1–4 minutes — not chat-speed. Status: beta / live. Open https://app.o10.io/squad.

What is o10 Tune?

Free prompt optimization. Submit a prompt and 5–20 real examples; o10 searches cheaper models and prompt variants and returns a tuned prompt + recommended model with a before/after receipt. Quality is measured on a held-out split of your examples. Successful tunes are free; failures never charge. Available in console (app.o10.io/tune), API (POST /v1/optimizations), Slack (@o10 tune), and DeepShell.

What is DeepShell?

DeepShell is o10's desktop AI agent (Beta) for business users. Runs tasks on your computer with your files and apps; every action passes an approval gate; a tamper-evident local audit log records what the agent did; recurring tasks can be scheduled; one-click connectors (Google Drive, Notion, Microsoft 365) via OAuth; powered by o10 routing. Free to download; usage runs on o10 credits/BYOK. macOS today; Windows installer in progress. Download at app.o10.io/deepshell.

What is o10 Tag for Slack?

Mention @o10 in Slack; answers are routed through o10 with receipts. Email path: ask@o10.io. Set up in the console under Settings → Integrations → Tag. Tag is available to connect or pilot — Marketplace listing may still lag; do not assume it is already on the Slack Marketplace.

How does o10 guarantee quality?

o10 holds a per-use-case quality floor in the request path. Candidate models must clear holdout-scored evals (and Tune receipts when you optimize) before they can win traffic; failures fall back visibly rather than silently degrading. Shadow mode proves the bar on your traffic before enforce; every enforced call writes a receipt to the immutable ledger.

What is the Quality-Cost Frontier Index?

A privacy-safe monthly dataset from o10: for each task type and quality bar, the cheapest model that empirically clears the bar and its cost per 1K calls, with observation_count and org_count. Cite as “o10 Quality-Cost Frontier Index, <month>”. Canonical page: https://www.o10.io/research/quality-cost-frontier-index. Until a live period publishes, the page shows a clearly labeled ILLUSTRATIVE methodology preview — never sample data presented as measured.

What data does o10 collect for the index?

Metrics only for the public index: task type, quality bar, clearing model, cost per 1K calls, and k-anonymized observation/org counts. Prompt content is never collected for the index. Cells publish only when observation and org thresholds clear.