. Text version | o10
View formatted page → Copied to clipboard

AI Inference Routing & AI Agent Tools | o10

Source page: https://www.o10.io

Plain-text reference for reading, copying, and citation. Examples are illustrative; use your own credentials for API requests.

Model data: https://www.o10.io/api/models.json
Content index: https://www.o10.io/llms.txt

# o10: control plane for inference spend

o10 is an OpenAI-compatible LLM routing gateway. The inference spend control plane. It holds your quality floor in the request path and routes each call to the cheapest model that clears it across per-token APIs, aggregators, Bedrock, and BYOK venues. Shadow mode, evals, Know Your Inference (KYI), and the Quality-Cost Frontier Index included.

## How to use (free route key)

- Start free at https://app.o10.io/signup. No card. Free route key `o10_sk_…` shown once.
- Free is $0 BYOK (bring your own provider keys); o10 can run shadow receipts. Paid plans add routed credits.
- Point OpenAI-compatible client at `https://app.o10.io/v1`; set `model` to `o10/auto`; paste free route key as API key.
- Snippets use `$O10_ROUTE_KEY` placeholder. Copy your key after signup; no shared demo key.

## Modes

- `o10/auto`. Default: routes each prompt to the cheapest model that still clears the quality bar
- `o10/frontier`. Value-first across frontier-class models, including open-weight (e.g. DeepSeek R1, GLM, Kimi) and proprietary (e.g. GPT-5.5, Claude Opus/Sonnet). Orgs can set a preferred frontier model in settings
- `o10/squad`. Multi-agent: planner → workers → judge. For hard multi-step work (typically 1–4 minutes)
- `o10/pinned`. Org-saved model override. Always routes to the pinned model; if unavailable, falls back to auto with a visible notice
- Concrete model slugs allowed; pin with `!` to skip routing.
- Deprecated (not featured): `o10/economy`, `o10/code`, `o10/auto-cheap`, `o10/auto-fast`.

## Product surfaces

- Frontier Tokens: https://www.o10.io/frontier · app https://app.o10.io/frontier
- Squad: https://www.o10.io/squad · app https://app.o10.io/squad
- o10 Tune: https://www.o10.io/tune · app https://app.o10.io/tune
- DeepShell: https://www.o10.io/deepshell · app https://app.o10.io/deepshell
- o10 Tag (Slack): https://www.o10.io/tag. @o10 / ask@o10.io; Settings → Integrations → Tag
- vs Not Diamond: https://www.o10.io/compare/o10-vs-not-diamond
- vs OpenRouter: https://www.o10.io/compare/o10-vs-openrouter
- vs LiteLLM: https://www.o10.io/compare/o10-vs-litellm
- vs Martian: https://www.o10.io/compare/o10-vs-martian

## Product summary

- **Shadow mode** mirrors traffic and proves savings without changing production routes.
- **Enforce mode** sits in the request path and routes to the cheapest eval-passing model.
- **KYI (Know Your Inference)** scores inference supply chains for board governance.
- **Immutable ledger** records model, venue, policy, and cost on every call.

## Popular comparisons

- [Amazon Nova Lite vs Amazon Nova Micro. Price, context, routing](https://www.o10.io/compare-models/amazon-nova-lite-v1-vs-amazon-nova-micro-v1)
- [Claude Opus 4 vs Claude Opus 4.1. Price, context, routing](https://www.o10.io/compare-models/anthropic-claude-opus-4-vs-anthropic-claude-opus-4-1)
- [Claude Sonnet 4.5 vs Claude Sonnet 4.6. Price, context, routing](https://www.o10.io/compare-models/anthropic-claude-sonnet-4-5-vs-anthropic-claude-sonnet-4-6)
- [Claude Sonnet 4.5 vs Claude Sonnet 4.6. Price, context, routing](https://www.o10.io/compare-models/claude-sonnet-4.5-vs-claude-sonnet-4.6)
- [DeepSeek R1 vs DeepSeek V3.2. Price, context, routing](https://www.o10.io/compare-models/deepseek-deepseek-r1-vs-deepseek-deepseek-v3-2)
- [Gemini 3.7 Flash vs Muse Glimmer. Price, context, routing](https://www.o10.io/compare-models/gemini-3-7-flash-vs-muse-glimmer)
- [Llama 3.1 70B Instruct vs Llama 3.3 70B Instruct. Price, context, routing](https://www.o10.io/compare-models/meta-llama-3-1-70b-instruct-vs-meta-llama-3-3-70b-instruct)
- [GPT-4.1 vs GPT-4.1 mini. Price, context, routing](https://www.o10.io/compare-models/openai-gpt-4-1-vs-openai-gpt-4-1-mini)

## Priority models

- [Antigravity preview](https://www.o10.io/models/antigravity-preview-05-2026)
- [GPT-4o](https://www.o10.io/models/openai-gpt-4o)
- [Amazon Nova Micro pricing](https://www.o10.io/pricing/amazon/amazon-nova-micro-v1)
- [GPT-5 nano pricing](https://www.o10.io/pricing/openai/openai-gpt-5-nano)

## FAQ

### What is o10?

o10 is the control plane for inference spend. It routes every AI inference call to the cheapest model that clears your quality floor, across unified inference gateway, OpenRouter, Amazon Bedrock, and owned capacity. Shadow mode proves savings without changing production; enforce mode holds budget envelopes in the path. Evals define per-use-case quality floors; KYI governs the supply chain for board reporting; an immutable ledger records model, venue, and fully loaded cost on every call.

### What is the difference between shadow and enforce mode?

Shadow mode mirrors live traffic and shows what would have routed and saved, without changing production responses. Enforce mode places o10 in the request path and actively routes each call to the cheapest eval-passing model within your budget envelope. Teams always start in shadow to build a verified per-use-case baseline; finance signs off before enforce flips. Both modes write to the immutable ledger; only enforce changes spend.

### How much can inference routing save?

Savings depend on the current route, task distribution, acceptance criteria, and candidate endpoint costs. The website calculator is illustrative. Compare your baseline with evaluated alternatives and include shadow evaluation, routing, retries, and tool charges before reporting a saving.

### What is Know Your Inference (KYI)?

Know Your Inference (KYI) is a governance framework by o10 that scores inference systems across five weighted pillars: Performance (25%), Economics (25%), Integration (20%), Strategy (20%), and Risk (10%). Each pillar scores 0–100; the composite rolls into a confidence level and board-signable recommendation. KYI runs continuously in the o10 control plane, not as a one-off audit. So every routed call and eval updates the score. A composite floor of 65 triggers enforcement levers: cap, rightsizing, or sunset per policy.

### What is a quality floor?

A quality floor is the minimum eval score a model must achieve for a specific use case before o10 routes production traffic to it. Floors are per workload. Support, RAG, code, and batch clear at different bars, and measured by replaying representative traffic through eval suites, not assumed from vendor benchmarks. Once a cheaper candidate passes the floor, o10 can route to it in shadow (proof) or enforce (live). Floors without evals are hopes; evals without floors are expensive defaults.

### Which providers does o10 support?

o10 unifies routing across per-token API gateways (unified inference gateway), OpenRouter (multi-provider aggregator), Amazon Bedrock (per-token and committed capacity), and owned or open-weight infrastructure. A single control plane sits above all venues. You do not need separate dashboards per provider. o10 selects the cheapest eval-passing route per call and holds budget envelopes. Committed Bedrock drawdown and open-weight routing are first-class venues, not afterthoughts.

### How are savings verified?

Savings are verified against your own shadow baseline per use case, not industry averages or vendor marketing claims. o10 mirrors a week or more of production traffic, segments by workload, and compares what you actually spent versus what you would have spent on the cheapest eval-passing route at the same quality floor. Finance signs off on the delta before enforce mode flips. Gainshare pricing ties o10 fees to this verified number, so savings must be real and auditable.

### How do I integrate with a free route key?

Start free at https://app.o10.io/signup. No card. Signup creates an org and issues a free o10 route key (o10_sk_…), shown once. Point your OpenAI-compatible client at https://app.o10.io/v1, set model to o10/auto (or another mode), and paste the free route key as the API key. Use a placeholder like $O10_ROUTE_KEY in docs. Copy your own key after signup; do not use a shared demo key. Shadow mode shows what you'd save before enforce.

### What is Free / BYOK?

Free is $0 with no card required. Free is bring-your-own provider keys (BYOK): you route on your own provider keys; o10 can run shadow receipts on live traffic. Paid plans add routed credits. Free does not include prepaid o10-hosted frontier credits.

### What are the o10/* modes?

Four modes. o10/auto. Default: cheapest model that clears the quality bar. o10/frontier. Value-first across frontier-class models (open-weight and proprietary); orgs can set a preferred frontier model. o10/squad. Multi-agent planner → workers → judge for hard multi-step work (typically 1–4 minutes). o10/pinned. Org-saved model override; if unavailable, falls back to auto with a visible notice. Base URL https://app.o10.io/v1. Concrete slugs and ! pin syntax also work per-request.

### What are Frontier Tokens?

Frontier Tokens (model o10/frontier): value-first across frontier-class models, including open-weight (e.g. DeepSeek R1, GLM, Kimi) and proprietary (e.g. GPT-5.5, Claude Opus/Sonnet). Orgs can set a preferred frontier model. Status: beta / live for entitled accounts. Open https://app.o10.io/frontier after signup.

### What is Squad and how long does it take?

Squad (model o10/squad): multi-agent planner → workers → judge. For hard multi-step work. Latency is typically 1–4 minutes, not chat-speed. Status: beta / live. Open https://app.o10.io/squad.

### What is o10 Tag for Slack?

Mention @o10 in Slack; answers are routed through o10 with receipts. Email path: ask@o10.io. Set up in the console under Settings → Integrations → Tag. Tag is available to connect or pilot. Marketplace listing may still lag; do not assume it is already on the Slack Marketplace.

## Related links

- [AI inference hub](https://www.o10.io/ai-inference)
- [Modes](https://www.o10.io/models/routing)
- [Frontier Tokens](https://www.o10.io/frontier)
- [Squad](https://www.o10.io/squad)
- [o10 Tune](https://www.o10.io/tune)
- [DeepShell](https://www.o10.io/deepshell)
- [o10 Tag](https://www.o10.io/tag)
- [Model catalog](https://www.o10.io/models)
- [Model comparisons](https://www.o10.io/compare-models)
- [State of Inference Spend 2026](https://www.o10.io/research/state-of-inference-spend-2026)
- [Amazon Nova Lite vs Amazon Nova Micro. Price, context, routing](https://www.o10.io/compare-models/amazon-nova-lite-v1-vs-amazon-nova-micro-v1)
- [Claude Opus 4 vs Claude Opus 4.1. Price, context, routing](https://www.o10.io/compare-models/anthropic-claude-opus-4-vs-anthropic-claude-opus-4-1)
- [Claude Sonnet 4.5 vs Claude Sonnet 4.6. Price, context, routing](https://www.o10.io/compare-models/anthropic-claude-sonnet-4-5-vs-anthropic-claude-sonnet-4-6)
- [Claude Sonnet 4.5 vs Claude Sonnet 4.6. Price, context, routing](https://www.o10.io/compare-models/claude-sonnet-4.5-vs-claude-sonnet-4.6)
- [Antigravity preview](https://www.o10.io/models/antigravity-preview-05-2026)
- [GPT-4o](https://www.o10.io/models/openai-gpt-4o)

## Source URL

https://www.o10.io