Llama 3.1 8B Instruct costs $0.120/1M input tokens and $0.120/1M output tokens as of 2026-06-27. Meta publishes this model with a 128K tokens context window. o10 lists Llama 3.1 8B Instruct on NVIDIA NIM at routing band C0. API slug meta-llama-3-1-8b-instruct. Verify list prices against the venue you call.

Last updated: 2026-06-27. Full catalog. Reference prices are dated snapshots. Compare total cost and measured task quality before choosing a route.

Llama 3.1 8B Instruct. Price, context, and routing

AnswersSelf-contained

What is Llama 3.1 8B Instruct?

Llama 3.1 8B Instruct is a chat model from meta (slug: meta-llama-3-1-8b-instruct). Pick Llama 3.1 8B Instruct when you want Meta Llama on a fulfillment host (or your own GPUs) instead of a closed lab SKU. Published input pricing starts at $0.120/1M tokens. Context window: 128K tokens.

Catalog snapshot 2026-06-27.

When should you pick Llama 3.1 8B Instruct?

Pick Llama 3.1 8B Instruct when you want Meta Llama on a fulfillment host (or your own GPUs) instead of a closed lab SKU. Llama 3.1 8B Instruct is a Llama instruct checkpoint. Smaller Llama SKUs cut cost; larger ones raise quality. Prove the swap in shadow.

How does Llama 3.1 8B Instruct differ from sibling models?

Llama 3.1 8B Instruct is a Llama instruct checkpoint. Smaller Llama SKUs cut cost; larger ones raise quality. Prove the swap in shadow.

How much does Llama 3.1 8B Instruct cost per million tokens?

Llama 3.1 8B Instruct costs $0.120/1M input tokens and $0.120/1M output tokens in the gateway catalog snapshot (2026-06-27). Endpoint pricing may differ by fulfillment host.

Sourced from gateway catalog snapshot.

Where can you run Llama 3.1 8B Instruct?

Llama 3.1 8B Instruct is configured on NVIDIA NIM. Known host IDs: nvidia_nim: nvidia_nim/meta/llama-3.1-8b-instruct. This model is routable through o10 eval-gated routing.

What is Llama 3.1 8B Instruct's context window?

Llama 3.1 8B Instruct supports a 128K tokens context window in the gateway catalog snapshot. Context limits may vary by fulfillment host endpoint.

What o10 routing band is Llama 3.1 8B Instruct?

Llama 3.1 8B Instruct is classified at o10 routing band C0. Bands group models by cost tier and eval profile: C0 (economy), C1 (standard), C2 (capable), C3/C4 (frontier). Production routing is enabled for this slug.

Who makes Llama 3.1 8B Instruct?

Llama 3.1 8B Instruct is attributed to source provider meta in the o10 catalog. Fulfillment hosts (NVIDIA NIM) supply API access; list pricing may differ by venue.

01Deep dive

Specifications

Llama 3.1 8B Instruct in the o10 unified catalog.

Llama 3.1 8B Instruct specifications
FieldValue
o10 slug`meta-llama-3-1-8b-instruct`
Typetext
BandC0
RoutableYes
Source providermeta
Primary hostNVIDIA NIM
Context window128,000 tokens
02Deep dive

Token pricing

Gateway snapshot pricing for Llama 3.1 8B Instruct.

Llama 3.1 8B Instruct token pricing
UnitPrice
Input$0.12/1M tokens
Output$0.12/1M tokens
03Deep dive

Fulfillment hosts

Upstream venues configured for Llama 3.1 8B Instruct.

Llama 3.1 8B Instruct fulfillment routes
HostUpstream model ID
nvidia_nimnvidia_nim/meta/llama-3.1-8b-instruct
04Deep dive

Routing with o10

Eval-gated model selection for this slug.

o10 slug: meta-llama-3-1-8b-instruct. Band: C0. Routable: yes.

See [/models/routing](/models/routing) for smart aliases (o10/auto, o10/frontier, o10/squad) and [/kyi](/kyi) for governance.

SourceMethodology

Fulfillment catalog 2026-06-27. Gateway pricing joined where slug match exists.

FAQFrequently asked questions

Common questions

How much does Llama 3.1 8B Instruct cost?

Llama 3.1 8B Instruct costs $0.120/1M input tokens and $0.120/1M output tokens in the gateway snapshot (2026-06-27). Verify against your provider's published pricing.

What is Llama 3.1 8B Instruct's context window?

Llama 3.1 8B Instruct supports 128K tokens in the gateway catalog snapshot. Host-specific endpoints may differ.

Which hosts serve Llama 3.1 8B Instruct?

Llama 3.1 8B Instruct (meta-llama-3-1-8b-instruct) is configured on: NVIDIA NIM. Upstream model IDs are listed in the fulfillment routes section.

Is Llama 3.1 8B Instruct routable through o10?

Yes. Llama 3.1 8B Instruct is routable at band C0. Send model: "meta-llama-3-1-8b-instruct" or use an o10 smart routing alias.

How do I call Llama 3.1 8B Instruct via o10?

Use o10 slug "meta-llama-3-1-8b-instruct" in the model field. See /models/routing for smart aliases (o10/auto, o10/frontier, o10/squad).

How does o10 route this model?

o10 routes Llama 3.1 8B Instruct only when per-use-case evals clear at your quality floor. Band C0 indicates cost tier. See /kyi for governance and /models/routing for smart aliases.

What is the API model ID for Llama 3.1 8B Instruct?

The o10 slug is "meta-llama-3-1-8b-instruct". Host IDs: nvidia_nim: nvidia_nim/meta/llama-3.1-8b-instruct. Pricing snapshot date: 2026-06-27.

Who hosts Llama 3.1 8B Instruct?

Llama 3.1 8B Instruct is hosted as nvidia_nim: nvidia_nim/meta/llama-3.1-8b-instruct. o10 maps those upstream IDs to slug "meta-llama-3-1-8b-instruct".