Muse Glimmer is Meta Superintelligence Labs' 30 billion parameter Apache 2.0 model for always-on local agents. Hosted API input starts at $0.30 to $0.35 per million tokens; local inference is $0. Context window is 131,072 tokens. o10 lists it at band C1 on Together, Fireworks, and OpenRouter.
Last updated: 2026-08-14. Full catalog. Reference prices are dated snapshots. Compare total cost and measured task quality before choosing a route.
Muse Glimmer is a 30-billion-parameter dense multimodal model from Meta Superintelligence Labs, distilled from Muse Spark and released under Apache 2.0 in August 2026. It is built for always-on local agents: sequential tool calls, failure recovery, and text-plus-image input on a Mac or a single consumer GPU. Weights are on Hugging Face as meta-models/Muse-Glimmer-30B.
Meta Superintelligence Labs, August 2026.
How much does Muse Glimmer cost per million tokens?
Local inference is $0 beyond hardware. Hosted: Together AI and Fireworks list $0.35 per million input, $0.04 cached, and $1.50 output. OpenRouter's DeepInfra route lists $0.30 input and $1.20 output. o10's catalog snapshot uses $0.35/$1.50 as the Together/Fireworks reference price; cheaper OpenRouter rows can win the route.
Together, Fireworks, OpenRouter, August 2026.
Can Muse Glimmer run on a laptop?
Yes. Meta sized it for a 24–32 GB GPU envelope once quantized (about 4-bit, under 20 GB for the language weights, plus KV cache, a ~1.8B perception encoder, and optional DFlash drafter). Runtimes called out: llama.cpp, MLX, ExecuTorch, Ollama, LM Studio, Unsloth, vLLM, and SGLang. Validate the quantized checkpoint on your evals; k-quants trade quality for fit.
Meta developer docs, August 2026.
What is Muse Glimmer's context window?
Default context is 131,072 tokens, with longer contexts supported. Architecture is a dense causal transformer with sliding-window local attention (2048) and periodic global layers, plus a ViT-G/14 perception encoder. Input is text and image; output is text.
Hugging Face model card.
Where can you run Muse Glimmer?
Hosted: Together (meta-models/Muse-Glimmer-30B), Fireworks (accounts/fireworks/models/muse-glimmer-30b), and OpenRouter (meta/muse-glimmer-30b) with DeepInfra, Together, and Fireworks backends. Local: Hugging Face weights plus Ollama, LM Studio, llama.cpp. o10 fulfillment currently lists Together, Fireworks, and OpenRouter.
o10 fulfillment overlay 2026-08-14.
What o10 routing band is Muse Glimmer?
Muse Glimmer is band C1. Hosted it undercuts Grok 4.6 and Gemini 3.7 Flash by a wide margin. Locally it is a $0 venue. It will not replace C3/C4 on hard coding floors; use evals. Meta cites MCP Atlas 75.5 and SWE-Bench Pro 51.2 in the size class versus Gemma4-31B and Qwen3.6-27B.
o10 band C1. OpenRouter launch note on evals.
01Deep dive
Specifications
Muse Glimmer in the o10 unified catalog.
Muse Glimmer specifications
Field
Value
o10 slug
`muse-glimmer`
Type
text
Band
C1
Routable
Yes
Source provider
meta
Primary host
Together AI
Gateway ID
`meta/muse-glimmer-30b`
Context window
131,072 tokens
Released
2026-08-10
02Deep dive
Token pricing
Gateway snapshot pricing for Muse Glimmer.
Muse Glimmer token pricing
Unit
Price
Input
$0.35/1M tokens
Output
$1.5/1M tokens
03Deep dive
Fulfillment hosts
Upstream venues configured for Muse Glimmer.
Muse Glimmer fulfillment routes
Host
Upstream model ID
together
together/meta-models/Muse-Glimmer-30B
fireworks
fireworks/muse-glimmer-30b
openrouter
openrouter/meta/muse-glimmer-30b
04Deep dive
Routing with o10
Eval-gated model selection for this slug.
o10 slug: muse-glimmer. Band C1 (economy capable). Routable on hosted venues. Owned/local GPUs are a $0 venue: treat BYO Muse Glimmer as open-weight capacity in the same eval ladder.
See [/models/routing](/models/routing) for smart aliases (o10/auto, o10/frontier, o10/squad) and [/kyi](/kyi) for governance.
05Deep dive
Local $0 versus hosted cents
The economic story is venue, not just the 30B card.
On a quantized 24 GB box, Muse Glimmer has no per-token bill and no data leaves the machine. That is a first-class o10 venue if you already own the GPU: fully loaded cost is power, depreciation, and engineering time.
On Together or Fireworks you pay about $0.35/$1.50 per million and buy burst capacity. OpenRouter can be cheaper on DeepInfra. Route on evals and fully loaded cost, not the Hugging Face headline.
SourceMethodology
Architecture and license from Meta / Hugging Face (Muse-Glimmer-30B, Apache 2.0, ~29.6B, 131,072 context, text+image in). Hosted prices: Together and Fireworks $0.35/$1.50 per 1M; OpenRouter DeepInfra $0.30/$1.20 as of 2026-08-14. Local run is $0 plus hardware. o10 overlay snapshot 2026-08-14.
Local: $0 plus hardware. Hosted reference: $0.35/1M input and $1.50/1M output on Together and Fireworks. OpenRouter DeepInfra: $0.30/$1.20. Cache reads about $0.04/1M.
What is Muse Glimmer's context window?
131,072 tokens by default, with longer contexts supported. Text and image in, text out.
Is Muse Glimmer open source?
Weights are Apache 2.0, Meta's most permissive open-model license to date. Commercial use, modification, and redistribution are allowed under that license. Confirm the Hugging Face card for the exact terms.
Who makes Muse Glimmer?
Meta Superintelligence Labs. It is distilled from Muse Spark, not trained from scratch. About 29.6B parameters including a ~1.8B vision encoder.
How do I call Muse Glimmer via o10?
Send model "muse-glimmer". Gateway/OpenRouter id "meta/muse-glimmer-30b". For owned GPUs, add the local/vLLM endpoint as a venue and keep the same eval floor.
How does o10 route Muse Glimmer?
Band C1, eval-gated. Hosted Muse Glimmer is a likely auto winner for agent-lite jobs that pass the floor. Do not assume it clears a Grok 4.6 or Gemini 3.7 Flash coding eval without a replay.
What is the API model ID for Muse Glimmer?
o10 slug "muse-glimmer"; OpenRouter "meta/muse-glimmer-30b"; Together "meta-models/Muse-Glimmer-30B". Snapshot 2026-08-14.