. Text version | o10
View formatted page → Copied to clipboard

Muse Glimmer. Specs, price, and routing

Source page: https://www.o10.io/models/muse-glimmer

Plain-text reference for reading, copying, and citation. Examples are illustrative; use your own credentials for API requests.

Model data: https://www.o10.io/api/models.json
Content index: https://www.o10.io/llms.txt
Prices are dated references, not live quotes. Check the current endpoint rate before use.

# Muse Glimmer. Specs, price, and routing

Muse Glimmer is Meta Superintelligence Labs' 30 billion parameter Apache 2.0 model for always-on local agents. Hosted API input starts at $0.30 to $0.35 per million tokens; local inference is $0. Context window is 131,072 tokens. o10 lists it at band C1 on Together, Fireworks, and OpenRouter.

## Key takeaways

### What is Muse Glimmer?

Muse Glimmer is a 30-billion-parameter dense multimodal model from Meta Superintelligence Labs, distilled from Muse Spark and released under Apache 2.0 in August 2026. It is built for always-on local agents: sequential tool calls, failure recovery, and text-plus-image input on a Mac or a single consumer GPU. Weights are on Hugging Face as meta-models/Muse-Glimmer-30B.

*Meta Superintelligence Labs, August 2026.*

### How much does Muse Glimmer cost per million tokens?

Local inference is $0 beyond hardware. Hosted: Together AI and Fireworks list $0.35 per million input, $0.04 cached, and $1.50 output. OpenRouter's DeepInfra route lists $0.30 input and $1.20 output. o10's catalog snapshot uses $0.35/$1.50 as the Together/Fireworks reference price; cheaper OpenRouter rows can win the route.

*Together, Fireworks, OpenRouter, August 2026.*

### Can Muse Glimmer run on a laptop?

Yes. Meta sized it for a 24–32 GB GPU envelope once quantized (about 4-bit, under 20 GB for the language weights, plus KV cache, a ~1.8B perception encoder, and optional DFlash drafter). Runtimes called out: llama.cpp, MLX, ExecuTorch, Ollama, LM Studio, Unsloth, vLLM, and SGLang. Validate the quantized checkpoint on your evals; k-quants trade quality for fit.

*Meta developer docs, August 2026.*

### What is Muse Glimmer's context window?

Default context is 131,072 tokens, with longer contexts supported. Architecture is a dense causal transformer with sliding-window local attention (2048) and periodic global layers, plus a ViT-G/14 perception encoder. Input is text and image; output is text.

*Hugging Face model card.*

### Where can you run Muse Glimmer?

Hosted: Together (meta-models/Muse-Glimmer-30B), Fireworks (accounts/fireworks/models/muse-glimmer-30b), and OpenRouter (meta/muse-glimmer-30b) with DeepInfra, Together, and Fireworks backends. Local: Hugging Face weights plus Ollama, LM Studio, llama.cpp. o10 fulfillment currently lists Together, Fireworks, and OpenRouter.

*o10 fulfillment overlay 2026-08-14.*

### What o10 routing band is Muse Glimmer?

Muse Glimmer is band C1. Hosted it undercuts Grok 4.6 and Gemini 3.7 Flash by a wide margin. Locally it is a $0 venue. It will not replace C3/C4 on hard coding floors; use evals. Meta cites MCP Atlas 75.5 and SWE-Bench Pro 51.2 in the size class versus Gemma4-31B and Qwen3.6-27B.

*o10 band C1. OpenRouter launch note on evals.*

## Citation
Muse Glimmer is Meta Superintelligence Labs' 30 billion parameter Apache 2.0 model for always-on local agents. Hosted API input starts at $0.30 to $0.35 per million tokens; local inference is $0. Context window is 131,072 tokens. o10 lists it at band C1 on Together, Fireworks, and OpenRouter.
## Methodology

Architecture and license from Meta / Hugging Face (Muse-Glimmer-30B, Apache 2.0, ~29.6B, 131,072 context, text+image in). Hosted prices: Together and Fireworks $0.35/$1.50 per 1M; OpenRouter DeepInfra $0.30/$1.20 as of 2026-08-14. Local run is $0 plus hardware. o10 overlay snapshot 2026-08-14.

## FAQ

### How much does Muse Glimmer cost?

Local: $0 plus hardware. Hosted reference: $0.35/1M input and $1.50/1M output on Together and Fireworks. OpenRouter DeepInfra: $0.30/$1.20. Cache reads about $0.04/1M.

### What is Muse Glimmer's context window?

131,072 tokens by default, with longer contexts supported. Text and image in, text out.

### Is Muse Glimmer open source?

Weights are Apache 2.0, Meta's most permissive open-model license to date. Commercial use, modification, and redistribution are allowed under that license. Confirm the Hugging Face card for the exact terms.

### Who makes Muse Glimmer?

Meta Superintelligence Labs. It is distilled from Muse Spark, not trained from scratch. About 29.6B parameters including a ~1.8B vision encoder.

### How do I call Muse Glimmer via o10?

Send model "muse-glimmer". Gateway/OpenRouter id "meta/muse-glimmer-30b". For owned GPUs, add the local/vLLM endpoint as a venue and keep the same eval floor.

### How does o10 route Muse Glimmer?

Band C1, eval-gated. Hosted Muse Glimmer is a likely auto winner for agent-lite jobs that pass the floor. Do not assume it clears a Grok 4.6 or Gemini 3.7 Flash coding eval without a replay.

### What is the API model ID for Muse Glimmer?

o10 slug "muse-glimmer"; OpenRouter "meta/muse-glimmer-30b"; Together "meta-models/Muse-Glimmer-30B". Snapshot 2026-08-14.

## Related links

- [Muse Glimmer HTML](https://www.o10.io/models/muse-glimmer)
- [Muse Glimmer pricing](https://www.o10.io/pricing/meta/muse-glimmer)
- [Model catalog](https://www.o10.io/models)
- [llms-models.txt](https://www.o10.io/llms-models.txt)

## Source URL

https://www.o10.io/models/muse-glimmer