Morning Edition · №
AI MENLO PARK, CALIF.

Meta Releases Muse Glimmer, a 30B Open-Weight Model Built to Run Locally

The Apache 2.0-licensed model, distilled from Meta's closed Muse Spark 1.2, is tuned for agent tasks and fits on a single consumer GPU with no cloud account required.

Meta Releases Muse Glimmer, a 30B Open-Weight Model Built to Run Locally
Code on a laptop screen. — Photograph: Glen Carrie / Unsplash
SHARE X f in ⧉

Meta Superintelligence Labs released an open-weight artificial intelligence model called Muse Glimmer on Monday, a 30-billion-parameter system the company says is built specifically to run local, always-on AI agents on ordinary consumer hardware rather than in the cloud.

The model is distilled from Muse Spark 1.2, Meta's closed flagship model that launched Aug. 5, using a technique called logit distillation that transfers knowledge from the larger teacher model into the smaller one. Meta has released Muse Glimmer's weights on Hugging Face under an Apache 2.0 license, meaning developers can download, modify and deploy it commercially without paying Meta or routing traffic through Meta's servers, according to the company's announcement.

Meta built Muse Glimmer around agentic work rather than open-ended chat: multi-step reasoning, reliable tool calling, coding and debugging, error recovery when a tool call fails, and multimodal reasoning over screenshots or documents. The company said the model supports more than 100 languages and is designed to handle tasks such as managing schedules and organizing files, in addition to writing and testing code.

Built to fit on one GPU

The release's most notable engineering choice is size. A 30-billion-parameter dense model — meaning all 30 billion parameters activate on every request, unlike sparser "mixture of experts" designs — would normally need more than 55GB of memory at full precision, out of reach for most consumer machines. Meta instead shipped an official 4-bit quantized version that fits under 20GB, small enough to run on a single consumer GPU such as Nvidia's RTX 3090, 4090 or 5090, or on a Mac, without a cloud account, subscription or metered API tokens, per Meta and reporting from VentureBeat.

To keep that compressed model fast, Meta paired it with block-level speculative decoding, using a smaller drafter model that proposes several tokens at once for the larger model to verify, rather than generating one token at a time. Meta reported speedups ranging from roughly 1.5 times on Apple's M4 Max chip to as much as 3.1 times on an Nvidia RTX 5090.

Independent testing published in the days after release found Muse Glimmer competitive with, though not uniformly ahead of, other open agentic models in its size class. On the MCP-Atlas agentic benchmark, Muse Glimmer scored 75.5 against 62.5 for Alibaba's Qwen3.6-27B and 54.2 for Google's Gemma4-31B, and it led on benchmarks including SWE-Bench Pro and GAIA2. Qwen3.6 came out ahead on others, including OSWorld and SWE-Bench Verified, and ran roughly twice as fast in one independently published head-to-head test of local agent tasks, even though Muse Glimmer completed more of them successfully.

The release continues a shift for Meta, which drew criticism from open-source advocates after keeping its most capable Muse Spark models closed. By open-sourcing a smaller, purpose-built distillation instead, Meta is betting developers will build agent tooling around a model they can run and modify freely, even as the underlying flagship model stays proprietary.

SHARE THIS ARTICLE X Facebook LinkedIn Copy link
Claire Fontaine · Technology & Regulation Correspondent

Reports on technology and its regulation for UBStandard, with a focus on Brussels, AI policy and Europe's digital economy.

[email protected]
Related coverage Front page →