Meta Launches Muse Glimmer, a 30-Billion-Parameter Local Agent Model
The model’s weights are available on Hugging Face and come with documentation that walks users through deployment on popular inference engines such as llama.cpp, MLX, and ExecuTorch. Meta says Muse Glimmer can power local coding assistants, function‑calling workflows, and LLM‑as‑a‑judge evaluations, and it supports OpenClaw and other agent‑orchestration patterns.
To fit the full‑precision 30‑B model into a consumer‑grade GPU, Meta applied 4‑bit weight quantisation, shrinking the 55 GB footprint to under 20 GB of VRAM. The quantised version still leaves space for a KV cache, a perception encoder for image‑text input, and a speculative‑decoding drafter that proposes token blocks for the main model to verify in parallel. Meta’s DFlash‑based drafter reportedly speeds generation without compromising output quality.
In hardware tests, Meta ran the 17‑GB quantised version on a MacBook M4‑Max, MacBook M5‑Max, and an RTX‑5090. The company described the experience as smooth for fluid conversation and real‑time agent interaction, and the public release is expected to support 24‑GB or 32‑GB memory envelopes.
Benchmark results show Muse Glimmer performing strongly against similarly sized models. On the MCP Atlas agentic benchmark, the model scored 75.5, compared with 54.2 for Gemma4‑31B and 62.5 for Qwen3.6‑27B. In DeepSearch QA, Muse Glimmer earned 74.6 against 61.7 for Gemma4‑31B and 71.1 for Qwen3.6‑27B. The model also led in τ²‑Banking (23.5 vs. 15.1 and 16.7) and WildClawBench (47.6 vs. 37.6 and 43.2). On GAIA2, Muse Glimmer scored 43.3, ahead of Gemma4‑31B’s 36.4 and Qwen3.6‑27B’s 40.0.
In coding‑specific tests, Muse Glimmer topped SWE‑Bench Pro with 51.2, slightly ahead of Gemma4‑31B (36.9) and Qwen3.6‑27B (50.2). It also scored 43.6 on SciCode, marginally above Gemma4‑31B’s 43.4. Qwen3.6‑27B led SWE‑Bench Verified (77.2 vs. 76.0) and TerminalBench 2.1 (60.7 vs. 51.7). Meta noted that a local coding agent requires a scaffold that defines which repositories, terminals, and test environments the model can access.
Multimodal benchmarks favour Qwen3.6‑27B in most tests. Muse Glimmer achieved 78.8 on Charxiv Reasoning, slightly higher than Gemma4‑31B (77.7) and Qwen3.6‑27B (78.4). It scored 75.4 on ScreenSpot Pro, just below Qwen3.6‑27B’s 76.1 and Gemma4‑31B’s 75.9. On OmniDocBench v1.5, Muse Glimmer earned 75.8, compared with 77.8 for Qwen3.6‑27B and 72.5 for Gemma4‑31B. The model also scored 74 on MMMU Pro, close to Qwen3.6‑27B’s 75.
Safety evaluations show Muse Glimmer with a lower reported attack success rate than Qwen3.6‑27B. In CI Memories, the model’s violation rate was 26.4 with a coverage score of 64.8, versus Qwen3.6‑27B’s 53.4 and 66.9. In Siren AgentDojo, Muse Glimmer’s attack success rate was 28.4 and utility 94.2, compared with Qwen3.6‑27B’s 40.3 and 92.7.
General reasoning benchmarks indicate Muse Glimmer competes closely with the other two models. It led IFBench (77.0 vs. 76.0 and 70.8), scored 94.7 on AIME 2026 (vs. 89.2 and 94.1), and achieved 80.0 on AA‑LCR. On Beam 128K, it earned 65.1, slightly ahead of Qwen3.6‑27B’s 63.0. Gemma4‑31B led GPQA Diamond (85.7 vs. 83.5 and 84.2) and Humanity’s Last Exam, Text No Tools (23.6 vs. 22.0).
Meta’s release tackles a common constraint for AI teams: cloud‑hosted models require network access and central infrastructure. By offering a local, open‑weight model that fits on consumer GPUs, Meta provides an option for privacy‑sensitive workloads, such as personal agents that access schedules, messages, files, and other private context.
The company said integrations with llama.cpp, MLX, and ExecuTorch will arrive in the coming days, and it plans to release an open‑weight version of Muse Spark in the near future. For developers interested in local deployment, the Hugging Face repository contains 4‑bit and 8‑bit quantised builds, and the model’s documentation outlines how to configure the DFlash drafter and set up retry training for failed tool calls.
In summary, Muse Glimmer is Meta’s first open‑weight model from its Superintelligence Labs division. It offers a 30‑billion‑parameter, multimodal, local‑agent solution that performs competitively across agentic, coding, multimodal, safety, and reasoning benchmarks. The model’s availability under an Apache 2.0 license and its compatibility with consumer GPUs make it a practical choice for developers seeking to run advanced AI workflows offline.