Meta releases Muse Glimmer, a 30B offline AI model for desktop GPUs
Meta has released Muse Glimmer, a fully offline AI model with 30 billion parameters that users can run directly on their desktop GPUs. The model is licensed under Apache 2.0, giving developers complete freedom to download, modify and build on the weights without restrictions. Muse Glimmer accepts both text and image inputs, supports over 100 languages and can maintain extremely long conversations up to 131,000 tokens. To make the model usable on consumer hardware, Meta relies on quantization, reducing VRAM requirements from 64GB to under 20GB for the language portion. Two versions are available: K‑Quant‑Dynamic, which needs 32GB VRAM with minimal accuracy loss, and K‑Quant‑17GB, which fits into 24GB VRAM with a slightly larger accuracy hit. Performance is boosted by Meta’s DFlash accelerator, which can triple token generation speeds on high‑end GPUs like the RTX 5090. While Muse Glimmer excels at planning and multi‑step reasoning, it trails competitors in desktop‑control tasks. The model is available now on Hugging Face and LM Studio for users with compatible hardware.
Read the full story on Digital Trends →