Muse Glimmer
Local multimodal 30B model for autonomous agents, coding, and tools
...Distilled from the larger Muse Spark, it combines multi-step reasoning, reliable tool use, coding, failure recovery, and image understanding in a dense 29.6B-parameter architecture with a dedicated 1.8B-parameter perception encoder. It supports more than 100 languages and a 131K+ token context window, allowing agents to maintain coherent plans across extended workflows. Muse Glimmer can interpret screenshots, charts, documents, and images alongside text, while configurable reasoning strength lets developers balance response quality and speed. Its quantized variants reduce the model below 20 GB for operation on systems with 24–32 GB of memory, and DFlash speculative decoding can substantially accelerate generation.