oMLX is a local LLM inference server optimized for Apple Silicon and managed from the macOS menu bar or command line. It serves text models, vision-language models, OCR models, embeddings, and rerankers through OpenAI- and Anthropic-compatible APIs. Continuous batching allows concurrent requests, while a tiered KV cache keeps active data in RAM and moves colder blocks to SSD for reuse. Models can be pinned, unloaded manually, or evicted automatically when memory runs low. A web dashboard provides monitoring, chat, downloads, benchmarks, integrations, and per-model settings. It also supports tool calling, structured output, MCP integration, and experimental multi-Mac inference.

Features

  • Continuous batching for concurrent inference
  • RAM and SSD tiered KV caching
  • Multi-model serving and automatic eviction
  • OpenAI and Anthropic API compatibility
  • Built-in administration and chat dashboard
  • Tool calling, MCP, and structured output

Project Samples

Project Activity

See All Activity >

License

Apache License V2.0

Follow oMLX

oMLX Web Site

Other Useful Business Software
MongoDB Atlas runs apps anywhere Icon
MongoDB Atlas runs apps anywhere

Deploy in 115+ regions with the modern database for every enterprise.

MongoDB Atlas gives you the freedom to build and run modern applications anywhere—across AWS, Azure, and Google Cloud. With global availability in over 115 regions, Atlas lets you deploy close to your users, meet compliance needs, and scale with confidence across any geography.
Start Free
Rate This Project
Login To Rate This Project

User Reviews

Be the first to post a review of oMLX!

Additional Project Details

Operating Systems

Mac

Programming Language

Python

Related Categories

Python Large Language Models (LLM)

Registered

21 hours ago