oMLX is a local LLM inference server optimized for Apple Silicon and managed from the macOS menu bar or command line. It serves text models, vision-language models, OCR models, embeddings, and rerankers through OpenAI- and Anthropic-compatible APIs. Continuous batching allows concurrent requests, while a tiered KV cache keeps active data in RAM and moves colder blocks to SSD for reuse. Models can be pinned, unloaded manually, or evicted automatically when memory runs low. A web dashboard provides monitoring, chat, downloads, benchmarks, integrations, and per-model settings. It also supports tool calling, structured output, MCP integration, and experimental multi-Mac inference.

Features

  • Continuous batching for concurrent inference
  • RAM and SSD tiered KV caching
  • Multi-model serving and automatic eviction
  • OpenAI and Anthropic API compatibility
  • Built-in administration and chat dashboard
  • Tool calling, MCP, and structured output

Project Samples

Project Activity

See All Activity >

License

Apache License V2.0

Follow oMLX

oMLX Web Site

Other Useful Business Software
Build Agents and Models on One Platform Icon
Build Agents and Models on One Platform

Everything you need to build production-ready agents and models. Access 200+ Google and third-party AI models and tools.

Gemini Enterprise Agent Platform is Google Cloud's comprehensive platform for developers to build, scale, govern, and optimize agents and models. Choose from Google's most advanced models and third-party models like Anthropic's Claude Model Family.
Try It Free
Rate This Project
Login To Rate This Project

User Reviews

Be the first to post a review of oMLX!

Additional Project Details

Operating Systems

Mac

Programming Language

Python

Related Categories

Python Large Language Models (LLM)

Registered

1 day ago