oMLX is a local LLM inference server optimized for Apple Silicon and managed from the macOS menu bar or command line. It serves text models, vision-language models, OCR models, embeddings, and rerankers through OpenAI- and Anthropic-compatible APIs. Continuous batching allows concurrent requests, while a tiered KV cache keeps active data in RAM and moves colder blocks to SSD for reuse. Models can be pinned, unloaded manually, or evicted automatically when memory runs low. A web dashboard provides monitoring, chat, downloads, benchmarks, integrations, and per-model settings. It also supports tool calling, structured output, MCP integration, and experimental multi-Mac inference.

Features

  • Continuous batching for concurrent inference
  • RAM and SSD tiered KV caching
  • Multi-model serving and automatic eviction
  • OpenAI and Anthropic API compatibility
  • Built-in administration and chat dashboard
  • Tool calling, MCP, and structured output

Project Samples

Project Activity

See All Activity >

License

Apache License V2.0

Follow oMLX

oMLX Web Site

Other Useful Business Software
Cut Data Warehouse Costs by 54% Icon
Cut Data Warehouse Costs by 54%

Easily migrate from Snowflake, Redshift, or Databricks with free tools.

BigQuery delivers 54% lower TCO with exabyte scale and flexible pricing. Free migration tools handle the SQL translation automatically.
Try Free
Rate This Project
Login To Rate This Project

User Reviews

Be the first to post a review of oMLX!

Additional Project Details

Operating Systems

Mac

Programming Language

Python

Related Categories

Python Large Language Models (LLM)

Registered

16 hours ago