oMLX
Ideal For
Run local coding agents on Apple Silicon
Test tool calling and MCP workflows offline
Serve local LLMs, embeddings and rerankers
Build private chat apps with a local backend
Key Strengths
Private on-device inference: data never leaves your machine
Ultra-low latency: local SSD KV caching
OpenAI- and Anthropic-compatible APIs: easy integration
Core Features
Paged SSD KV caching: Reduces time to first token
Continuous batching: Higher throughput
OpenAI-compatible APIs: Easy cloud compatibility
Anthropic-compatible APIs: Broad model access
Native macOS menu bar app: Quick control