oMLX: Why This Mac Inference Server Caches to Your SSD (And Why That's Brilliant)
Hook
Most local LLM servers throw away your computed KV cache when memory runs low. oMLX serializes it to your SSD in safetensors format and restores it on the next request—even after a restart. For coding agents that repeatedly query the same massive codebase, this changes everything.
Context
Running LLMs locally on Apple Silicon has become viable thanks to MLX, but serving them efficiently is a different problem. Tools like llama.cpp and Ollama made it trivial to load a model and chat with it, but they fall apart when you need to handle concurrent requests or work with long contexts. Every new request recomputes the entire context window from scratch, and there's no batching to amortize prefill costs across users.
The real pain point emerges when you're running coding agents—tools like Cursor, Aider, or custom RAG pipelines that send 50,000+ token prompts containing your entire codebase as context. You iterate on a feature, send ten requests in an hour, and each one burns 30 seconds just reprocessing the same code files. Cloud APIs like Claude or GPT-4 handle this gracefully with prompt caching, but local tools treat every request as a blank slate. oMLX was built to solve this: a production-grade inference server that implements continuous batching for concurrency and a novel two-tier KV cache that persists computed context to disk, making context-heavy workflows dramatically faster.
Technical Insight
At its core, oMLX wraps mlx-lm's BatchGenerator with a tiered caching layer that treats memory and storage as cooperative resources rather than forcing an either/or choice. The hot tier lives in RAM using block-based prefix sharing with Copy-on-Write semantics—when multiple requests share a common prefix (like the same system prompt or codebase context), they reference the same underlying KV cache blocks until generation diverges. This is standard vLLM-style PagedAttention, adapted for MLX's unified memory model.
The breakthrough is the cold tier: when memory pressure triggers eviction, oMLX doesn't discard those computed KV blocks. Instead, it serializes the least-recently-used blocks to disk in safetensors format, storing them in ~/.omlx/cache/ with filenames keyed by content hash. When a future request matches that prefix, the server deserializes from SSD instead of recomputing. This survives server restarts, making it viable to build up a persistent "context library" for your most-used prompts.
Here's what the cache restoration flow looks like in practice:
# Simplified from omlx/cache/kv_cache.py
class TieredKVCache:
def get_or_compute(self, prompt_tokens, model):
prefix_hash = self._hash_tokens(prompt_tokens)
# Check hot tier (in-memory blocks)
if prefix_hash in self.hot_cache:
return self.hot_cache[prefix_hash]
# Check cold tier (SSD-backed safetensors)
cache_path = self.cache_dir / f"{prefix_hash}.safetensors"
if cache_path.exists():
kv_blocks = self._deserialize_from_disk(cache_path)
self.hot_cache[prefix_hash] = kv_blocks # Promote to hot
return kv_blocks
# Cache miss: compute fresh and persist both tiers
kv_blocks = model.prefill(prompt_tokens)
self.hot_cache[prefix_hash] = kv_blocks
self._serialize_to_disk(kv_blocks, cache_path)
return kv_blocks
The economics here are fascinating. On an M3 Ultra with NVMe storage, deserializing 50K tokens worth of KV cache from disk takes ~400ms. Recomputing that same prefix takes 4-8 seconds depending on the model. You pay the serialization cost once, then every subsequent request with that prefix gets an 10-20x speedup. For coding agents that iterate on the same codebase, this compounds—your fifth request of the day might hit 100% cached and skip prefill entirely.
Continuous batching is handled via mlx-lm's BatchGenerator, which pools incoming requests and processes them in a single forward pass. Unlike vLLM's sophisticated scheduler, this is relatively simple—no speculative decoding, no chunked prefill, no priority queues. But it's enough to handle 4-8 concurrent requests without latency cliffs, and it leverages all of MLX's Metal kernel optimizations without requiring custom CUDA-style kernels.
The server exposes OpenAI and Anthropic-compatible REST endpoints, which means drop-in replacement for existing agent codebases:
# Standard OpenAI chat completion
curl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "qwen2.5-coder-7b",
"messages": [{"role": "user", "content": "Explain this code..."}],
"stream": true
}'
But the real power user feature is model profiles. You can load a single set of weights (say, qwen2.5-coder-7b) and expose it under multiple API names with different sampling parameters, chat templates, or tool-calling configurations:
# ~/.omlx/profiles.yaml
profiles:
- name: "qwen3-8b:thinking"
base_model: "qwen2.5-coder-7b"
temperature: 0.9
top_p: 0.95
system_prompt: "Think step-by-step before answering."
- name: "qwen3-8b:precise"
base_model: "qwen2.5-coder-7b"
temperature: 0.1
top_k: 10
system_prompt: "Be concise and technical."
Now API consumers can request model: "qwen3-8b:thinking" vs model: "qwen3-8b:precise" and get completely different behavior with zero additional memory overhead—the same loaded weights, just different inference-time parameters applied. This is shockingly useful for A/B testing agent prompting strategies without spinning up multiple model instances.
The native Swift menubar app handles lifecycle management—auto-starting the Python server, crash recovery, log aggregation, and a web-based admin dashboard. The dashboard isn't just monitoring; it's a full CRUD interface for adjusting sampling params, setting per-model TTLs, and inspecting cache hit rates. This is Mac-native UX that makes Python CLI tools feel primitive by comparison.
Gotcha
The SSD cache performance assumption is that you have fast NVMe storage. If you're running this on a Mac with SATA SSDs, external drives, or network-attached storage, the cache restoration overhead can exceed recomputation time. There are no documented benchmarks comparing cache miss vs. restore latency across different storage tiers, so you're flying blind until you profile your own hardware. A 4TB external Thunderbolt drive might sound appealing for cache capacity, but if it can't sustain 1GB/s reads, you've just made your inference slower.
Tool calling support claims to handle 11+ model families, but the implementation is brittle. It works by regex-parsing chat template strings and matching against hard-coded model name patterns. Custom-trained models or chat templates with non-standard formatting will silently fail—your model will return tool calls that the parser can't extract, and debugging requires reading through omlx/tools/parser.py to understand which regex failed. There's no schema validation or fallback to raw text extraction. If you're using anything outside the mainstream model families (Qwen, Llama, Mistral, DeepSeek), expect to spend time debugging tool call extraction.
The memory guard tiers (safe, balanced, aggressive) abstract eviction behavior but document nothing about the actual formulas. The default safe mode reserves 8GB for macOS, which is reasonable for a Mac Studio but borderline reckless for a MacBook Air with 16GB total RAM. Under concurrent load, you'll hit swap, and SSD thrashing will destroy inference throughput. The admin dashboard shows memory usage graphs but no predictive warnings before you hit the cliff.
Verdict
Use if: You're running coding agents locally on Apple Silicon with long, repetitive context (same codebase across multiple requests), you have NVMe storage and 32GB+ RAM, and you want a polished Mac-native server that handles concurrency without manual batching gymnastics. The persistent SSD cache genuinely changes the economics of local inference for iterative workflows. Also use if you need zero-config OpenAI API compatibility and don't want to fight Docker or environment setup—the menubar app makes this the smoothest local inference deployment on macOS.
Skip if: You need maximum batching throughput for high-concurrency workloads (vLLM on NVIDIA hardware will smoke this), you're on slower storage (SATA, external drives), or you need advanced features like speculative decoding or chunked prefill. Also skip if you're working with custom or niche models outside the mainstream families—tool calling auto-detection will break, and you'll spend more time debugging parsers than running inference. Finally, skip if you're not on Apple Silicon; this is a Mac-first tool with no cross-platform ambitions.