Run the model where it makes sense. Fast and resident on the laptop, big and always-on for the LAN, served from a desktop-class Apple Silicon host the laptops call as clients.“Local LLM” here means two distinct stacks that answer to the same OpenAI-shaped API, plus the strategy that decides which one runs what. Neither is an agent. Both are just the model plus a serving stack. The agent layer (Claude Code, Antigravity, the routines) lives above them and calls in over HTTP.
- The client laptops:
mlx-lmworkers behindllama-swapon a laptop-class M4 Max, one resident model, tuned to coexist with a working desktop. This is a laptop’s own private model for delegated edits, drafts, and “don’t burn cloud tokens on this” tasks. - The always-on large tier: a desktop-class Apple Silicon host, always on and reachable across the LAN, holds the big resident model the laptops call as clients. One place serves the large tier; everything else is a client of it. This supersedes the earlier dedicated-GPU approach. See Homelab GPU for the path it replaces.
One gateway, one endpoint
A single self-hosted, OpenAI-compatible LLM gateway unifies both stacks (the local models and the always-on workstation-hosted models) behind one endpoint. Every client speaks the same OpenAI API to one URL and one key; the gateway routes to whichever model answers. Model serving spans the always-on workstation-class host for larger models and a separate GPU host for a smaller, faster tier. The gateway selects between them automatically, so no client needs to know which one served the request. This GPU host is distinct from the older dedicated-GPU approach for large models noted earlier, which the second workstation superseded. The homelab’s self-hosted autonomous agents run continuously and connect to that same endpoint as their model backend. A free hosted tier sits behind both local tiers in the fallback order. It is reached only when local capacity is unavailable, never in front of it. See Model routing for the general policy this follows.Current shape, and what’s still gated
The always-on large-tier host (Mac Studio) is live, tuned, and benchmarked, serving two models resident concurrently (see Homelab GPU). It is the intelligence tier: it optimises for answer quality, not throughput. A boot warmup LaunchAgent (mlx-warmup) faults model weights into memory at
boot so the first request does not pay the cold-start cost.
For what that host actually loads right now, read
What is actually loaded, a
dated, live-read roster. No throughput figures are quoted here on purpose:
every earlier number was decode-only and carried no prompt regime, which makes
it uncitable. The re-measured, regime-stated figures live in
Re-measured in isolation.
The historical tuning report is in nix-darwin
MACOS-LLM-PERFORMANCE-TUNING-REPORT.md.
Note that this does not turn two machines into one 256 GB pool. Apple
Silicon unified memory can’t be merged across a cable.
The separate, still-gated direction is combined capacity by sharding. It
means a model too big for either machine alone, run across both over a fast
interconnect, unattended, by morning. That’s a capacity win, not a speed win,
and it remains measurement-gated. See
Distributed & multi-Mac for the honest picture.
The principles that hold across both stacks
- A small resident set, not a rotation. Swapping a multi-GB model evicts
wired GPU memory and reloads another. That’s the slowest thing you can do. A
laptop holds one resident model; the serving host holds two. Capability-role
aliases (
default,coding,quickest,tool-calling, …) resolve onto that resident set, so a role never triggers a swap. How many are resident is a property of the host, read from the registry, not a fixed number. - The registry is the source of truth. Which physical model is resident
lives in one place: the AI-stack registry that
nix-aiwrites at activation, read as~/.config/ai-stack/registry.json. Docs describe the strategy; the current id is a registry value, never hard-coded here. See Models & quantization. - Measure, don’t claim; then verify the instrument. No tuning change ships on a marketing number, only a measured one. But a measurement is only as good as proof of which weights the process loaded, so every number carries that provenance. See Verifying the instrument.
- MoE for throughput. A sparse mixture-of-experts model with a few billion active parameters decodes far faster than a dense model of the same total size. That’s the lever that makes a big model usable on a laptop.
- Capacity, not speed, across machines. Sharding one model over two Macs is communication-bound; reserve it for models that don’t fit, and run two independent workers for everything that does.
In this section
Apple Silicon stack
The M4 Max
mlx-lm + llama-swap stack and every non-secret tuning knob, and why each one is set the way it is.Mac Studio serving
The always-on LAN-shared large-tier model serving host: models, use cases, and headline performance outcomes.
Models & quantization
One-resident posture, MoE vs dense, OptiQ / DWQ / mxfp4, and the fast-vs-overnight model tiers.
Backends & tool calling
Why
vllm-mlx, how it compares to Ollama / llama.cpp / mlx-lm / Rapid-MLX, and the tool-calling reliability problem.Distributed & multi-Mac
The honest two-Mac story: combined capacity via sharding, two-workers-vs-shard, and what’s measurement-gated.
Homelab GPU
The always-on, LAN-shared model on a dedicated GPU: a different machine, a bigger model.
Verifying the instrument
How a server serves the wrong weights silently, the one check that catches it, and the 7x error it hid.
Benchmarking
The reproducible harness and public dataset that every tuning decision is measured against.
How it connects
nix-ai
Packages the inference stack: the MLX model-server LaunchAgent,
llama-swap, the MLX modules, and the AI-stack registry.AI development pipeline
Where local models sit in the bigger picture: routed alongside Claude, Antigravity, and Copilot by task class.
Local AI isolation
Why a local model and the agents calling it still can’t read protected secrets.
Operational reference (private)
Host-specific values, real topology, and incident history live in the gated companion docs.