Skip to main content
Two memory behaviours ride on one flag, over a wired ceiling that moves with the serving mode. The stack tunes them at runtime. A model or mode swap needs no rebuild. But the GPU cap alone is necessary, not sufficient.
A local MLX serving stack is a swapping proxy in front of one or more inference workers. It has a single flag, --gpu-memory-utilization, that drives two different memory behaviours at once. Reading it as one knob is the usual reason a well-sized host still ends up paging.

The two behaviours

The flag is a fraction. It feeds two independent calculations:
  • allocation_limit is the per-worker Metal allocation cap. A worker that tries to allocate past it fails the allocation.
  • trip is the engine’s emergency KV-cache clear. When resident memory crosses it, the engine drops the cache to recover headroom.
The two use different denominators, and that is the whole story: Raising the wired ceiling raises allocation_limit but leaves trip exactly where it was.

Insight one: it is a cap, not a reservation

--gpu-memory-utilization does not pre-allocate anything. Setting it to a large fraction does not hand that memory to the worker. Setting it small does not hold memory back for the desktop either. It only says how far an allocation is allowed to go before it is refused. The knob that actually decides how much of a host goes to inference is the OS wired ceiling. Treat --gpu-memory-utilization as a safety rail on top of that decision, not as the decision itself.

Insight two: the two knobs move together

The invariant to hold is:
The trip point is a soft valve. It has to fire before the host reaches its hard ceiling and starts paging. It also has to clear the worker’s steady resident footprint so it does not fire during normal work. Move either knob alone and one side of that invariant breaks:

Lowering utilization alone

The trip point drops below the resident working set. The engine clears the KV cache constantly, every turn re-prefills from scratch, and multi-turn sessions degrade badly while the GPU looks busy.

Raising the ceiling alone

The trip point stays put while the hard ceiling climbs past it. Or the trip point already lands past the ceiling entirely. Either way, the soft valve never fires, and the host reaches its hard limit first.
Neither knob is safe to change by itself. Any change to the wired ceiling needs a matching re-derivation of --gpu-memory-utilization, and the reverse.

Re-deriving the utilization

Rearranging the invariant gives a symbolic band for the flag:
The lower bound keeps the allocation cap higher than what the worker actually holds. The upper bound keeps the trip point below the hard ceiling, with the engine’s own 0.05 offset subtracted out. Pick a value inside the band with margin on both sides, and re-derive it whenever the ceiling, the model, or the batch shape changes. If the band is empty (the lower bound exceeds the upper), no value of the flag satisfies the invariant on that host. Raise the ceiling, shrink the footprint (smaller cache reservation, fewer concurrent sequences, a smaller model), or both.

Measuring the footprint: use wired, not RSS

On macOS, ps RSS does not attribute kernel-wired Metal allocations to the process that caused them. A worker holding tens of GB of wired GPU memory can look small in ps while top shows the real figure. Size against the wired number; RSS quietly tells you the host has room it does not have. Wired is the right number for the GPU footprint, but it is not the whole process. On unified memory a worker also holds CPU-side tensors, memory-mapped weight pages, and load-time copies that the wired cap does not charge. Bounding wired memory is necessary but not sufficient. See the honest guarantee below, which bounds the whole process instead.

Finding the real ceiling

The default wired ceiling on Apple Silicon is commonly quoted as about 75% of RAM. In practice it can be a substantially larger fraction than that. The rule of thumb is a guess; measure the machine instead.
1

Read the ceiling from MLX

Query max_recommended_working_set_size through MLX. That is the value the allocator actually enforces, whatever the sysctl is currently set to. Read it before changing anything. The default may already be higher than you were about to set.
2

Set the sysctl, then re-read

Change iogpu.wired_limit_mb and read max_recommended_working_set_size again to confirm the new ceiling took effect.
3

Re-derive utilization

Feed the confirmed ceiling back through the preceding band and set --gpu-memory-utilization to a value inside it.
4

Confirm the engine actually applied it

Read back the limit the engine reports at startup, not just the flag you passed. Serving stacks build their engine down more than one code path. A flag honored on one path can be silently ignored on another, leaving the default in force while your command line says otherwise. Trust the engine’s own log line, then check the resulting trip point still satisfies the invariant.
Passing --gpu-memory-utilization is not proof it took effect. If an idle auto-unload or lazy-load feature routes model construction through a secondary path, the value can revert to the engine default there. The only reliable signal is the allocation limit the engine prints once the model loads. If it does not match your flag, the invariant you derived is not the one running.

The unit is MiB, and it is exact

iogpu.wired_limit_mb is in MiB, and MLX reports the ceiling as:
There is no rounding and no decimal conversion in that step. So a sysctl value read as “roughly N thousand MB” produces a ceiling about 7.4% larger than the same number read as decimal GB. That’s because 1024² bytes per MiB is that much more than 1000² bytes per MB. That gap matters because it lands on the wrong side of safety. The ceiling you get is bigger than the one you thought you set, so util × ceiling leaves more allocation headroom than intended. Meanwhile trip is computed against total_ram, untouched by the sysctl, so it stays exactly where it was. Always re-derive from the value MLX reports back, never from the number you typed into the sysctl. Weights carry the mirror-image trap. Hugging Face reports a model’s size in decimal GB (×1e9 bytes); every budget in the admission formula is in GiB (×1024³). Convert weightsBytes with ×1e9, not ×1024³. Conflate the two and the subtracted weight inflates by 7.37%, silently shrinking the KV pool you thought you had.

Bounding the footprint: which flag is the hard cap

The invariant needs a real bound on footprint. Two flags look like they set it; only one does.
  • --max-cache-blocks is the hard cap. It bounds the active paged-KV pool, the memory that grows as concurrent requests and longer contexts arrive. When the pool runs out, the worker rejects the request instead of allocating past the bound. This is the flag that keeps footprint under trip.
  • --cache-memory-mb bounds only the prefix cache. The prefix cache is a separate store that reuses shared prompt prefixes across requests. Sizing against it leaves the paged-KV pool unbounded, and that pool is the part that actually grows under load.
Size the footprint against --max-cache-blocks. A host tuned on --cache-memory-mb alone can still cross trip under concurrency, because the pool that grew was never the one it capped.

A context limit includes the answer

The context number exposed by a serving profile is a total admission limit. It is not an input-only promise. Before a worker accepts a request, it must reserve room for both the prompt and the allowed completion:
Keep the values separate in configuration, runtime records, and benchmark reports:
  • Model-native maximum: what the checkpoint architecture can support.
  • Configured total context: what this serving profile is willing to admit.
  • Actual prompt tokens: what the server received for one request.
  • Output reservation: the completion room protected at admission.
  • Actual total context: what the request used after completion.
This distinction prevents two common mistakes. First, a 32k prompt with a 512-token reservation does not fit a 32k total limit. It needs a larger total budget or a smaller prompt. Second, enabling a 192k total limit does not turn a 32k input into a 192k benchmark. It creates a useful window-sensitivity case: measure the same 32k input under both configurations, then report the actual prompt and configured limit together. Context is only one part of memory use. The active paged-KV pool grows with actual admitted requests. The prefix cache is separate shared storage. Model weights, operating-system cache, and load-time workspace use also affect headroom without changing a request’s token count. A safe long-context policy therefore reserves output first, admits against the global KV pool, and records which cache states applied to the measurement.

Safety is runtime, not a build assert

Holding footprint < trip < ceiling by hand means two coupled knobs move together: the wired ceiling and --gpu-memory-utilization. Both must be re-derived every time the model, the mode, or the batch shape changes. A build could derive them once and assert the result, but that only covers the config it was built with. Models and modes change while the host runs, and rebuilding to swap a model is exactly the friction the stack avoids. So the safety lives at runtime, at two enforcement points. Neither needs a darwin-rebuild. Swapping a model or switching mode is as cheap as changing a model in a proxy router.

One budget input per host, per mode

A host declares one number, maxLocalLlmGb: the total memory the local LLM stack may use. It carries two values: one for standalone serving, and a larger one for clustered mode. In clustered mode the desktop is quiesced, so more memory is free for a shard. The operator sets the numbers; the running mode picks which one applies.

Enforcement point one: the live mode-aware sysctl cap

The OS wired ceiling, iogpu.wired_limit_mb, is a live sysctl. A link watcher flips it to the current mode’s budget the moment the cluster cable is plugged or unplugged. Plugging in applies the clustered budget; unplugging applies the standalone budget. No reboot, no rebuild. In clustered mode both Macs raise the same higher cap. This is the hard kernel backstop on GPU-wired memory, and it moves with the mode. It stays high, set to the mode’s budget, and is never lowered as a quick fix.

Enforcement point two: per-model load-time admission control

When the swap proxy spawns a worker, a wrapper computes the safe caps from live inputs. Those inputs are the current sysctl value, the model’s architecture, and the requested concurrency and context. Swapping a model needs no nix change:
--max-cache-blocks is that global pool count: the whole paged-KV pool, shared across every concurrent sequence, not a per-request budget. So the pool size does not divide by concurrency. Concurrency enters at admission instead:
The wrapper sets --max-cache-blocks to poolBlocks and derives --gpu-memory-utilization from the preceding live ceiling band. If the requested concurrency and context do not fit the pool, the wrapper rejects or clamps the load and fails loud. It shrinks concurrency or context until the admission check holds. A model too big to fit even one sequence is refused, not paged to disk.
kvLayers is the count of layers that actually hold a KV cache, not the total layer count. A hybrid-attention model charges KV only for its full-attention layers. qwen3-next-80b runs 12 full-attention layers of 48, so its KV costs about 24 KiB/token. That’s cheaper than a dense 30 B at about 96 KiB/token, despite the far larger weights. Read kvLayers from the model’s architecture; never assume it equals the depth.
One function covers every case by changing only the model or the mode:
  • a tiny model at high concurrency: small weights leave a large KV budget, so many sequences fit.
  • one large model with a long context: big weights, concurrency of one or two, so blocks size to the remaining budget and an over-long context is clamped.
  • clustered, one large model sharded at low concurrency: each rank holds its shard and KV under that machine’s clustered cap, and the two machines pool their memory across the cluster.
transientReserve is the load’s insurance. It is a measured high-water fraction of the budget. It covers the load-time and first-request spikes the formula alone cannot see, not a slice of the weights. Re-measure it whenever the engine, model, or KV dtype changes.

The build assert is only a policy check

A compile-time assert stays, but demoted. It checks that the declared budgets are internally consistent and that each leaves an OS floor free for the desktop and compositor. That’s a policy sanity check, not a per-model guarantee:
The per-model guarantee is the preceding runtime admission, not this assert. You never hand-tune the ceiling, the utilization, or the per-model caps.

The honest guarantee

The live sysctl cap plus per-load admission make overload virtually impossible across model and mode changes with no rebuild. But “never” is a stronger claim, and on unified memory the GPU-wired cap is necessary, not sufficient. CPU-side tensors, memory-mapped weight pages, load-time conversion copies, graph, and workspace compilation, and cache or queue growth all pressure RAM. None of it is charged to the wired cap. Bounding wired memory alone leaves that hole open. Closing it takes seven invariants, all of them:
  1. Per-process total-memory guard. Cap each worker’s whole-process footprint, not just its GPU-wired share, and fail closed before pressure becomes paging. This is the real answer to the unified-memory hole.
  2. Measured-peak accounting. Admit on observed load-plus-first-request peaks, fault-injected, not on weights + KV formula. Treat unknown allocations as reserve, never as free capacity.
  3. No-overlap transitions. Unload and verify release before loading a replacement. Never hold old and new at once. On a mode downshift, quiesce and evict until observed memory fits the new envelope before trusting the lowered sysctl; a lowered cap does not reclaim already-wired memory.
  4. Strict caps on every cache and queue. Prefix cache, input queue, output queue, workspace cache, mmap prefetch, CPU copies: each bounded, all counted in the budget.
  5. Atomic distributed admission. In clustered mode, admit per-rank on each rank’s measured peak, not weights split evenly. The layer split, the embedding and output layers, and the communication buffers make ranks uneven. Two-phase: preflight and reserve all ranks, then commit all or roll back all. Any rank that runs out tears down every peer.
  6. Memory-pressure tripwire. Watch process and system pressure continuously; on rising pressure stop admission, cancel work, and unload before sustained paging. A backstop, not a substitute for admission.
  7. Cap-comprehensiveness test. Deliberately fault-inject GPU, CPU, mmap, load-time, prefill, and first-inference spikes while measuring resident memory, wired memory, swap, page-outs, and pressure. Pass means zero page-outs and controlled rejection before the OS reaches pressure.
Invariant 7 is the load-bearing one. Until it passes on the real macOS and MLX version in use, “never overload the OS” is a target, not a proven property. The design earns the word “never” only by passing that test.
“Virtually impossible” is the honest claim today: a live mode-aware ceiling and per-load admission control catch every configuration the formula can see. “Never” is the target the fault-injection test verifies.

Where this applies

The same arithmetic holds in both serving modes. In standalone mode, one host serves itself, and the footprint is a single worker plus whatever else is running on the desktop. In clustered mode, several hosts serve a shared endpoint, and each host is derived independently. A headless desktop-class member can carry a much higher ceiling than a laptop that also runs an interactive session. That’s because the laptop must leave room for the compositor and the user’s own work. Host-specific values live in the gated operational reference, alongside the rest of the per-host numbers. That includes real ceilings, measured footprints, and the utilization figures derived from them.

Nix serving stack

How the Nix modules ship the mode-aware budgets, the load-time admission wrapper, and the link watcher — then launch the worker.

Apple Silicon stack

The serving stack these knobs belong to, and the rest of the tuning playbook.

Models and quantization

What sets the footprint side of the invariant in the first place.

Distributed serving

How clustered mode spreads work across more than one host.

nix-ai

The module that generates the serving config these flags land in.