Two memory behaviours ride on one flag, over a wired ceiling that moves with the serving mode. The stack tunes them at runtime. A model or mode swap needs no rebuild. But the GPU cap alone is necessary, not sufficient.A local MLX serving stack is a swapping proxy in front of one or more inference workers. It has a single flag,
--gpu-memory-utilization, that
drives two different memory behaviours at once. Reading it as one knob is the
usual reason a well-sized host still ends up paging.
The two behaviours
The flag is a fraction. It feeds two independent calculations:allocation_limitis the per-worker Metal allocation cap. A worker that tries to allocate past it fails the allocation.tripis the engine’s emergency KV-cache clear. When resident memory crosses it, the engine drops the cache to recover headroom.
Raising the wired ceiling raises
allocation_limit but leaves trip exactly
where it was.
Insight one: it is a cap, not a reservation
--gpu-memory-utilization does not pre-allocate anything. Setting it to a
large fraction does not hand that memory to the worker. Setting it small does
not hold memory back for the desktop either. It only says how far an
allocation is allowed to go before it is refused.
The knob that actually decides how much of a host goes to inference is the OS
wired ceiling. Treat --gpu-memory-utilization as a safety rail on top of that
decision, not as the decision itself.
Insight two: the two knobs move together
The invariant to hold is:Lowering utilization alone
The trip point drops below the resident working set. The engine clears the
KV cache constantly, every turn re-prefills from scratch, and multi-turn
sessions degrade badly while the GPU looks busy.
Raising the ceiling alone
The trip point stays put while the hard ceiling climbs past it. Or the
trip point already lands past the ceiling entirely. Either way, the soft
valve never fires, and the host reaches its hard limit first.
Neither knob is safe to change by itself. Any change to the wired ceiling
needs a matching re-derivation of
--gpu-memory-utilization, and the reverse.Re-deriving the utilization
Rearranging the invariant gives a symbolic band for the flag:0.05 offset subtracted out. Pick a value inside the band with margin on
both sides, and re-derive it whenever the ceiling, the model, or the batch
shape changes.
If the band is empty (the lower bound exceeds the upper), no value of the flag
satisfies the invariant on that host. Raise the ceiling, shrink the footprint
(smaller cache reservation, fewer concurrent sequences, a smaller model), or
both.
Measuring the footprint: use wired, not RSS
On macOS,ps RSS does not attribute kernel-wired Metal allocations to the
process that caused them. A worker holding tens of GB of wired GPU memory can
look small in ps while top shows the real figure. Size against the wired
number; RSS quietly tells you the host has room it does not have.
Wired is the right number for the GPU footprint, but it is not the whole
process. On unified memory a worker also holds CPU-side tensors, memory-mapped
weight pages, and load-time copies that the wired cap does not charge. Bounding
wired memory is necessary but not sufficient. See the honest
guarantee below, which bounds the whole process
instead.
Finding the real ceiling
The default wired ceiling on Apple Silicon is commonly quoted as about 75% of RAM. In practice it can be a substantially larger fraction than that. The rule of thumb is a guess; measure the machine instead.1
Read the ceiling from MLX
Query
max_recommended_working_set_size through MLX. That is the value the
allocator actually enforces, whatever the sysctl is currently set to. Read
it before changing anything. The default may already be higher than you
were about to set.2
Set the sysctl, then re-read
Change
iogpu.wired_limit_mb and read max_recommended_working_set_size
again to confirm the new ceiling took effect.3
Re-derive utilization
Feed the confirmed ceiling back through the preceding band and set
--gpu-memory-utilization to a value inside it.4
Confirm the engine actually applied it
Read back the limit the engine reports at startup, not just the flag you
passed. Serving stacks build their engine down more than one code path. A
flag honored on one path can be silently ignored on another, leaving the
default in force while your command line says otherwise. Trust the engine’s
own log line, then check the resulting trip point still satisfies the
invariant.
The unit is MiB, and it is exact
iogpu.wired_limit_mb is in MiB, and MLX reports the ceiling as:
util × ceiling leaves
more allocation headroom than intended. Meanwhile trip is computed against
total_ram, untouched by the sysctl, so it stays exactly where it was. Always
re-derive from the value MLX reports back, never from the number you typed
into the sysctl.
Weights carry the mirror-image trap. Hugging Face reports a model’s size in
decimal GB (×1e9 bytes); every budget in the admission formula is in
GiB (×1024³). Convert weightsBytes with ×1e9, not ×1024³. Conflate
the two and the subtracted weight inflates by 7.37%, silently shrinking the KV
pool you thought you had.
Bounding the footprint: which flag is the hard cap
The invariant needs a real bound onfootprint. Two flags look like they set
it; only one does.
--max-cache-blocksis the hard cap. It bounds the active paged-KV pool, the memory that grows as concurrent requests and longer contexts arrive. When the pool runs out, the worker rejects the request instead of allocating past the bound. This is the flag that keepsfootprintundertrip.--cache-memory-mbbounds only the prefix cache. The prefix cache is a separate store that reuses shared prompt prefixes across requests. Sizing against it leaves the paged-KV pool unbounded, and that pool is the part that actually grows under load.
--max-cache-blocks. A host tuned on
--cache-memory-mb alone can still cross trip under concurrency, because the
pool that grew was never the one it capped.
A context limit includes the answer
The context number exposed by a serving profile is a total admission limit. It is not an input-only promise. Before a worker accepts a request, it must reserve room for both the prompt and the allowed completion:- Model-native maximum: what the checkpoint architecture can support.
- Configured total context: what this serving profile is willing to admit.
- Actual prompt tokens: what the server received for one request.
- Output reservation: the completion room protected at admission.
- Actual total context: what the request used after completion.
Safety is runtime, not a build assert
Holdingfootprint < trip < ceiling by hand means two coupled knobs move
together: the wired ceiling and --gpu-memory-utilization. Both must be
re-derived every time the model, the mode, or the batch shape changes. A build
could derive them once and assert
the result, but that only covers the config it was built with. Models and modes
change while the host runs, and rebuilding to swap a model is exactly the
friction the stack avoids. So the safety lives at runtime, at two enforcement
points. Neither needs a darwin-rebuild. Swapping a model or switching mode is
as cheap as changing a model in a proxy router.
One budget input per host, per mode
A host declares one number,maxLocalLlmGb: the total memory the local LLM
stack may use. It carries two values: one for standalone serving, and a larger
one for clustered mode. In clustered mode the desktop is quiesced, so more
memory is free for a shard. The operator sets the numbers; the running mode
picks which one applies.
Enforcement point one: the live mode-aware sysctl cap
The OS wired ceiling,iogpu.wired_limit_mb, is a live sysctl. A link watcher
flips it to the current mode’s budget the moment the cluster cable is plugged
or unplugged. Plugging in applies the clustered budget; unplugging applies the
standalone budget. No reboot, no rebuild. In clustered mode both Macs raise the
same higher cap.
This is the hard kernel backstop on GPU-wired memory, and it moves with the
mode. It stays high, set to the mode’s budget, and is never lowered as a quick
fix.
Enforcement point two: per-model load-time admission control
When the swap proxy spawns a worker, a wrapper computes the safe caps from live inputs. Those inputs are the current sysctl value, the model’s architecture, and the requested concurrency and context. Swapping a model needs no nix change:--max-cache-blocks is that global pool count: the whole paged-KV pool, shared
across every concurrent sequence, not a per-request budget. So the pool size
does not divide by concurrency. Concurrency enters at admission instead:
--max-cache-blocks to poolBlocks and derives
--gpu-memory-utilization from the preceding live ceiling band. If the
requested concurrency and context do not fit the pool, the wrapper rejects
or clamps the load and fails loud. It shrinks concurrency or context until
the admission check holds. A model too big to fit even one sequence is
refused, not paged to disk.
kvLayers is the count of layers that actually hold a KV cache, not the total
layer count. A hybrid-attention model charges KV only for its full-attention
layers. qwen3-next-80b runs 12 full-attention layers of 48, so its KV costs
about 24 KiB/token. That’s cheaper than a dense 30 B at about 96 KiB/token,
despite the far larger weights. Read kvLayers from the model’s
architecture; never assume it equals the depth.- a tiny model at high concurrency: small weights leave a large KV budget, so many sequences fit.
- one large model with a long context: big weights, concurrency of one or two, so blocks size to the remaining budget and an over-long context is clamped.
- clustered, one large model sharded at low concurrency: each rank holds its shard and KV under that machine’s clustered cap, and the two machines pool their memory across the cluster.
transientReserve is the load’s insurance. It is a measured high-water
fraction of the budget. It covers the load-time and first-request spikes the
formula alone cannot see, not a slice of the weights. Re-measure it whenever
the engine, model, or KV dtype changes.
The build assert is only a policy check
A compile-time assert stays, but demoted. It checks that the declared budgets are internally consistent and that each leaves an OS floor free for the desktop and compositor. That’s a policy sanity check, not a per-model guarantee:The honest guarantee
The live sysctl cap plus per-load admission make overload virtually impossible across model and mode changes with no rebuild. But “never” is a stronger claim, and on unified memory the GPU-wired cap is necessary, not sufficient. CPU-side tensors, memory-mapped weight pages, load-time conversion copies, graph, and workspace compilation, and cache or queue growth all pressure RAM. None of it is charged to the wired cap. Bounding wired memory alone leaves that hole open. Closing it takes seven invariants, all of them:- Per-process total-memory guard. Cap each worker’s whole-process footprint, not just its GPU-wired share, and fail closed before pressure becomes paging. This is the real answer to the unified-memory hole.
- Measured-peak accounting. Admit on observed load-plus-first-request
peaks, fault-injected, not on
weights + KV formula. Treat unknown allocations as reserve, never as free capacity. - No-overlap transitions. Unload and verify release before loading a replacement. Never hold old and new at once. On a mode downshift, quiesce and evict until observed memory fits the new envelope before trusting the lowered sysctl; a lowered cap does not reclaim already-wired memory.
- Strict caps on every cache and queue. Prefix cache, input queue, output queue, workspace cache, mmap prefetch, CPU copies: each bounded, all counted in the budget.
- Atomic distributed admission. In clustered mode, admit per-rank on each rank’s measured peak, not weights split evenly. The layer split, the embedding and output layers, and the communication buffers make ranks uneven. Two-phase: preflight and reserve all ranks, then commit all or roll back all. Any rank that runs out tears down every peer.
- Memory-pressure tripwire. Watch process and system pressure continuously; on rising pressure stop admission, cancel work, and unload before sustained paging. A backstop, not a substitute for admission.
- Cap-comprehensiveness test. Deliberately fault-inject GPU, CPU, mmap, load-time, prefill, and first-inference spikes while measuring resident memory, wired memory, swap, page-outs, and pressure. Pass means zero page-outs and controlled rejection before the OS reaches pressure.
“Virtually impossible” is the honest claim today: a live mode-aware ceiling
and per-load admission control catch every configuration the formula can see.
“Never” is the target the fault-injection test verifies.
Where this applies
The same arithmetic holds in both serving modes. In standalone mode, one host serves itself, and the footprint is a single worker plus whatever else is running on the desktop. In clustered mode, several hosts serve a shared endpoint, and each host is derived independently. A headless desktop-class member can carry a much higher ceiling than a laptop that also runs an interactive session. That’s because the laptop must leave room for the compositor and the user’s own work. Host-specific values live in the gated operational reference, alongside the rest of the per-host numbers. That includes real ceilings, measured footprints, and the utilization figures derived from them.Related
Nix serving stack
How the Nix modules ship the mode-aware budgets, the load-time admission wrapper, and the link watcher — then launch the worker.
Apple Silicon stack
The serving stack these knobs belong to, and the rest of the tuning playbook.
Models and quantization
What sets the footprint side of the invariant in the first place.
Distributed serving
How clustered mode spreads work across more than one host.
nix-ai
The module that generates the serving config these flags land in.