Decode reads the whole weight file to produce one token. Speculative decoding is the one trick that gets you more than one token per read.
Why decode is slow, and the only two ways out
Generating a token means streaming every active weight through the GPU. So single-stream decode is memory-bandwidth bound, not compute bound, and the ceiling is close to:
Quantization is the familiar one and hits diminishing returns fast. The second
is where multi-token prediction lives.
What MTP is, and how it differs from a drafter
Ordinary speculative decoding needs two models: a small draft model proposes tokens, the big target model verifies them in a single batched pass. Tokens the target agrees with are kept; the first disagreement discards the rest. You win when drafts are accepted often enough to beat the cost of proposing them. The weak point is the pairing. A separate small model was trained on different data with a different tokenizer distribution, so its guesses diverge from the target’s and acceptance falls. MTP removes the pairing problem. Qwen 3.8 was trained with a multi-token prediction head. This is a small module trained alongside the model, on the same data, to predict the next several positions. The draft comes from the target’s own learned distribution, so acceptance is structurally higher. On MLX, that head ships as a separate Hub artifact. It uses the standalone drafter format the server expects, and loads as the draft model beside the target:lukaskremla/Qwen3.8-27B-MTP-6bit-MLX: the extracted MTP tensorsmlx-community/Qwen3.8-27B-MTP-{4bit,8bit,mxfp4,bf16}: MTP-bearing targets
The drafter is quantized independently of the target. A 6-bit drafter
works against 2-, 3-, 4-, 5-, 6- and 8-bit targets. So adopting MTP does not
mean re-downloading your weights. The target you already serve stays put,
and you add the drafter beside it.
Draft depth is the tunable
Depth (how many tokens the head proposes per step) is the knob worth sweeping. Deeper drafts win more per accepted run and waste more per rejection, so the best value depends on how predictable your text is:- Code accepts deeper drafts. Syntax is constrained, so long runs survive.
- Prose accepts shallower ones. More branch points, earlier rejection.
Concurrency is a coupling, not a cap
The most common misreading. In the Nix serving stack two values must be equal:- the worker’s batch width (
maxNumSeqson the MTP profile), and - the proxy-side admission gate (
modelConcurrencyLimits, elseproxy.concurrencyLimit).
429s.
The rule is that both move together, not that either is pinned.
Don’t infer configuration from a throughput ratio
A tempting shortcut: measure with the feature off, measure with it on, divide, and read the draft depth off the multiplier. It doesn’t work at this precision. On one estate’s own hardware, the same model serving byte-identical output measured anywhere from 17.4 to 27.3 tok/s across runs. A ratio built from two single runs sits well inside that spread, so it cannot identify a parameter. The noise is larger than the effect being attributed. Two runs and a divide is not a measurement. Before recording any number:- Prove which weights the process actually loaded. See Verifying the instrument.
- Repeat each configuration enough times to see the spread, and report the spread, not a single figure.
- Fix the prompt set, and separate code from prose. They have different acceptance rates, and averaging them hides both.
- Restart the server between configurations. Most of these settings are read at startup only, so an in-place edit measures the old config and looks like a null result.
Evaluating a change here
1
Confirm the option is reachable
Enable the profile and evaluate the configuration. A feature can be fully
wired and still be unreachable if a policy assertion elsewhere forbids the
backend it needs.
2
Serve it on an isolated worker
Never measure on the box handling live traffic. Reconcile the engine batch
width with the profile’s allowed range first.
3
Prove the weights, then sweep
Read the loaded model off the worker’s own process command line, then sweep
draft depth across the range, prose and code separately.
4
Decide on the spread
Adopt only if the win clears the run-to-run noise floor by a margin you’d
defend. Otherwise record what you measured and leave the default.
Related
- Verifying the instrument: proving which weights a server actually loaded
- Models and quantization: the bit-depth axis
- Prompt cache and concurrency: the admission gate this coupling touches
- Backends and tool calling: why tok/s is the wrong headline metric