Skip to main content
Decode reads the whole weight file to produce one token. Speculative decoding is the one trick that gets you more than one token per read.

Why decode is slow, and the only two ways out

Generating a token means streaming every active weight through the GPU. So single-stream decode is memory-bandwidth bound, not compute bound, and the ceiling is close to:
That leaves exactly two levers: Quantization is the familiar one and hits diminishing returns fast. The second is where multi-token prediction lives.

What MTP is, and how it differs from a drafter

Ordinary speculative decoding needs two models: a small draft model proposes tokens, the big target model verifies them in a single batched pass. Tokens the target agrees with are kept; the first disagreement discards the rest. You win when drafts are accepted often enough to beat the cost of proposing them. The weak point is the pairing. A separate small model was trained on different data with a different tokenizer distribution, so its guesses diverge from the target’s and acceptance falls. MTP removes the pairing problem. Qwen 3.8 was trained with a multi-token prediction head. This is a small module trained alongside the model, on the same data, to predict the next several positions. The draft comes from the target’s own learned distribution, so acceptance is structurally higher. On MLX, that head ships as a separate Hub artifact. It uses the standalone drafter format the server expects, and loads as the draft model beside the target:
  • lukaskremla/Qwen3.8-27B-MTP-6bit-MLX: the extracted MTP tensors
  • mlx-community/Qwen3.8-27B-MTP-{4bit,8bit,mxfp4,bf16}: MTP-bearing targets
The drafter is quantized independently of the target. A 6-bit drafter works against 2-, 3-, 4-, 5-, 6- and 8-bit targets. So adopting MTP does not mean re-downloading your weights. The target you already serve stays put, and you add the drafter beside it.
This independence matters more than it looks. Quantizing a draft head as hard as the body is a classic way to lose the speedup without noticing. The body still answers correctly, but the drafts stop being accepted, and the win quietly evaporates. Keeping the drafter at higher precision than the target is cheap, since it is a small fraction of the weights, and it protects acceptance.

Draft depth is the tunable

Depth (how many tokens the head proposes per step) is the knob worth sweeping. Deeper drafts win more per accepted run and waste more per rejection, so the best value depends on how predictable your text is:
  • Code accepts deeper drafts. Syntax is constrained, so long runs survive.
  • Prose accepts shallower ones. More branch points, earlier rejection.
Sweep it per workload rather than adopting a number from someone else’s report. Acceptance depends on the model, the quant pairing, and the text.

Concurrency is a coupling, not a cap

The most common misreading. In the Nix serving stack two values must be equal:
  • the worker’s batch width (maxNumSeqs on the MTP profile), and
  • the proxy-side admission gate (modelConcurrencyLimits, else proxy.concurrencyLimit).
Both accept 1–4. Nothing forces MTP to concurrency 1. A matched pair at 2, 3, or 4 is equally valid. The equality is enforced because the split is what actually broke. A proxy admitting 4 requests while the server served 1 turned the excess into 429s. The rule is that both move together, not that either is pinned.
Because proxy.concurrencyLimit defaults to 1, an MTP profile at maxNumSeqs = 1 looks like proof that MTP must be serialized. It isn’t. Read the assertion as “these two must agree,” and raise both together when you want concurrency.

Don’t infer configuration from a throughput ratio

A tempting shortcut: measure with the feature off, measure with it on, divide, and read the draft depth off the multiplier. It doesn’t work at this precision. On one estate’s own hardware, the same model serving byte-identical output measured anywhere from 17.4 to 27.3 tok/s across runs. A ratio built from two single runs sits well inside that spread, so it cannot identify a parameter. The noise is larger than the effect being attributed. Two runs and a divide is not a measurement. Before recording any number:
  1. Prove which weights the process actually loaded. See Verifying the instrument.
  2. Repeat each configuration enough times to see the spread, and report the spread, not a single figure.
  3. Fix the prompt set, and separate code from prose. They have different acceptance rates, and averaging them hides both.
  4. Restart the server between configurations. Most of these settings are read at startup only, so an in-place edit measures the old config and looks like a null result.
Step 4 is the quiet one. A config file edited without a restart is the most common way an A/B test produces a confident, wrong answer.

Evaluating a change here

1

Confirm the option is reachable

Enable the profile and evaluate the configuration. A feature can be fully wired and still be unreachable if a policy assertion elsewhere forbids the backend it needs.
2

Serve it on an isolated worker

Never measure on the box handling live traffic. Reconcile the engine batch width with the profile’s allowed range first.
3

Prove the weights, then sweep

Read the loaded model off the worker’s own process command line, then sweep draft depth across the range, prose and code separately.
4

Decide on the spread

Adopt only if the win clears the run-to-run noise floor by a margin you’d defend. Otherwise record what you measured and leave the default.