license: other
license_name: qwen-community-license-1.0
license_link: LICENSE
base_model: dealignai/Qwen3.8-Flash-Next-ABLITERATED-NVFP4
library_name: vllm
tags:
- speculative-decoding
- mtp
- nvfp4
- qwen4_exp
Qwen3.8-Flash-Next-ABLITERATED — MTP drafter, NVFP4 experts + 98K FP8 head
A drop-in speculative-decoding drafter (the model's own MTP layer) fordealignai/Qwen3.8-Flash-Next-ABLITERATED-NVFP4
on vLLM. It is not a standalone model: it only contains mtp.* and needs that checkpoint as the target.
Compared with the stock BF16 MTP layer shipped inside the checkpoint:
| stock MTP (BF16) | this drafter | |
|---|---|---|
| routed experts | BF16, 4.69 GiB | NVFP4 W4A4 g16 (ModelOpt layout), ~1.35 GiB |
| draft head | shares the target's BF16 lm_head (248,320 × 2560, 1.27 GB read per draft step) |
dedicated FP8 head over the 98,304 most frequent tokens (0.25 GB) |
| attention, router, shared expert, hyper-connections, norms | BF16 | BF16 (unchanged) |
| file size | — | 1.72 GiB |
The target model still verifies every token against its full vocabulary, so drafting can only change speed,
never which tokens are accepted.
Measured (1× RTX PRO 6000 Blackwell 96 GB, vLLM nightly af7f9488, FP8 KV, YaRN ×2 → 512K context, --gpu-memory-utilization 0.97, k = 3)
Mean decode tok/s over code / tool-call / reasoning / prose prompts (8 each, greedy); 1 request = per-stream,
2–8 = aggregate.
| 1 req | 2 req | 4 req | 8 req | KV cache pool | |
|---|---|---|---|---|---|
| MTP off | 80 | 138 | 223 | 335 | 1,258,883 tokens |
| stock MTP, k=3 | 124 | 224 | 339 | 520 | 688,659 |
| this drafter, k=3 | 145 | 249 | 377 | 602 | 904,042 |
- +17 % over the stock drafter at 1 request, +16 % at 8, and +215K KV tokens (the NVFP4 experts free ~3.3 GiB).
- Per-position acceptance is unchanged vs stock (code 0.98 / 0.89 / 0.73, tool calls 0.96 / 0.95 / 0.89).
- 98K-token head covers 99.8–99.98 % of held-out chat, reasoning, code, tool-call and Croatian text.
- Greedy divergence vs MTP-off is at the same rate as MTP-off vs itself (this stack is not bit-reproducible
across server starts); quality battery: see below.
Quality (n = 250, thinking off, same items as the MTP-off baseline): HumanEval 93/100, debug 9/10,
GSM8K 98/100, tool calls 10/10, IFEval-lite 9/10, refusals 0/20 — total 239/250 vs 239/250 for MTP off
(2 items flipped each way, exact McNemar p = 1.00).
Usage
Apply the draft-head patch to vLLM (the dedicated reduced-vocab head is not upstream):
SITE=$(python3 -c 'import vllm, os; print(os.path.dirname(os.path.dirname(vllm.__file__)))') patch -p1 -d "$SITE" < qwen4-mtp-draft-head.patch python3 test_draft_head.py # unit test: scatter + wiringThe patch is against vLLM
af7f9488(vllm/models/qwen4_exp/nvidia/mtp.py, ~80 lines) and is inert unless the
draft config setsmtp_draft_vocab_size.Serve the target with this repo as the drafter:
hf download soppyleon/Qwen3.8-Flash-Next-ABLITERATED-MTP-NVFP4-h98k --local-dir /models/mtp-h98k vllm serve /models/dealignai-Qwen3.8-Flash-Next-ABLITERATED-NVFP4 \ --speculative-config '{"method":"mtp","num_speculative_tokens":3,"model":"/models/mtp-h98k"}' \ ...your usual flags
How it was built
mtp_quantize.py (included) from the mtp.* tensors of dealignai@be794b99:
- experts: per-expert NVFP4 (global scale = amax / (6·448), E4M3 block scales per 16, e2m1 round-to-nearest-even),
layout verified bit-exact against RadixArk's own NVFP4 experts;input_scale= 2 × the target's last-layer max
(no calibration pass); - head: rows of the target
lm_headfor the 98,304 most frequent tokens of an output-side corpus (UltraChat
assistant turns, GSM8K/MATH solutions, open-source Python/Rust/TS code, SWE-bench patches, syntheticqwen3_xmltool calls, Croatian Wikipedia) plus all special and single-character tokens, stored as E4M3 with a
per-tensor scale.
License
Derived from Qwen3.8-Flash-Next via dealignai's abliterated NVFP4 checkpoint; distributed under the
Qwen Community License (see LICENSE). Built with Qwen.