← back to catalog · registered 2026-09-29 21:58

soppyleon/Qwen3.8-Flash-Next-ABLITERATED-MTP-NVFP4-h98k

soppyleon second-order
curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/soppyleon%2FQwen3.8-Flash-Next-ABLITERATED-MTP-NVFP4-h98k"
Response includes
  • classification unknown
  • files 18
  • author_summary 3 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
?
Primary method

Unclassified

No clear signals of an abliteration technique in this model.
Confidence
UNKNOWN
Why this label 1 signal
No classification signals present. This may not be an abliterated model at all - it could be a repackaging, a merge with unrelated goals, or unrelated content that mentions the term.
  • no classification signals present (no abliterated, uncensored, or known producer/method markers)
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-09-29

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
other
Tags
vllm safetensors qwen4_exp speculative-decoding mtp nvfp4 base_model:dealignai/Qwen3.8-Flash-Next-ABLITERATED-NVFP4 base_model:quantized:dealignai/Qwen3.8-Flash-Next-ABLITERATED-NVFP4 license:other 8-bit modelopt region:us

Related

Total size
1.72 GB
Files
18
Quantizations
1
Registered
2026-09-29 21:58
Last updated on HF
2026-09-29 21:21

Files by quantization

Auxiliary files 18 files 1.74 GB
model.safetensors 1.72 GB fbb5a217 download
tokenizer.json 12.2 MB 0997f410 download
vocab.json 6.41 MB 0aa0ce06 download
merges.txt 3.20 MB a494e019 download
model.safetensors.index.json 457 KB 80481387 download
tokenizer_config.json 17.5 KB 5de744b3 download
mtp_quantize.py 11.6 KB 3f0fc26a download
chat_template.jinja 8.74 KB c0c686f9 download
README.md 4.18 KB d3419ffc download
config.json 4.07 KB 7aa10827 download
qwen4-mtp-draft-head.patch 3.45 KB 8f779456 download
LICENSE 3.16 KB 9557a896 download
test_draft_head.py 2.19 KB 6d46dc04 download
.gitattributes 1.53 KB 52373fe2 download
hf_quant_config.json 423 B cdfa8ad8 download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
generation_config.json 202 B 023756cf download

README current version from Hugging Face


license: other
license_name: qwen-community-license-1.0
license_link: LICENSE
base_model: dealignai/Qwen3.8-Flash-Next-ABLITERATED-NVFP4
library_name: vllm
tags:

  • speculative-decoding
  • mtp
  • nvfp4
  • qwen4_exp

Qwen3.8-Flash-Next-ABLITERATED — MTP drafter, NVFP4 experts + 98K FP8 head

A drop-in speculative-decoding drafter (the model's own MTP layer) for
dealignai/Qwen3.8-Flash-Next-ABLITERATED-NVFP4
on vLLM. It is not a standalone model: it only contains mtp.* and needs that checkpoint as the target.

Compared with the stock BF16 MTP layer shipped inside the checkpoint:

stock MTP (BF16) this drafter
routed experts BF16, 4.69 GiB NVFP4 W4A4 g16 (ModelOpt layout), ~1.35 GiB
draft head shares the target's BF16 lm_head (248,320 × 2560, 1.27 GB read per draft step) dedicated FP8 head over the 98,304 most frequent tokens (0.25 GB)
attention, router, shared expert, hyper-connections, norms BF16 BF16 (unchanged)
file size — 1.72 GiB

The target model still verifies every token against its full vocabulary, so drafting can only change speed,
never which tokens are accepted.

Measured (1× RTX PRO 6000 Blackwell 96 GB, vLLM nightly af7f9488, FP8 KV, YaRN ×2 → 512K context, --gpu-memory-utilization 0.97, k = 3)

Mean decode tok/s over code / tool-call / reasoning / prose prompts (8 each, greedy); 1 request = per-stream,
2–8 = aggregate.

1 req 2 req 4 req 8 req KV cache pool
MTP off 80 138 223 335 1,258,883 tokens
stock MTP, k=3 124 224 339 520 688,659
this drafter, k=3 145 249 377 602 904,042
  • +17 % over the stock drafter at 1 request, +16 % at 8, and +215K KV tokens (the NVFP4 experts free ~3.3 GiB).
  • Per-position acceptance is unchanged vs stock (code 0.98 / 0.89 / 0.73, tool calls 0.96 / 0.95 / 0.89).
  • 98K-token head covers 99.8–99.98 % of held-out chat, reasoning, code, tool-call and Croatian text.
  • Greedy divergence vs MTP-off is at the same rate as MTP-off vs itself (this stack is not bit-reproducible
    across server starts); quality battery: see below.

Quality (n = 250, thinking off, same items as the MTP-off baseline): HumanEval 93/100, debug 9/10,
GSM8K 98/100, tool calls 10/10, IFEval-lite 9/10, refusals 0/20 — total 239/250 vs 239/250 for MTP off
(2 items flipped each way, exact McNemar p = 1.00).

Usage

  1. Apply the draft-head patch to vLLM (the dedicated reduced-vocab head is not upstream):

    SITE=$(python3 -c 'import vllm, os; print(os.path.dirname(os.path.dirname(vllm.__file__)))')
    patch -p1 -d "$SITE" < qwen4-mtp-draft-head.patch
    python3 test_draft_head.py     # unit test: scatter + wiring
    

    The patch is against vLLM af7f9488 (vllm/models/qwen4_exp/nvidia/mtp.py, ~80 lines) and is inert unless the
    draft config sets mtp_draft_vocab_size.

  2. Serve the target with this repo as the drafter:

    hf download soppyleon/Qwen3.8-Flash-Next-ABLITERATED-MTP-NVFP4-h98k --local-dir /models/mtp-h98k
    vllm serve /models/dealignai-Qwen3.8-Flash-Next-ABLITERATED-NVFP4 \
      --speculative-config '{"method":"mtp","num_speculative_tokens":3,"model":"/models/mtp-h98k"}' \
      ...your usual flags
    

How it was built

mtp_quantize.py (included) from the mtp.* tensors of dealignai@be794b99:

  • experts: per-expert NVFP4 (global scale = amax / (6·448), E4M3 block scales per 16, e2m1 round-to-nearest-even),
    layout verified bit-exact against RadixArk's own NVFP4 experts; input_scale = 2 × the target's last-layer max
    (no calibration pass);
  • head: rows of the target lm_head for the 98,304 most frequent tokens of an output-side corpus (UltraChat
    assistant turns, GSM8K/MATH solutions, open-source Python/Rust/TS code, SWE-bench patches, synthetic
    qwen3_xml tool calls, Croatian Wikipedia) plus all special and single-character tokens, stored as E4M3 with a
    per-tensor scale.

License

Derived from Qwen3.8-Flash-Next via dealignai's abliterated NVFP4 checkpoint; distributed under the
Qwen Community License (see LICENSE). Built with Qwen.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.