← back to catalog · registered 2026-08-22 13:56

THe-Plague/Qwen3.6-35B-A3B-abliterated-NVFP4-MTP

THe-Plague Qwen 17B MoE
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/THe-Plague%2FQwen3.6-35B-A3B-abliterated-NVFP4-MTP"
Response includes
  • classification m1
  • files 12
  • hub_downloads_all_time 33,925
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
34K
2K last 30d - cooling
Likes
10
Model age
3mo ago
created 2026-06-28
Downloads over time
Now34.4K→from344↑9,901%
012.6K25.2K37.8K344 on Jul 134.4K on Oct 11JulAugSepOct
Jul 1 → Oct 11 · 54 snapshots · spans 102 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh
Tags
vllm safetensors qwen3_5_moe qwen3 moe abliterated uncensored nvfp4 mtp speculative-decoding blackwell dgx-spark
Total size
22.0 GB
Files
12
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-06-28 20:00

Files by quantization

Auxiliary files 12 files 22.0 GB
model.safetensors 20.4 GB ce41feb8 download
model-mtp.safetensors 1.57 GB 68b2cabe download
tokenizer.json 19.1 MB 6a0316e3 download
model.safetensors.index.json 11.9 MB ff3c44b4 download
config.json 13.6 KB 149dffa9 download
chat_template.jinja 10.7 KB 74741eb8 download
README.md 3.59 KB 1462fca9 download
.gitattributes 1.60 KB aa7aacd0 download
processor_config.json 1.16 KB 33818c7f download
tokenizer_config.json 1.07 KB e15d4cc3 download
recipe.yaml 237 B 2ed68f09 download
generation_config.json 213 B 23a0a961 download

README current version from Hugging Face


license: apache-2.0
language:

  • en
  • zh
    base_model:
  • Qwen/Qwen3-35B-A3B
  • huihui-ai/Huihui-Qwen3.6-35B-A3B-abliterated
  • sakamakismile/Huihui-Qwen3.6-35B-A3B-abliterated-NVFP4
    library_name: vllm
    pipeline_tag: text-generation
    tags:
  • qwen3
  • moe
  • abliterated
  • uncensored
  • nvfp4
  • mtp
  • speculative-decoding
  • blackwell
  • dgx-spark
  • gb10

Qwen3.6-35B-A3B-abliterated-NVFP4-MTP

The fastest abliterated Qwen3.6 for NVIDIA Blackwell / DGX Spark — NVFP4 quantized with MTP (Multi-Token Prediction) layers for speculative decoding.

95.2 tok/s decode on a single DGX Spark (GB10) with MTP=3.

What Makes This Different

No other model on HuggingFace combines all three:

Feature This Model sakamakismile NVFP4 RedHatAI NVFP4
NVFP4 weights ✅ ✅ ✅
MTP layers ✅ BF16 ❌ ✅
Abliterated ✅ ✅ ❌
Speculative decode ✅ 95 tok/s ❌ ~70 tok/s ✅ but censored

Without MTP layers, you're limited to ~70 tok/s on GB10. MTP enables speculative decoding which gives a +35% speedup.

Model Details

  • Architecture: Qwen3 MoE (Mixture of Experts), 35B total / 3B active parameters
  • Main weights: NVFP4 quantized (21.9 GB) — optimized for NVIDIA Blackwell CUTLASS kernels
  • MTP layers: BF16 (1.6 GB, 19 tensors in model-mtp.safetensors)
  • Total size: ~23.5 GB
  • Context length: 131,072 tokens
  • Abliteration: Refusal direction removed via representation engineering (from huihui-ai)

How It Was Made

  1. Started with sakamakismile/Huihui-Qwen3.6-35B-A3B-abliterated-NVFP4 (NVFP4 main weights, no MTP)
  2. Extracted BF16 MTP layers from huihui-ai/Huihui-Qwen3.6-35B-A3B-abliterated (full BF16 model)
  3. Injected MTP tensors into a separate model-mtp.safetensors file
  4. Updated model.safetensors.index.json with MTP weight mappings
  5. Verified with vLLM speculative decoding (MTP=1 through MTP=4)

Benchmark Results

Hardware: NVIDIA DGX Spark (GB10 Blackwell, sm_121, ~128GB unified memory)
Engine: vLLM v0.23.1 with FlashInfer + CUTLASS NVFP4 MoE kernels
Benchmark: 256 generated tokens at various context lengths

Config 1K ctx 4K ctx 16K ctx 32K ctx 64K ctx
MTP=3, seqs=1 (best) 92.2 95.2 90.7 89.7 84.4
MTP=2, seqs=4 87.2 75.5 82.8 81.6 74.4
MTP=4, seqs=1 75.6 71.3 74.0 77.6 74.4
No MTP (NVFP4 only) ~71 ~70 ~74 ~68 ~70

MTP=3 is the sweet spot. MTP=4 regresses due to speculation overhead exceeding acceptance rate.

Recommended vLLM Launch Config

docker run -d --gpus all --network host --ipc host --shm-size=16g \
  --name vllm-qwen36 \
  -e VLLM_MARLIN_USE_ATOMIC_ADD=1 \
  -e TORCH_MATMUL_PRECISION=high \
  -v /path/to/models:/models \
  ghcr.io/spark-arena/dgx-vllm-eugr-nightly:latest \
  vllm serve /models/Qwen3.6-35B-A3B-abliterated-NVFP4-MTP \
    --host 0.0.0.0 --port 8000 --max-model-len 131072 \
    --max-num-batched-tokens 32768 --max-num-seqs 1 \
    --trust-remote-code --gpu-memory-utilization 0.7 \
    --reasoning-parser qwen3 --kv-cache-dtype fp8 \
    --load-format instanttensor --attention-backend flashinfer \
    --speculative-config '{"method": "mtp", "num_speculative_tokens": 3}' \
    --enable-prefix-caching -tp 1

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-28Update README.md39cd0673.6 KB
    Loading...
  2. 2026-06-28initial commitd979eb528 B
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration