← back to catalog · registered 2026-08-22 13:56

coolthor/Huihui-Qwen3.6-35B-A3B-abliterated-NVFP4-W4A4

coolthor Qwen 17B MoE multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/coolthor%2FHuihui-Qwen3.6-35B-A3B-abliterated-NVFP4-W4A4"
Response includes
  • classification m1
  • files 27
  • hub_downloads_all_time 1,576
  • author_summary 7 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
2K
295 last 30d - stable
Likes
0
Model age
4mo ago
created 2026-06-02
Downloads over time
Now1.7K→from320↑422%
2527691.3K1.8K320 on Jun 101.7K on Oct 11JunJulAugSepOct
Jun 10 → Oct 11 · 57 snapshots · spans 123 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh
Tags
safetensors qwen3_5_moe abliterated uncensored nvfp4 w4a4 fp4 compressed-tensors llm-compressor vllm dgx-spark gb10

Related

Total size
22.0 GB
Files
27
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-06-02 09:04

Files by quantization

Auxiliary files 27 files 22.0 GB
model-00010-of-00012.safetensors 1.86 GB 9bf01bdb download
model-00005-of-00012.safetensors 1.86 GB 9df5cff9 download
model-00008-of-00012.safetensors 1.86 GB c4c3cb40 download
model-00003-of-00012.safetensors 1.86 GB 6df4c8c2 download
model-00007-of-00012.safetensors 1.86 GB f926fff8 download
model-00009-of-00012.safetensors 1.86 GB e02dc39b download
model-00006-of-00012.safetensors 1.86 GB b73c2d2b download
model-00004-of-00012.safetensors 1.86 GB 7fae6c90 download
model-00002-of-00012.safetensors 1.86 GB f010a80d download
model-00011-of-00012.safetensors 1.86 GB 76bb77c7 download
model_mtp.safetensors 1.57 GB 7d0030f5 download
model-00001-of-00012.safetensors 970 MB 6abf5b16 download
model-00012-of-00012.safetensors 832 MB 8fb51895 download
model.safetensors.index.json 13.7 MB aa4dd647 download
tokenizer.json 12.2 MB 5f9e4d49 download
vocab.json 6.41 MB 0aa0ce06 download
merges.txt 3.20 MB a494e019 download
tokenizer_config.json 16.3 KB 28d96ff3 download
config.json 13.5 KB 6470ffef download
chat_template.jinja 7.58 KB a8755d82 download
README.md 6.63 KB 7893a330 download
.gitattributes 1.60 KB a09db2ea download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
recipe.yaml 259 B ce528af5 download
generation_config.json 202 B 023756cf download
configuration.json 58.0 B d24dba94 download

README current version from Hugging Face


license: apache-2.0
base_model: huihui-ai/Huihui-Qwen3.6-35B-A3B-abliterated
base_model_relation: quantized
tags:

  • abliterated
  • uncensored
  • nvfp4
  • w4a4
  • fp4
  • compressed-tensors
  • llm-compressor
  • vllm
  • dgx-spark
  • gb10
  • sm121
  • moe
  • cudagraph
    language:
  • en
  • zh
    pipeline_tag: image-text-to-text

Huihui-Qwen3.6-35B-A3B-abliterated-NVFP4-W4A4

NVFP4 W4A4 quantization (4-bit weights and 4-bit activations) of huihui-ai/Huihui-Qwen3.6-35B-A3B-abliterated — an uncensored / abliterated Qwen3.6-35B-A3B hybrid MoE (~3B active, GDN linear-attention + full-attention interleave). Self-quantized with llm-compressor (NVFP4 scheme) in compressed-tensors nvfp4-pack-quantized format.

The point of this checkpoint: on a single NVIDIA DGX Spark (GB10, SM121) it is servable in vLLM and decodes faster than the FP8 build — which most NVFP4 MoE quants on consumer/SM121 Blackwell are not (they typically fall back to a dequant slow path or refuse to compile the block-scaled MMA). With the right toolchain it lands at 66.9 tok/s vs 52.0 tok/s for FP8 (+29%) while using 16 GB less weight memory.

Lineage: BF16 abliterated base (huihui-ai) → NVFP4 W4A4 (this repo). Vision tower preserved and the model.language_model.visual.* → visual.* loader-prefix fix already applied (see note below). MTP heads preserved but off by default (see below).

Quick stats

Metric Value
Base huihui-ai/Huihui-Qwen3.6-35B-A3B-abliterated (BF16)
Format compressed-tensors nvfp4-pack-quantized
Scheme NVFP4 — W4A4 (weights FP4 + activations FP4), group size 16, FP8 e4m3fn scales
KV cache scheme fp8
Disk size ~22 GB
Kept higher precision lm_head, visual.*, mlp.gate (router), shared_expert_gate
MTP weights Preserved (model_mtp.safetensors) — keep OFF, see note
Native context 256K (max_position_embeddings 262144), no YaRN
Modality Text + Image (vision tower preserved, loader-prefix fixed)

Performance on DGX Spark (GB10, vLLM)

Paired, same-harness, single-stream decode, KV-cache fp8, warm:

Configuration Weight mem tok/s vs FP8
FP8 dynamic ~38 GB 52.0 1.00×
This NVFP4 W4A4 (cudagraph on) ~22 GB 66.9 1.29×
This NVFP4 W4A4, --enforce-eager ~22 GB ~23 0.44×

The whole speedup lives in CUDA graphs. With --enforce-eager this model is only ~23 tok/s. On a hybrid MoE, eager-mode launch overhead from the many small per-expert kernels dominates wall-clock; capturing CUDA graphs removes it and recovers the full 66.9 tok/s. Do not pass --enforce-eager.

Full write-up (method, sweep, traces): DGX Spark Part 33 — NVFP4 W4A4 MoE + CUDA graphs.

Serving (vLLM, GB10 / SM121)

Toolchain requirement. You need a vLLM build where the SM121 NVFP4 block-scaled MMA actually compiles: cutlass-dsl 4.5.x + flashinfer 0.6.11+. Do not use any sm121-b12x style mod that downgrades cutlass-dsl to 4.4.2 — that breaks the compiled NVFP4 path and silently forces enforce-eager behaviour (back to ~23 tok/s).

vllm serve coolthor/Huihui-Qwen3.6-35B-A3B-abliterated-NVFP4-W4A4 \
  --served-model-name qwen36-abliterated \
  --moe-backend flashinfer_cutlass \
  --kv-cache-dtype fp8 \
  --gpu-memory-utilization 0.50 \
  --max-model-len 262144 \
  --max-num-seqs 4 \
  --max-num-batched-tokens 8192 \
  --reasoning-parser qwen3 \
  --enable-auto-tool-choice \
  --tool-call-parser qwen3_coder \
  --enable-prefix-caching \
  --trust-remote-code

Gotchas

  • Never set --enforce-eager — it drops you from 66.9 → ~23 tok/s (kills the CUDA-graph win).
  • --max-num-batched-tokens >= 2096 is required — the hybrid Mamba/GDN cache block alignment needs it; smaller values fail to allocate the conv/state cache.
  • The flashinfer_cutlass NVFP4 MoE backend wants W4A4 (this model). A W4A16 weight-only NVFP4 export is rejected by this backend — the activations must be FP4 too, which is why this checkpoint is W4A4.
  • MTP heads are preserved but should stay OFF. Swept num_speculative_tokens n=0–4; the no-spec baseline wins. On this hybrid at GB10 memory bandwidth, NVFP4 weight loading and MTP speculation are substitutes competing for the same bandwidth — stacking them does not help. Just don't pass a --speculative-config.

Vision loader-prefix fix (carried forward)

Qwen3_5MoeForConditionalGeneration saves vision weights under model.language_model.visual.*, but vLLM's loader looks for visual.*, so all vision tensors silently skip-load and image input degenerates into !!!!! loops (text-only unaffected). This checkpoint already has the prefix stripped in the affected shard + model.safetensors.index.json, so vision works out of the box.

Quantization recipe

# recipe.yaml
default_stage:
  default_modifiers:
    QuantizationModifier:
      targets: [Linear]
      ignore: ['re:.*lm_head', 're:visual.*', 're:model.visual.*', 're:.*mlp.gate$', 're:.*shared_expert_gate$']
      scheme: NVFP4
      bypass_divisibility_checks: false
from compressed_tensors.utils import save_mtp_tensors_to_checkpoint
from transformers import AutoProcessor, Qwen3_5MoeForConditionalGeneration
from llmcompressor import oneshot
from llmcompressor.modifiers.quantization import QuantizationModifier

MODEL_PATH = "huihui-ai/Huihui-Qwen3.6-35B-A3B-abliterated"
SAVE_DIR = "Huihui-Qwen3.6-35B-A3B-abliterated-NVFP4-W4A4"

model = Qwen3_5MoeForConditionalGeneration.from_pretrained(
    MODEL_PATH, dtype="auto", low_cpu_mem_usage=True,
)
processor = AutoProcessor.from_pretrained(MODEL_PATH)

recipe = QuantizationModifier(
    targets="Linear",
    scheme="NVFP4",  # W4A4: 4-bit weights + 4-bit activations
    ignore=[
        "re:.*lm_head",
        "re:visual.*",
        "re:model.visual.*",
        "re:.*mlp.gate$",
        "re:.*shared_expert_gate$",
    ],
)

oneshot(model=model, recipe=recipe)
model.save_pretrained(SAVE_DIR, max_shard_size="2GB", safe_serialization=True)
processor.save_pretrained(SAVE_DIR)
save_mtp_tensors_to_checkpoint(source_model=MODEL_PATH, dest_dir=SAVE_DIR)

License

Apache-2.0, inherited from huihui-ai/Huihui-Qwen3.6-35B-A3B-abliterated and the upstream Qwen/Qwen3.6-35B-A3B (Apache-2.0). Redistribution permitted.


☕ If this saved you GPU hours, you can buy me a coffee.

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-02docs: add Buy Me a Coffee link to model card1d54e906.6 KB
    Loading...
  2. 2026-06-02Upload folder using huggingface_hub41d776c6.5 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration