← back to catalog · registered 2026-08-22 13:56

YuYu1015/Huihui-Qwen3-30B-A3B-Thinking-2507-abliterated-NVFP4

YuYu1015 Qwen 15B MoE second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/YuYu1015%2FHuihui-Qwen3-30B-A3B-Thinking-2507-abliterated-NVFP4"
Response includes
  • classification m1
  • files 17
  • benchmarks 11 entries
  • hub_downloads_all_time 1,123
  • author_summary 23 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
1K
29 last 30d - cooling
Likes
0
Model age
6mo ago
created 2026-04-08
Downloads over time
Now1.1K→from901↑26%
8899791.1K1.2K901 on Apr 151.1K on Oct 111.1K on Oct 9AprMayJunJulAugSepOct
Apr 15 → Oct 11 · 65 snapshots · spans 179 days

Benchmarks

Benchmark Score Source
Entertainment 1.3 UGI
Hazardous 1.2 UGI
Natural Intelligence 15.27 UGI
Political lean -7.8% UGI
Sensitive-Info 10.63 UGI
SocPol 0.7 UGI
UGI 33.75 UGI
Willingness (10) 8 UGI
W10-Adherence 8 UGI
W10-Direct 8 UGI
Writing 31.14 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh
Tags
transformers safetensors qwen3_moe text-generation qwen3 moe nvfp4 4-bit quantized abliterated dgx-spark blackwell

Related

Total size
16.9 GB
Files
17
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-04-12 18:10

Files by quantization

Auxiliary files 17 files 16.9 GB
model-00002-of-00004.safetensors 4.66 GB 6ee7067b download
model-00001-of-00004.safetensors 4.66 GB 60984b67 download
model-00003-of-00004.safetensors 4.66 GB fd44c16a download
model-00004-of-00004.safetensors 2.88 GB bc466658 download
tokenizer.json 10.9 MB c0acdaba download
model.safetensors.index.json 7.10 MB 9ec9293f download
vocab.json 2.65 MB 4783fe10 download
merges.txt 1.59 MB 31349551 download
README.md 11.9 KB 49067616 download
tokenizer_config.json 5.28 KB 1d4fba2d download
chat_template.jinja 3.95 KB 2e2f69c3 download
config.json 3.86 KB 25c6214b download
.gitattributes 1.53 KB 52373fe2 download
added_tokens.json 707 B b54f9135 download
special_tokens_map.json 613 B ac23c0aa download
recipe.yaml 308 B 1b23e516 download
generation_config.json 214 B 9e289a1a download

README current version from Hugging Face


license: apache-2.0
base_model:

  • huihui-ai/Huihui-Qwen3-30B-A3B-Thinking-2507-abliterated
    base_model_relation: quantized
    library_name: transformers
    pipeline_tag: text-generation
    tags:
  • safetensors
  • qwen3
  • moe
  • nvfp4
  • 4-bit
  • quantized
  • abliterated
  • dgx-spark
  • blackwell
  • gb10
  • sm121
  • vllm
  • llm-compressor
    language:
  • en
  • zh

Huihui-Qwen3-30B-A3B-Thinking-2507-abliterated-NVFP4

English | 繁體中文


English

[!TIP]
Re-quantized on 2026-04-13 with corrected ignore list (mlp.gate + embed_tokens now preserved in BF16), fixing routing quality issues in the previous release.

[!WARNING]
NVIDIA DGX Spark (GB10 SM121) — Driver 590.48+ / CUDA 13.1+

As of April 2026, NVFP4 software support on SM121 is still incomplete. The native W4A4 compute path is not yet functional on this hardware — the runtime silently falls back to W4A16 (BF16 activations), negating the theoretical throughput advantage of FP4.

If accuracy and inference speed are your priority, we recommend the INT4 AutoRound version:
👉 YuYu1015/Huihui-Qwen3.5-35B-A3B-Claude-4.6-Opus-abliterated-int4-AutoRound

INT4 AutoRound leverages the mature W4A16 Marlin kernel path on DGX Spark, offering more thorough calibration (~99.5% quality retention) and significantly more stable performance. The full potential of NVFP4 will only be unlocked once NVIDIA delivers complete W4A4 kernel support for SM121.

NVFP4 quantization of huihui-ai/Huihui-Qwen3-30B-A3B-Thinking-2507-abliterated, optimized for NVIDIA DGX Spark (GB10 SM121).

Model Details

Item Value
Architecture MoE (30B total, 3B active), 48 layers, 128 experts, top-8 routing
Base model Qwen/Qwen3-30B-A3B
Fine-tuned by huihui-ai (Thinking 2507 + abliteration)
Quantized by YuYu1015
Model size ~18.1 GB (NVFP4, vs ~60 GB BF16 original)
Context length Up to 131,072 tokens
Thinking mode Built-in Chain-of-Thought reasoning (enabled by default)
Tool calling Supported (qwen3_coder parser)

Quantization Details

Item Value
Method llm-compressor v0.10.0.1
Scheme NVFP4 (E2M1 + FP8 per-group scaling, group size 16)
Format compressed-tensors v0.14.0.1
Calibration dataset HuggingFaceH4/ultrachat_200k (train_sft split)
Calibration samples 512
Calibration sequence length 2048
MoE expert calibration moe_calibrate_all_experts=True (all experts receive calibration data)
Hardware NVIDIA DGX Spark (GB10, 128GB unified memory)
Environment transformers==4.57.1 + llm-compressor==0.10.0.1

Layers Preserved in BF16

The following layers are not quantized to preserve model quality:

Layer Reason
lm_head Output head, sensitive to quantization noise
re:.*mlp.gate$ MoE routing gate — critical for expert selection accuracy
re:.*embed_tokens$ Input embeddings

Serving with vLLM

vllm serve /path/to/model \
    --quantization compressed-tensors \
    --served-model-name qwen3-30b \
    --reasoning-parser qwen3 \
    --enable-auto-tool-choice \
    --tool-call-parser qwen3_coder \
    --kv-cache-dtype fp8 \
    --gpu-memory-utilization 0.90 \
    --max-model-len 32768 \
    --enable-prefix-caching \
    --enable-chunked-prefill \
    --trust-remote-code

DGX Spark (SM121) Compatibility Notes

  • NVFP4 on SM121 falls back to W4A16 (native W4A4 path not yet supported, missing cvt.e2m1x2 instruction)
  • Qwen3 (non-3.5) has no Mamba layers, so FP8 KV cache works safely
  • Qwen3 has no GDN, so linear_attn does not need to be excluded
  • Clear page cache before starting on UMA: sudo sh -c 'echo 3 > /proc/sys/vm/drop_caches'

Safety Warning

This model has safety filtering removed (abliterated) and may generate inappropriate content. Users are solely responsible for all consequences arising from its use.

Credits


繁體中文

[!TIP]
2026-04-13 重新量化上傳,修正先前版本的 ignore list(mlp.gate 與 embed_tokens 現在保留 BF16),解決 MoE 路由品質問題。

[!WARNING]
NVIDIA DGX Spark (GB10 SM121) 使用者 — Driver 590.48+ / CUDA 13.1+

截至 2026 年 4 月,NVFP4 在 SM121 上的軟體支援仍不完整。原生 W4A4 運算路徑尚未在此硬體上就緒——執行時會靜默退回 W4A16(BF16 activation),FP4 的理論吞吐量優勢無法發揮。

若精度與推理速度為首要考量,建議改用 INT4 AutoRound 版本:
👉 YuYu1015/Huihui-Qwen3.5-35B-A3B-Claude-4.6-Opus-abliterated-int4-AutoRound

INT4 AutoRound 在 DGX Spark 上使用成熟的 W4A16 Marlin kernel 路徑,校準更完整(品質保留約 99.5%),效能顯著更穩定。待 NVIDIA 為 SM121 提供完整的 W4A4 kernel 支援後,NVFP4 的真正優勢才能發揮。

huihui-ai/Huihui-Qwen3-30B-A3B-Thinking-2507-abliterated 的 NVFP4 量化版本,針對 NVIDIA DGX Spark (GB10 SM121) 最佳化。

模型資訊

項目 數值
架構 MoE(30B 總參數, 3B 活躍),48 層,128 experts,top-8 routing
基礎模型 Qwen/Qwen3-30B-A3B
微調者 huihui-ai(Thinking 2507 + abliteration)
量化者 YuYu1015
模型大小 ~18.1 GB(NVFP4,原版 BF16 約 60 GB)
Context 長度 最高 131,072 tokens
思考模式 內建思維鏈推理(預設啟用)
工具呼叫 支援(qwen3_coder parser)

量化詳情

項目 數值
方法 llm-compressor v0.10.0.1
方案 NVFP4(E2M1 + FP8 逐群縮放,群組大小 16)
格式 compressed-tensors v0.14.0.1
校準資料集 HuggingFaceH4/ultrachat_200k (train_sft 分割)
校準樣本數 512
校準序列長度 2048
MoE 專家校準 moe_calibrate_all_experts=True(所有專家都接收校準資料)
量化硬體 NVIDIA DGX Spark(GB10, 128GB 統一記憶體)
環境 transformers==4.57.1 + llm-compressor==0.10.0.1

保留 BF16 的層

以下層未被量化以保持模型品質:

層 原因
lm_head 輸出頭,對量化雜訊敏感
re:.*mlp.gate$ MoE 路由閘——對專家選擇精度至關重要
re:.*embed_tokens$ 輸入嵌入

vLLM 部署

vllm serve /path/to/model \
    --quantization compressed-tensors \
    --served-model-name qwen3-30b \
    --reasoning-parser qwen3 \
    --enable-auto-tool-choice \
    --tool-call-parser qwen3_coder \
    --kv-cache-dtype fp8 \
    --gpu-memory-utilization 0.90 \
    --max-model-len 32768 \
    --enable-prefix-caching \
    --enable-chunked-prefill \
    --trust-remote-code

DGX Spark (SM121) 相容性說明

  • NVFP4 在 SM121 上會退回 W4A16(原生 W4A4 路徑尚未支援,缺少 cvt.e2m1x2 指令)
  • Qwen3(非 3.5)沒有 Mamba 層,FP8 KV cache 可以安全使用
  • Qwen3 沒有 GDN,linear_attn 不需要排除
  • UMA 架構啟動前請先清除 page cache:sudo sh -c 'echo 3 > /proc/sys/vm/drop_caches'

安全警告

此模型已移除安全過濾機制(abliterated),可能產生不當內容。使用者須自行承擔所有風險與法律責任。

致謝

README history 4 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-04-12Update README.mdd375e9e11.9 KB
    Loading...
  2. 2026-04-11Update README.mdb2ddb0515.1 KB
    Loading...
  3. 2026-04-11Update README.mdb2e26e811.2 KB
    Loading...
  4. 2026-04-08Create README.mda2bfdb99 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration