← back to catalog · registered 2026-08-22 13:56

YuYu1015/Huihui-Qwen3-30B-A3B-Instruct-2507-abliterated-NVFP4

YuYu1015 Qwen 15B MoE second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/YuYu1015%2FHuihui-Qwen3-30B-A3B-Instruct-2507-abliterated-NVFP4"
Response includes
  • classification m1
  • files 17
  • benchmarks 11 entries
  • hub_downloads_all_time 4,445
  • author_summary 23 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
4K
320 last 30d - cooling
Likes
0
Model age
6mo ago
created 2026-04-12
Downloads over time
Now4.6K→from137↑3,279%
01.7K3.4K5.1K137 on Apr 154.6K on Oct 11AprMayJunJulAugSepOct
Apr 15 → Oct 11 · 66 snapshots · spans 179 days

Benchmarks

Benchmark Score Source
Entertainment 0.9 UGI
Hazardous 2.4 UGI
Natural Intelligence 15.65 UGI
Political lean -21.0% UGI
Sensitive-Info 15.14 UGI
SocPol 1.6 UGI
UGI 38.43 UGI
Willingness (10) 8.5 UGI
W10-Adherence 9 UGI
W10-Direct 8 UGI
Writing 30.69 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh
Tags
transformers safetensors qwen3_moe text-generation qwen3 moe nvfp4 4-bit quantized abliterated dgx-spark blackwell

Related

Total size
16.9 GB
Files
17
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-04-12 23:00

Files by quantization

Auxiliary files 17 files 16.9 GB
model-00002-of-00004.safetensors 4.66 GB 7dbcff32 download
model-00001-of-00004.safetensors 4.66 GB 332797a9 download
model-00003-of-00004.safetensors 4.66 GB c370e0a6 download
model-00004-of-00004.safetensors 2.88 GB 78f4bfec download
tokenizer.json 10.9 MB c0acdaba download
model.safetensors.index.json 7.10 MB 9ec9293f download
vocab.json 2.65 MB 4783fe10 download
merges.txt 1.59 MB 31349551 download
README.md 11.8 KB 37c8a8d5 download
tokenizer_config.json 5.28 KB 1d4fba2d download
chat_template.jinja 3.95 KB a18870ad download
config.json 3.86 KB 25c6214b download
.gitattributes 1.53 KB 52373fe2 download
added_tokens.json 707 B b54f9135 download
special_tokens_map.json 613 B ac23c0aa download
recipe.yaml 308 B 1b23e516 download
generation_config.json 213 B bdb4e037 download

README current version from Hugging Face


license: apache-2.0
base_model:

  • huihui-ai/Huihui-Qwen3-30B-A3B-Instruct-2507-abliterated
    base_model_relation: quantized
    library_name: transformers
    pipeline_tag: text-generation
    tags:
  • safetensors
  • qwen3
  • moe
  • nvfp4
  • 4-bit
  • quantized
  • abliterated
  • dgx-spark
  • blackwell
  • gb10
  • sm121
  • vllm
  • llm-compressor
    language:
  • en
  • zh

Huihui-Qwen3-30B-A3B-Instruct-2507-abliterated-NVFP4

English | 繁體中文


English

[!TIP]
Quantized on 2026-04-13 with mlp.gate + embed_tokens preserved in BF16 for MoE routing accuracy.

[!WARNING]
NVIDIA DGX Spark (GB10 SM121) — Driver 590.48+ / CUDA 13.1+

As of April 2026, NVFP4 software support on SM121 is still incomplete. The native W4A4 compute path is not yet functional on this hardware — the runtime silently falls back to W4A16 (BF16 activations), negating the theoretical throughput advantage of FP4.

If accuracy and inference speed are your priority, we recommend the INT4 AutoRound version:
👉 YuYu1015/Huihui-Qwen3.5-35B-A3B-Claude-4.6-Opus-abliterated-int4-AutoRound

INT4 AutoRound leverages the mature W4A16 Marlin kernel path on DGX Spark, offering more thorough calibration (~99.5% quality retention) and significantly more stable performance. The full potential of NVFP4 will only be unlocked once NVIDIA delivers complete W4A4 kernel support for SM121.

NVFP4 quantization of huihui-ai/Huihui-Qwen3-30B-A3B-Instruct-2507-abliterated, optimized for NVIDIA DGX Spark (GB10 SM121).

Model Details

Item Value
Architecture MoE (30B total, 3B active), 48 layers, 128 experts, top-8 routing
Base model Qwen/Qwen3-30B-A3B
Fine-tuned by huihui-ai (Instruct 2507 + abliteration)
Quantized by YuYu1015
Model size ~18.1 GB (NVFP4, vs ~57 GB BF16 original)
Context length Up to 262,144 tokens
Thinking mode Disabled (instruction-tuned, direct responses)
Tool calling Supported (qwen3_coder parser)

Quantization Details

Item Value
Method llm-compressor v0.10.0.1
Scheme NVFP4 (E2M1 + FP8 per-group scaling, group size 16)
Format compressed-tensors v0.14.0.1
Calibration dataset HuggingFaceH4/ultrachat_200k (train_sft split)
Calibration samples 512
Calibration sequence length 2048
MoE expert calibration moe_calibrate_all_experts=True (all experts receive calibration data)
Hardware NVIDIA DGX Spark (GB10, 128GB unified memory)
Environment transformers==4.57.1 + llm-compressor==0.10.0.1

Layers Preserved in BF16

The following layers are not quantized to preserve model quality:

Layer Reason
lm_head Output head, sensitive to quantization noise
re:.*mlp.gate$ MoE routing gate — critical for expert selection accuracy
re:.*embed_tokens$ Input embeddings

Serving with vLLM

vllm serve /path/to/model \
    --quantization compressed-tensors \
    --served-model-name qwen3-30b \
    --enable-auto-tool-choice \
    --tool-call-parser qwen3_coder \
    --kv-cache-dtype fp8 \
    --gpu-memory-utilization 0.90 \
    --max-model-len 32768 \
    --enable-prefix-caching \
    --enable-chunked-prefill \
    --trust-remote-code

DGX Spark (SM121) Compatibility Notes

  • NVFP4 on SM121 falls back to W4A16 (native W4A4 path not yet supported, missing cvt.e2m1x2 instruction)
  • Qwen3 (non-3.5) has no Mamba layers, so FP8 KV cache works safely
  • Qwen3 has no GDN, so linear_attn does not need to be excluded
  • Clear page cache before starting on UMA: sudo sh -c 'echo 3 > /proc/sys/vm/drop_caches'

Safety Warning

This model has safety filtering removed (abliterated) and may generate inappropriate content. Users are solely responsible for all consequences arising from its use.

Credits


繁體中文

[!TIP]
2026-04-13 量化上傳,mlp.gate 與 embed_tokens 保留 BF16 以確保 MoE 路由品質。

[!WARNING]
NVIDIA DGX Spark (GB10 SM121) 使用者 — Driver 590.48+ / CUDA 13.1+

截至 2026 年 4 月,NVFP4 在 SM121 上的軟體支援仍不完整。原生 W4A4 運算路徑尚未在此硬體上就緒——執行時會靜默退回 W4A16(BF16 activation),FP4 的理論吞吐量優勢無法發揮。

若精度與推理速度為首要考量,建議改用 INT4 AutoRound 版本:
👉 YuYu1015/Huihui-Qwen3.5-35B-A3B-Claude-4.6-Opus-abliterated-int4-AutoRound

INT4 AutoRound 在 DGX Spark 上使用成熟的 W4A16 Marlin kernel 路徑,校準更完整(品質保留約 99.5%),效能顯著更穩定。待 NVIDIA 為 SM121 提供完整的 W4A4 kernel 支援後,NVFP4 的真正優勢才能發揮。

huihui-ai/Huihui-Qwen3-30B-A3B-Instruct-2507-abliterated 的 NVFP4 量化版本,針對 NVIDIA DGX Spark (GB10 SM121) 最佳化。

模型資訊

項目 數值
架構 MoE(30B 總參數, 3B 活躍),48 層,128 experts,top-8 routing
基礎模型 Qwen/Qwen3-30B-A3B
微調者 huihui-ai(Instruct 2507 + abliteration)
量化者 YuYu1015
模型大小 ~18.1 GB(NVFP4,原版 BF16 約 57 GB)
Context 長度 最高 262,144 tokens
思考模式 停用(指令微調版,直接回應)
工具呼叫 支援(qwen3_coder parser)

量化詳情

項目 數值
方法 llm-compressor v0.10.0.1
方案 NVFP4(E2M1 + FP8 逐群縮放,群組大小 16)
格式 compressed-tensors v0.14.0.1
校準資料集 HuggingFaceH4/ultrachat_200k (train_sft 分割)
校準樣本數 512
校準序列長度 2048
MoE 專家校準 moe_calibrate_all_experts=True(所有專家都接收校準資料)
量化硬體 NVIDIA DGX Spark(GB10, 128GB 統一記憶體)
環境 transformers==4.57.1 + llm-compressor==0.10.0.1

保留 BF16 的層

以下層未被量化以保持模型品質:

層 原因
lm_head 輸出頭,對量化雜訊敏感
re:.*mlp.gate$ MoE 路由閘——對專家選擇精度至關重要
re:.*embed_tokens$ 輸入嵌入

vLLM 部署

vllm serve /path/to/model \
    --quantization compressed-tensors \
    --served-model-name qwen3-30b \
    --enable-auto-tool-choice \
    --tool-call-parser qwen3_coder \
    --kv-cache-dtype fp8 \
    --gpu-memory-utilization 0.90 \
    --max-model-len 32768 \
    --enable-prefix-caching \
    --enable-chunked-prefill \
    --trust-remote-code

DGX Spark (SM121) 相容性說明

  • NVFP4 在 SM121 上會退回 W4A16(原生 W4A4 路徑尚未支援,缺少 cvt.e2m1x2 指令)
  • Qwen3(非 3.5)沒有 Mamba 層,FP8 KV cache 可以安全使用
  • Qwen3 沒有 GDN,linear_attn 不需要排除
  • UMA 架構啟動前請先清除 page cache:sudo sh -c 'echo 3 > /proc/sys/vm/drop_caches'

安全警告

此模型已移除安全過濾機制(abliterated),可能產生不當內容。使用者須自行承擔所有風險與法律責任。

致謝

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-04-12Update README.mdaf962b811.8 KB
    Loading...
  2. 2026-04-12Update README.mdccbc58611.8 KB
    Loading...
  3. 2026-04-12Create README.md7e355ac11.8 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration