← back to catalog · registered 2026-08-22 13:56

YuYu1015/Huihui-Qwopus3.5-27B-v3-abliterated-NVFP4

YuYu1015 Qwen 9.4B second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/YuYu1015%2FHuihui-Qwopus3.5-27B-v3-abliterated-NVFP4"
Response includes
  • classification m1
  • files 9
  • hub_downloads_all_time 209
  • author_summary 23 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
209
29 last 30d - stable
Likes
1
Model age
6mo ago
created 2026-04-13
Downloads over time
Now223→from122↑83%
117156194233122 on Apr 15223 on Oct 11223 on Oct 10AprMayJunJulAugSepOct
Apr 15 → Oct 11 · 65 snapshots · spans 179 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh
Tags
transformers safetensors qwen3_5_text text-generation qwen3.5 dense nvfp4 4-bit quantized abliterated dgx-spark blackwell

Related

Total size
24.9 GB
Files
9
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-04-13 14:15

Files by quantization

Auxiliary files 9 files 25.0 GB
model.safetensors 24.9 GB e80b3d02 download
tokenizer.json 19.1 MB 6a0316e3 download
config.json 15.0 KB 9ab77abf download
README.md 12.2 KB cd8997e3 download
chat_template.jinja 3.95 KB 609532bf download
.gitattributes 1.53 KB 52373fe2 download
tokenizer_config.json 1.14 KB acca40e2 download
recipe.yaml 249 B a736efa4 download
generation_config.json 142 B 1e990d5f download

README current version from Hugging Face


license: apache-2.0
base_model:

  • huihui-ai/Huihui-Qwopus3.5-27B-v3-abliterated
    base_model_relation: quantized
    library_name: transformers
    pipeline_tag: text-generation
    tags:
  • safetensors
  • qwen3.5
  • dense
  • nvfp4
  • 4-bit
  • quantized
  • abliterated
  • dgx-spark
  • blackwell
  • gb10
  • sm121
  • vllm
  • llm-compressor
    language:
  • en
  • zh

Huihui-Qwopus3.5-27B-v3-abliterated-NVFP4

English | 繁體中文


English

[!TIP]
Quantized on 2026-04-13 with lm_head / linear_attn (GDN) / mtp / visual / embed_tokens preserved in BF16.

[!WARNING]
NVIDIA DGX Spark (GB10 SM121) — Driver 590.48+ / CUDA 13.1+

As of April 2026, NVFP4 software support on SM121 is still incomplete. The native W4A4 compute path is not yet functional on this hardware — the runtime silently falls back to W4A16 (BF16 activations), negating the theoretical throughput advantage of FP4.

If accuracy and inference speed are your priority, we recommend the INT4 AutoRound version:
👉 YuYu1015/Huihui-Qwopus3.5-27B-v3-abliterated-int4-AutoRound

INT4 AutoRound leverages the mature W4A16 Marlin kernel path on DGX Spark, offering more thorough calibration (~99.5% quality retention) and significantly more stable performance. The full potential of NVFP4 will only be unlocked once NVIDIA delivers complete W4A4 kernel support for SM121.

NVFP4 quantization of huihui-ai/Huihui-Qwopus3.5-27B-v3-abliterated, optimized for NVIDIA DGX Spark (GB10 SM121).

Model Details

Item Value
Architecture Dense 27B + GDN (Mamba) + Attention hybrid
Base model Qwen/Qwen3.5-27B
Fine-tuned by huihui-ai (Qwopus v3 distillation + abliteration)
Quantized by YuYu1015
Model size ~25 GB (NVFP4, vs ~51 GB BF16 original)
Context length Up to 262,144 tokens
Thinking mode Supported (enable_thinking: true/false)
Tool calling Supported (qwen3_coder parser)
MTP Built-in MTP weights included (preserved in BF16)

Quantization Details

Item Value
Method llm-compressor (main branch, PR #2608)
Scheme NVFP4 (E2M1 + FP8 per-group scaling, group size 16)
Format compressed-tensors (main branch)
Calibration dataset HuggingFaceH4/ultrachat_200k (train_sft split)
Calibration samples 512
Calibration sequence length 2048
Hardware NVIDIA DGX Spark (GB10, 128GB unified memory)
Environment transformers>=5.0 + llm-compressor main (Qwen3.5 qwen3_5 model_type requires tf5)

Layers Preserved in BF16

The following layers are not quantized to preserve model quality:

Layer Reason
lm_head Output head, sensitive to quantization noise
re:.*linear_attn.* GDN/DeltaNet (Mamba) layers — may output zeros if quantized
re:.*mtp\..* Multi-Token Prediction weights
re:.*visual\..* Vision encoder
re:.*embed_tokens$ Input embeddings

Serving with vLLM

vllm serve /path/to/model \
    --quantization compressed-tensors \
    --served-model-name qwen3.5-27b \
    --reasoning-parser qwen3 \
    --enable-auto-tool-choice \
    --tool-call-parser qwen3_coder \
    --kv-cache-dtype auto \
    --gpu-memory-utilization 0.90 \
    --max-model-len 65536 \
    --enable-prefix-caching \
    --enable-chunked-prefill \
    --trust-remote-code \
    --language-model-only

DGX Spark (SM121) Compatibility Notes

  • NVFP4 on SM121 falls back to W4A16 (native W4A4 path not yet supported, missing cvt.e2m1x2 instruction)
  • FP8 KV cache is not compatible with GDN non-causal attention layers; use --kv-cache-dtype auto
  • --language-model-only skips vision encoder profiling for text-only inference
  • Clear page cache before starting on UMA: sudo sh -c 'echo 3 > /proc/sys/vm/drop_caches'

Safety Warning

This model has safety filtering removed (abliterated) and may generate inappropriate content. Users are solely responsible for all consequences arising from its use.

Credits


繁體中文

[!TIP]
2026-04-13 量化上傳,lm_head / linear_attn (GDN) / mtp / visual / embed_tokens 保留 BF16。

[!WARNING]
NVIDIA DGX Spark (GB10 SM121) 使用者 — Driver 590.48+ / CUDA 13.1+

截至 2026 年 4 月,NVFP4 在 SM121 上的軟體支援仍不完整。原生 W4A4 運算路徑尚未在此硬體上就緒——執行時會靜默退回 W4A16(BF16 activation),FP4 的理論吞吐量優勢無法發揮。

若精度與推理速度為首要考量,建議改用 INT4 AutoRound 版本:
👉 YuYu1015/Huihui-Qwopus3.5-27B-v3-abliterated-int4-AutoRound

INT4 AutoRound 在 DGX Spark 上使用成熟的 W4A16 Marlin kernel 路徑,校準更完整(品質保留約 99.5%),效能顯著更穩定。待 NVIDIA 為 SM121 提供完整的 W4A4 kernel 支援後,NVFP4 的真正優勢才能發揮。

huihui-ai/Huihui-Qwopus3.5-27B-v3-abliterated 的 NVFP4 量化版本,針對 NVIDIA DGX Spark (GB10 SM121) 最佳化。

模型資訊

項目 數值
架構 Dense 27B + GDN (Mamba) + Attention 混合
基礎模型 Qwen/Qwen3.5-27B
微調者 huihui-ai(Qwopus v3 蒸餾 + abliteration)
量化者 YuYu1015
模型大小 ~14 GB(NVFP4,原版 BF16 約 54 GB)
Context 長度 最高 262,144 tokens
思考模式 支援(enable_thinking: true/false)
工具呼叫 支援(qwen3_coder parser)
MTP 內建 MTP 權重(保留 BF16)

量化詳情

項目 數值
方法 llm-compressor(main 分支,PR #2608)
方案 NVFP4(E2M1 + FP8 逐群縮放,群組大小 16)
格式 compressed-tensors(main 分支)
校準資料集 HuggingFaceH4/ultrachat_200k (train_sft 分割)
校準樣本數 512
校準序列長度 2048
量化硬體 NVIDIA DGX Spark(GB10, 128GB 統一記憶體)
環境 transformers>=5.0 + llm-compressor main(Qwen3.5 qwen3_5 model_type 需要 tf5)

保留 BF16 的層

以下層未被量化以保持模型品質:

層 原因
lm_head 輸出頭,對量化雜訊敏感
re:.*linear_attn.* GDN/DeltaNet (Mamba) 層,量化後可能輸出零
re:.*mtp\..* Multi-Token Prediction 權重
re:.*visual\..* 視覺編碼器
re:.*embed_tokens$ 輸入嵌入

vLLM 部署

vllm serve /path/to/model \
    --quantization compressed-tensors \
    --served-model-name qwen3.5-27b \
    --reasoning-parser qwen3 \
    --enable-auto-tool-choice \
    --tool-call-parser qwen3_coder \
    --kv-cache-dtype auto \
    --gpu-memory-utilization 0.90 \
    --max-model-len 65536 \
    --enable-prefix-caching \
    --enable-chunked-prefill \
    --trust-remote-code \
    --language-model-only

DGX Spark (SM121) 相容性說明

  • NVFP4 在 SM121 上會退回 W4A16(原生 W4A4 路徑尚未支援,缺少 cvt.e2m1x2 指令)
  • FP8 KV cache 與 GDN non-causal attention 不相容,請使用 --kv-cache-dtype auto
  • --language-model-only 跳過視覺編碼器 profiling,加速純文字推理啟動
  • UMA 架構啟動前請先清除 page cache:sudo sh -c 'echo 3 > /proc/sys/vm/drop_caches'

安全警告

此模型已移除安全過濾機制(abliterated),可能產生不當內容。使用者須自行承擔所有風險與法律責任。

致謝

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-04-13Update README.mdeb94ae812.2 KB
    Loading...
  2. 2026-04-13Create README.md72f2bbc12.2 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration