← back to catalog · registered 2026-08-22 13:56

YuYu1015/Huihui-Qwopus3.5-27B-v3-abliterated-int4-AutoRound

YuYu1015 Qwen 8.6B second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/YuYu1015%2FHuihui-Qwopus3.5-27B-v3-abliterated-int4-AutoRound"
Response includes
  • classification m1
  • files 25
  • hub_downloads_all_time 164
  • author_summary 23 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
164
17 last 30d - stable
Likes
2
Model age
6mo ago
created 2026-04-13
Downloads over time
Now174→from97↑79%
9312315218297 on Apr 15174 on Oct 11AprMayJunJulAugSepOct
Apr 15 → Oct 11 · 65 snapshots · spans 179 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh
Tags
transformers safetensors qwen3_5 image-text-to-text qwen3.5 4-bit int4 auto-round gptq quantized abliterated dgx-spark

Related

Total size
25.0 GB
Files
25
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-04-20 02:42

Files by quantization

Auxiliary files 25 files 25.1 GB
model-00012-of-00014.safetensors 2.37 GB 2fcb3673 download
model-00014-of-00014.safetensors 2.37 GB 86df4251 download
model-00005-of-00014.safetensors 2.00 GB cb369408 download
model-00007-of-00014.safetensors 2.00 GB 30251fa4 download
model-00003-of-00014.safetensors 2.00 GB f1725dbb download
model-00010-of-00014.safetensors 1.99 GB a2071b9a download
model-00004-of-00014.safetensors 1.99 GB 85b223d1 download
model-00006-of-00014.safetensors 1.99 GB 222b1f39 download
model-00002-of-00014.safetensors 1.98 GB 6a52dca8 download
model-00001-of-00014.safetensors 1.96 GB add58a57 download
model-00008-of-00014.safetensors 1.93 GB e7fc0744 download
model-00009-of-00014.safetensors 1.91 GB 75607c21 download
model-00011-of-00014.safetensors 560 MB 9bbc301c download
model-00013-of-00014.safetensors 10.1 KB daacfe4f download
tokenizer.json 19.1 MB 87a7830d download
model.safetensors.index.json 156 KB 55d15296 download
config.json 32.0 KB 5ce94637 download
quantization_config.json 26.4 KB 4ba5aa41 download
README.md 6.23 KB f08979a2 download
chat_template.jinja 3.95 KB 609532bf download
.gitattributes 1.63 KB b0ac67fe download
processor_config.json 1.27 KB 8af8110f download
tokenizer_config.json 1.14 KB acca40e2 download
preprocessor_config.json 478 B d7c5bf00 download
generation_config.json 163 B 87f878f9 download

README current version from Hugging Face


license: apache-2.0
base_model:

  • huihui-ai/Huihui-Qwopus3.5-27B-v3-abliterated
    base_model_relation: quantized
    library_name: transformers
    pipeline_tag: text-generation
    tags:
  • safetensors
  • qwen3.5
  • 4-bit
  • int4
  • auto-round
  • gptq
  • quantized
  • abliterated
  • dgx-spark
  • dflash
  • vllm
    language:
  • en
  • zh

Huihui-Qwopus3.5-27B-v3-abliterated-int4-AutoRound

English | 繁體中文


English

INT4 AutoRound quantization of huihui-ai/Huihui-Qwopus3.5-27B-v3-abliterated, optimized for NVIDIA DGX Spark (GB10 SM121) with Marlin INT4 kernel acceleration.

Model Details

Item Value
Architecture Dense 27B + GDN (Mamba) + Attention hybrid
Base model Qwen/Qwen3.5-27B
Fine-tuned by huihui-ai (Qwopus v3 distillation + abliteration)
Quantized by YuYu1015
Model size ~26 GB (vs ~51 GB BF16 original)
Context length Up to 65,536 tokens
Thinking mode Supported (enable_thinking: true/false)
Tool calling Supported (qwen3_coder parser)

Quantization Details

Item Value
Method Intel AutoRound v0.12.2
Bits 4
Group size 128
Format auto_round (GPTQ-compatible)
Iterations 200
Calibration samples 512
Calibration sequence length 2048
Hardware NVIDIA DGX Spark (GB10, 128GB unified memory)

Layers Preserved in BF16

The following layers are not quantized to preserve model quality:

Layer Reason
lm_head Output head, sensitive to quantization noise
embed_tokens Input embeddings (auto-excluded by shape)
linear_attn.* GDN/DeltaNet layers, may output zeros if quantized
model.visual.* Vision encoder (auto-excluded by shape)

Speculative Decoding

DFlash (requires separate drafter model):

--speculative-config '{"method": "dflash", "model": "z-lab/Qwen3.5-27B-DFlash", "num_speculative_tokens": 16}'

Note: The DFlash drafter was trained on the original Qwen3.5-27B. Acceptance rate on the abliterated/distilled variant may be lower than on the original model.

Serving with vLLM

vllm serve /path/to/model \
    --quantization gptq_marlin \
    --served-model-name qwen3.5-27b \
    --reasoning-parser qwen3 \
    --enable-auto-tool-choice \
    --tool-call-parser qwen3_coder \
    --kv-cache-dtype auto \
    --gpu-memory-utilization 0.90 \
    --max-model-len 65536 \
    --enable-prefix-caching \
    --enable-chunked-prefill \
    --trust-remote-code \
    --language-model-only

DGX Spark (SM121) Compatibility Notes

  • Use --quantization gptq_marlin for Marlin INT4 kernel (Dense model, not MoE)
  • FP8 KV cache is not compatible with GDN non-causal attention layers; use --kv-cache-dtype auto
  • NVFP4 is not supported on SM121 (missing cvt.e2m1x2 instruction)
  • Runtime FP8 (--quantization fp8) is not compatible with DFlash
  • --language-model-only skips vision encoder profiling for text-only inference
  • Clear page cache before starting on UMA: sudo sh -c 'echo 3 > /proc/sys/vm/drop_caches'

Safety Warning

This model has safety filtering removed (abliterated) and may generate inappropriate content. Users are solely responsible for all consequences arising from its use.

Credits


繁體中文

huihui-ai/Huihui-Qwopus3.5-27B-v3-abliterated 的 INT4 AutoRound 量化版本,針對 NVIDIA DGX Spark (GB10 SM121) 最佳化。

模型資訊

項目 數值
架構 Dense 27B + GDN (Mamba) + Attention 混合
基礎模型 Qwen/Qwen3.5-27B
微調者 huihui-ai(Qwopus v3 蒸餾 + abliteration)
量化者 YuYu1015
模型大小 ~26 GB(原版 BF16 約 51 GB)
Context 長度 最高 65,536 tokens
思考模式 支援(enable_thinking: true/false)
工具呼叫 支援(qwen3_coder parser)

量化詳情

項目 數值
方法 Intel AutoRound v0.12.2
位元數 4
Group size 128
格式 auto_round(GPTQ 相容)
迭代次數 200
校準樣本數 512
校準序列長度 2048
量化硬體 NVIDIA DGX Spark(GB10, 128GB 統一記憶體)

保留 BF16 的層

以下層未被量化以保持模型品質:

層 原因
lm_head 輸出頭,對量化雜訊敏感
embed_tokens 輸入嵌入(因 shape 自動排除)
linear_attn.* GDN/DeltaNet 層,量化後可能輸出零
model.visual.* 視覺編碼器(因 shape 自動排除)

DGX Spark (SM121) 相容性說明

  • 使用 --quantization gptq_marlin 啟用 Marlin INT4 kernel(Dense 模型,非 MoE)
  • FP8 KV cache 與 GDN non-causal attention 不相容,請使用 --kv-cache-dtype auto
  • NVFP4 在 SM121 上不支援(缺少 cvt.e2m1x2 指令)
  • Runtime FP8(--quantization fp8)與 DFlash 不相容
  • UMA 架構啟動前請先清除 page cache:sudo sh -c 'echo 3 > /proc/sys/vm/drop_caches'

安全警告

此模型已移除安全過濾機制(abliterated),可能產生不當內容。使用者須自行承擔所有風險與法律責任。

致謝

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-04-20Update README.md61b22606.2 KB
    Loading...
  2. 2026-04-13Update README.mdc66697b12.2 KB
    Loading...
  3. 2026-04-13Create README.md4b2132d7.6 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration