← back to catalog · registered 2026-08-22 13:56

YuYu1015/Huihui-Qwen3.5-27B-Claude-4.6-Opus-abliterated-int4-AutoRound

YuYu1015 Qwen 8.6B second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/YuYu1015%2FHuihui-Qwen3.5-27B-Claude-4.6-Opus-abliterated-int4-AutoRound"
Response includes
  • classification m1
  • files 26
  • hub_downloads_all_time 365
  • author_summary 23 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
365
21 last 30d - cooling
Likes
1
Model age
6mo ago
created 2026-04-11
Downloads over time
Now377→from32↑1,078%
1514727941232 on Apr 15377 on Oct 11AprMayJunJulAugSepOct
Apr 15 → Oct 11 · 65 snapshots · spans 179 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh
Tags
transformers safetensors qwen3_5 image-text-to-text qwen3.5 4-bit int4 auto-round gptq quantized abliterated dgx-spark

Related

Total size
25.3 GB
Files
26
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-04-12 18:11

Files by quantization

Auxiliary files 26 files 25.3 GB
model-00012-of-00014.safetensors 2.37 GB 2fcb3673 download
model-00014-of-00014.safetensors 2.37 GB 86df4251 download
model-00005-of-00014.safetensors 2.00 GB bdca9202 download
model-00007-of-00014.safetensors 2.00 GB a613d69e download
model-00003-of-00014.safetensors 2.00 GB 2a5b3ca0 download
model-00010-of-00014.safetensors 1.99 GB 7f7f9d4c download
model-00004-of-00014.safetensors 1.99 GB 5cadccd9 download
model-00006-of-00014.safetensors 1.99 GB c4238422 download
model-00002-of-00014.safetensors 1.98 GB ae2961a6 download
model-00001-of-00014.safetensors 1.96 GB 6596486d download
model-00008-of-00014.safetensors 1.93 GB 96ee120b download
model-00009-of-00014.safetensors 1.91 GB b1ae08db download
model-00011-of-00014.safetensors 560 MB 9bbc301c download
model_extra_tensors.safetensors 210 MB e7dbe627 download
model-00013-of-00014.safetensors 10.1 KB daacfe4f download
tokenizer.json 19.1 MB 87a7830d download
model.safetensors.index.json 159 KB 8e61f49b download
config.json 32.0 KB dfb096e2 download
quantization_config.json 26.4 KB 4ba5aa41 download
README.md 6.30 KB d40674aa download
chat_template.jinja 3.95 KB 609532bf download
.gitattributes 1.64 KB f5394cd9 download
processor_config.json 1.27 KB 8af8110f download
tokenizer_config.json 1.14 KB aeb7593d download
preprocessor_config.json 478 B d7c5bf00 download
generation_config.json 163 B 1a7183a4 download

README current version from Hugging Face


license: apache-2.0
base_model:

  • huihui-ai/Huihui-Qwen3.5-27B-Claude-4.6-Opus-abliterated
    base_model_relation: quantized
    library_name: transformers
    pipeline_tag: text-generation
    tags:
  • safetensors
  • qwen3.5
  • 4-bit
  • int4
  • auto-round
  • gptq
  • quantized
  • abliterated
  • dgx-spark
  • dflash
  • vllm
    language:
  • en
  • zh

Huihui-Qwen3.5-27B-Claude-4.6-Opus-abliterated-int4-AutoRound

English | 繁體中文


English

INT4 AutoRound quantization of huihui-ai/Huihui-Qwen3.5-27B-Claude-4.6-Opus-abliterated, optimized for NVIDIA DGX Spark (GB10 SM121) with Marlin INT4 kernel acceleration.

Model Details

Item Value
Architecture Dense 27B + GDN (Mamba) + Attention hybrid
Base model Qwen/Qwen3.5-27B
Fine-tuned by huihui-ai (Claude 4.6 Opus distillation + abliteration)
Quantized by YuYu1015
Model size ~26 GB (vs ~54 GB BF16 original)
Context length Up to 65,536 tokens
Thinking mode Supported (enable_thinking: true/false)
Tool calling Supported (qwen3_coder parser)

Quantization Details

Item Value
Method Intel AutoRound v0.12.2
Bits 4
Group size 128
Format auto_round (GPTQ-compatible)
Iterations 200
Calibration samples 512
Calibration sequence length 2048
Hardware NVIDIA DGX Spark (GB10, 128GB unified memory)

Layers Preserved in BF16

The following layers are not quantized to preserve model quality:

Layer Reason
lm_head Output head, sensitive to quantization noise
embed_tokens Input embeddings (auto-excluded by shape)
linear_attn.* GDN/DeltaNet layers, may output zeros if quantized
model.visual.* Vision encoder (auto-excluded by shape)

Speculative Decoding

DFlash (requires separate drafter model):

--speculative-config '{"method": "dflash", "model": "z-lab/Qwen3.5-27B-DFlash", "num_speculative_tokens": 16}'

Note: The DFlash drafter was trained on the original Qwen3.5-27B. Acceptance rate on the abliterated/distilled variant may be lower than on the original model.

Serving with vLLM

vllm serve /path/to/model \
    --quantization gptq_marlin \
    --served-model-name qwen3.5-27b \
    --reasoning-parser qwen3 \
    --enable-auto-tool-choice \
    --tool-call-parser qwen3_coder \
    --kv-cache-dtype auto \
    --gpu-memory-utilization 0.90 \
    --max-model-len 65536 \
    --enable-prefix-caching \
    --enable-chunked-prefill \
    --trust-remote-code \
    --language-model-only

DGX Spark (SM121) Compatibility Notes

  • Use --quantization gptq_marlin for Marlin INT4 kernel (Dense model, not MoE)
  • FP8 KV cache is not compatible with GDN non-causal attention layers; use --kv-cache-dtype auto
  • NVFP4 is not supported on SM121 (missing cvt.e2m1x2 instruction)
  • Runtime FP8 (--quantization fp8) is not compatible with DFlash
  • --language-model-only skips vision encoder profiling for text-only inference
  • Clear page cache before starting on UMA: sudo sh -c 'echo 3 > /proc/sys/vm/drop_caches'

Safety Warning

This model has safety filtering removed (abliterated) and may generate inappropriate content. Users are solely responsible for all consequences arising from its use.

Credits


繁體中文

huihui-ai/Huihui-Qwen3.5-27B-Claude-4.6-Opus-abliterated 的 INT4 AutoRound 量化版本,針對 NVIDIA DGX Spark (GB10 SM121) 最佳化。

模型資訊

項目 數值
架構 Dense 27B + GDN (Mamba) + Attention 混合
基礎模型 Qwen/Qwen3.5-27B
微調者 huihui-ai(Claude 4.6 Opus 蒸餾 + abliteration)
量化者 YuYu1015
模型大小 ~26 GB(原版 BF16 約 54 GB)
Context 長度 最高 65,536 tokens
思考模式 支援(enable_thinking: true/false)
工具呼叫 支援(qwen3_coder parser)

量化詳情

項目 數值
方法 Intel AutoRound v0.12.2
位元數 4
Group size 128
格式 auto_round(GPTQ 相容)
迭代次數 200
校準樣本數 512
校準序列長度 2048
量化硬體 NVIDIA DGX Spark(GB10, 128GB 統一記憶體)

保留 BF16 的層

以下層未被量化以保持模型品質:

層 原因
lm_head 輸出頭,對量化雜訊敏感
embed_tokens 輸入嵌入(因 shape 自動排除)
linear_attn.* GDN/DeltaNet 層,量化後可能輸出零
model.visual.* 視覺編碼器(因 shape 自動排除)

DGX Spark (SM121) 相容性說明

  • 使用 --quantization gptq_marlin 啟用 Marlin INT4 kernel(Dense 模型,非 MoE)
  • FP8 KV cache 與 GDN non-causal attention 不相容,請使用 --kv-cache-dtype auto
  • NVFP4 在 SM121 上不支援(缺少 cvt.e2m1x2 指令)
  • Runtime FP8(--quantization fp8)與 DFlash 不相容
  • UMA 架構啟動前請先清除 page cache:sudo sh -c 'echo 3 > /proc/sys/vm/drop_caches'

安全警告

此模型已移除安全過濾機制(abliterated),可能產生不當內容。使用者須自行承擔所有風險與法律責任。

致謝

README history 4 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-04-12Update README.md15eb97e6.3 KB
    Loading...
  2. 2026-04-11Update README.md8bd2ebe6.7 KB
    Loading...
  3. 2026-04-11Update README.md7cda5b26.7 KB
    Loading...
  4. 2026-04-11Create README.md3f99a5c6.7 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration