← back to catalog · registered 2026-08-22 13:56

Yirasumi/Huihui-Qwen3.8-27B-abliterated-INT4-W4A16

Yirasumi Qwen 26B second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/Yirasumi%2FHuihui-Qwen3.8-27B-abliterated-INT4-W4A16"
Response includes
  • classification m1
  • files 16
  • hub_downloads_all_time 1,598
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
2K
414 last 30d - stable
Likes
0
Model age
7w ago
created 2026-08-22
Downloads over time
Now1.7K→from0↑0%
06391.3K1.9K0 on Aug 191.7K on Oct 11AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
safetensors qwen3_5 abliterated uncensored compressed-tensors int4 w4a16 vllm quantized base_model:huihui-ai/Huihui-Qwen3.8-27B-abliterated base_model:quantized:huihui-ai/Huihui-Qwen3.8-27B-abliterated license:apache-2.0

Related

Total size
16.4 GB
Files
16
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-22 07:02

Files by quantization

Auxiliary files 16 files 16.4 GB
model-00001-of-00004.safetensors 4.66 GB a10df850 download
model-00002-of-00004.safetensors 4.66 GB 10397fff download
model-00003-of-00004.safetensors 4.63 GB 7e743385 download
model-00004-of-00004.safetensors 1.62 GB 693c6f7b download
model_extra_tensors.safetensors 810 MB 1d8268aa download
tokenizer.json 19.1 MB 06b95093 download
model.safetensors.index.json 194 KB 146b44b5 download
config.json 20.9 KB 5358558a download
quantization_config.json 15.7 KB 4c7c2843 download
chat_template.jinja 8.74 KB c0c686f9 download
README.md 3.73 KB 21b0cbf9 download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.16 KB 33818c7f download
tokenizer_config.json 1.14 KB 1d134cd2 download
preprocessor_config.json 390 B 2ea84a43 download
generation_config.json 214 B 3f9de11a download

README current version from Hugging Face


license: apache-2.0
tags:

  • qwen3_5
  • abliterated
  • uncensored
  • compressed-tensors
  • int4
  • w4a16
  • vllm
  • quantized
    base_model: huihui-ai/Huihui-Qwen3.8-27B-abliterated

Huihui-Qwen3.8-27B-abliterated-INT4-W4A16

INT4 W4A16 quantization of huihui-ai/Huihui-Qwen3.8-27B-abliterated, produced with Intel AutoRound and exported in compressed-tensors (pack-quantized) format.

Key facts

Property Value
Base model huihui-ai/Huihui-Qwen3.8-27B-abliterated
Quantization INT4 weight-only (W4A16), group_size=128, symmetric
Algorithm AutoRound (Intel)
Format compressed-tensors / pack-quantized (auto-detected by vLLM)
lm_head Quantized to INT4
Weight size ~15.6 GiB (4 shards)
Architecture Qwen3_5ForConditionalGeneration (multimodal, vision tower kept BF16)
Context 262,144 native
License Apache-2.0

Why this quant

This checkpoint is specifically sized to run fast on a single 24 GB GPU (e.g. RTX 3090):

  • Quantizing lm_head to INT4 frees ~1.9 GB of VRAM compared to BF16-lm_head quants.
  • That headroom is what enables speculative decoding (MTP / DFlash2) and CUDA graphs on 24 GB cards.
  • Weight-only INT4 keeps activations in BF16 → minimal quality loss.

The language-model transformer layers are quantized; the vision tower, linear_attn.in_proj_a/b and the MTP fusion layer (mtp.fc) remain in BF16.

Usage with vLLM

vllm serve Yirasumi/Huihui-Qwen3.8-27B-abliterated-INT4-W4A16 \
  --served-model-name qwen3.8-27b \
  --max-model-len 40000 \
  --gpu-memory-utilization 0.93 \
  --kv-cache-dtype fp8_e5m2 \
  --reasoning-parser qwen3 \
  --enable-auto-tool-choice --tool-call-parser qwen3_xml \
  --enable-prefix-caching --enable-chunked-prefill

MTP speculative decoding (the checkpoint ships the Qwen3.5 MTP head in model_extra_tensors.safetensors):

vllm serve Yirasumi/Huihui-Qwen3.8-27B-abliterated-INT4-W4A16 \
  --speculative-config '{"method":"mtp","num_speculative_tokens":4}' \
  # ... same flags as above

On a 24 GB RTX 3090 the tuned serving stack from syv-ai/qwen38-27b-rtx3090 applies directly — see that repo's prepare/ scripts for the optional lm_head/embed requantization and draft-vocab steps.

Quantization details

  • Toolchain: AutoRound 0.14.2, transformers 5.15.x, torch 2.13 (cu13)
  • Calibration: NeelNanda/pile-10k, 128 samples × 2048 tokens, 200 iters
  • Scheme: W4A16, group_size=128, symmetric (auto_round:llm_compressor export)
  • Hardware: NVIDIA H100 80 GB, ~55 minutes total

Verification

  • model.safetensors.index.json + 4 shards + model_extra_tensors.safetensors (MTP head) all present
  • Loads cleanly with vLLM 0.27.1 via the compressed-tensors path (Marlin INT4 kernels)
  • quant_method: compressed-tensors, format: pack-quantized

Credits

Usage warning: This is an abliterated model. Safety filtering has been significantly reduced. Use responsibly and in compliance with applicable laws.

See also

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-22Update README.md58b7a8d3.7 KB
    Loading...
  2. 2026-08-22Upload README.md with huggingface_hub940a2fb3.9 KB
    Loading...
  3. 2026-08-22initial commit21fcf4628 B
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration