← back to catalog · registered 2026-08-22 13:56

lued/Qwen3.8-27B-huihui-abliterated-INT8-W8A16-DFlash2

lued Qwen 26B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/lued%2FQwen3.8-27B-huihui-abliterated-INT8-W8A16-DFlash2"
Response includes
  • classification m1
  • files 18
  • hub_downloads_all_time 1,131
  • author_summary 3 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
1K
395 last 30d - stable
Likes
6
Model age
7w ago
created 2026-08-21
Downloads over time
Now1.3K→from33↑3,788%
04699391.4K33 on Aug 191.3K on Oct 11AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
vllm safetensors qwen3_5 qwen3.8 abliterated uncensored compressed-tensors w8a16 int8 quantized dflash2 speculative-decoding

Related

Total size
27.5 GB
Files
18
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-21 05:59

Files by quantization

Auxiliary files 18 files 27.5 GB
model-00002-of-00006.safetensors 4.98 GB b91e0aa3 download
model-00005-of-00006.safetensors 4.95 GB a671cac2 download
model-00003-of-00006.safetensors 4.94 GB 6e9e4f50 download
model-00001-of-00006.safetensors 4.94 GB d8ffd688 download
model-00004-of-00006.safetensors 4.92 GB 6c0cfd6b download
model-00006-of-00006.safetensors 2.77 GB d7f4646b download
tokenizer.json 19.1 MB 06b95093 download
model.safetensors.index.json 193 KB fa0c42df download
TECHNICAL.md 22.5 KB 48366bb5 download
config.json 20.8 KB 607c894f download
chat_template.jinja 8.74 KB c0c686f9 download
README.md 6.29 KB 01c440ac download
.gitattributes 1.53 KB 52373fe2 download
tokenizer_config.json 1.10 KB 6913705f download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
recipe.yaml 269 B 833a3de7 download
generation_config.json 202 B 023756cf download

README current version from Hugging Face


license: apache-2.0
base_model: huihui-ai/Huihui-Qwen3.8-27B-abliterated
library_name: vllm
pipeline_tag: image-text-to-text
tags:

  • qwen3_5
  • qwen3.8
  • abliterated
  • uncensored
  • compressed-tensors
  • w8a16
  • int8
  • quantized
  • vllm
  • dflash2
  • speculative-decoding
  • vision
  • conversational
    base_model_relation: quantized

Qwen

Huihui-Qwen3.8-27B-abliterated · INT8 W8A16 · DFlash2

Fast, near-lossless Huihui abliterated Qwen3.8-27B for dual RTX 3090s.

Base model · Qwen3.8-27B · vLLM · llm-compressor

Base Huihui abliterated Format W8A16 Weights INT8 Activations FP16 or BF16 Target Ampere License Apache 2.0

[!NOTE]
A numerical INT8 W8A16 quantization of huihui-ai/Huihui-Qwen3.8-27B-abliterated, an abliterated fine-tune of Qwen/Qwen3.8-27B, served with the DFlash2 drafter for speculative decoding. All model credit belongs to huihui-ai and Qwen; this repository changes numerics only. The detailed engineering notes live in TECHNICAL.md.

The short version

Huihui-Qwen3.8-27B-abliterated is a big model. This build makes it fast on consumer hardware:

  • Up to ~117 t/s decode on two RTX 3090s at full 262K context (family baseline)
  • ~1.4x to 1.5x faster than the same model with native MTP4 speculation (family baseline, pending own bench)
  • ~2.2x faster than plain autoregressive decoding (family baseline, pending own bench)
  • 98.5% top-1 agreement with the abliterated BF16 model (measured: mean KLD 0.000694)
  • 28 GiB on disk (6 shards), 2.02 GiB drafter, full 262K context on 2×24 GB

The KLD and audit numbers above are this build's own measured evidence. Decode
throughput figures are the family baseline from the identical sibling build;
this build's own serve run is pending (see Status).

How fast

Measured on this build: KLD 0.000694 nats mean, 98.5% top-1. Decode numbers
below are the family baseline (single stream, same engine, same 262K context,
cold cache). Generation tokens per second.

Prompt tokens Autoregressive MTP4 DFlash2 DFlash2 vs MTP4
128 47 77 117 1.5x
2,048 47 73 103 1.4x
8,192 47 75 102 1.4x

DFlash2 proposes 7 draft tokens per step and the target verifies them all in
one pass. Native MTP proposes 4. More drafts per verification step means more
accepted tokens and less idle GPU time.

How close to the original

BF16 original This build
Size ~56 GiB 28 GiB (6 shards)
Mean KLD 0 0.000694
Top-1 agreement 100% 98.5%

KLD is a fancy way of asking "does the quantized model pick the same next
token as the original?" Lower is closer. A mean of 0.000694 nats means the
INT8 weights and the BF16 weights are nearly interchangeable. Halving the size
is what lets the full 262K context fit on two 3090s.

What you get

  • Full 262,144-token context on 2×RTX 3090 (intended, pending serve)
  • Vision tower, thinking controls, and tool calling intact, same as upstream huihui-ai/Huihui-Qwen3.8-27B-abliterated
  • Native MTP removed: DFlash2 replaces it, so there is no dead weight in the repo
  • Runs on vLLM with the club-3090 patch set (DFlash2 is a new spec-decoder, still in PR review upstream)

Run it

Three steps. The full command with every flag is in TECHNICAL.md.

# 1. Models (pulls into ~/.cache/huggingface)
hf download lued/Qwen3.8-27B-huihui-abliterated-INT8-W8A16-DFlash2
hf download lued/Qwen3.8-27B-DFlash2-W8

# 2. Patches (DFlash2 support is not in a released vLLM yet)
git clone https://github.com/noonghunna/club-3090.git

# 3. Serve (condensed; full command in TECHNICAL.md)
export CLUB3090="$HOME/club-3090"
podman run --rm --replace --device nvidia.com/gpu=all --ipc=host -p 8080:8080 \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  -v "$CLUB3090/models/qwen3.6-27b/vllm/patches/vllm-pr52816-dflash2":/etc/club3090/pr52816:ro \
  -v "$CLUB3090/models/qwen3.6-27b/vllm/patches/vllm-pr48375-mamba-drop-eagle-block":/etc/club3090/pr48375:ro \
  --entrypoint bash docker.io/vllm/vllm-openai:nightly-5a4c8d99242e9e069b604d0e9b969e77f7dd501d \
  -c 'bash /etc/club3090/pr48375/install.sh || exit 1; bash /etc/club3090/pr52816/install.sh || exit 1; exec vllm serve "$@"' -- \
  lued/Qwen3.8-27B-huihui-abliterated-INT8-W8A16-DFlash2 \
  --tensor-parallel-size 2 --max-model-len 262144 --kv-cache-dtype fp8_e4m3 \
  --speculative-config '{"method":"dflash","model":"lued/Qwen3.8-27B-DFlash2-W8","num_speculative_tokens":7}'

Status

Built, audited, and published 2026-08-20 to
lued/Qwen3.8-27B-huihui-abliterated-INT8-W8A16-DFlash2
at commit 26c9c9bc. STRUCTURAL AUDIT PASS (401 packed / 783 preserved,
independently re-verified tensor-by-tensor), DFLASH2 AUDIT PASS (15 MTP
stripped, embed packed), shared drafter audit PASS. KLD measured: mean
0.000694 nats, top-1 0.9850, sanity gate PASS. Registered in the HF cache and
verified to resolve offline. Serve-validation on this rig is pending (family
baseline decode from the identical sibling build). The drafter ships
separately as
lued/Qwen3.8-27B-DFlash2-W8.
All engineering detail, audit evidence, and measured tables are in
TECHNICAL.md.

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-21Upload folder using huggingface_hub5ee1c626.3 KB
    Loading...
  2. 2026-08-21Upload folder using huggingface_hub26c9c9b6 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration