← back to catalog · registered 2026-08-22 13:56

lued/Qwen3.8-27B-huihui-abliterated-INT8-W8A16-MTP

lued Qwen 24B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/lued%2FQwen3.8-27B-huihui-abliterated-INT8-W8A16-MTP"
Response includes
  • classification m1
  • files 19
  • hub_downloads_all_time 6,018
  • author_summary 3 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
6K
1K last 30d - stable
Likes
4
Model age
7w ago
created 2026-08-17
Downloads over time
Now6.5K→from494↑1,209%
1952.5K4.8K7.1K494 on Aug 196.5K on Oct 11AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
vllm safetensors qwen3_5 qwen3.8 abliterated uncensored huihui compressed-tensors w8a16 int8 quantized mtp

Related

Total size
29.4 GB
Files
19
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-18 19:48

Files by quantization

Auxiliary files 19 files 29.5 GB
model-00001-of-00007.safetensors 4.98 GB a891b4f8 download
model-00005-of-00007.safetensors 4.97 GB 65117574 download
model-00002-of-00007.safetensors 4.96 GB b1f26d6f download
model-00003-of-00007.safetensors 4.95 GB e2b19689 download
model-00006-of-00007.safetensors 4.94 GB 5c22f0e5 download
model-00004-of-00007.safetensors 3.67 GB 8bbe7b9c download
model_mtp.safetensors 810 MB 1d8268aa download
model-00007-of-00007.safetensors 188 MB 3dab30e3 download
tokenizer.json 19.1 MB 06b95093 download
model.safetensors.index.json 193 KB 8efa23ca download
config.json 20.3 KB 0c20d293 download
chat_template.jinja 8.74 KB c0c686f9 download
README.md 7.22 KB 53484848 download
.gitattributes 1.53 KB 52373fe2 download
tokenizer_config.json 1.10 KB 6913705f download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
recipe.yaml 269 B 833a3de7 download
generation_config.json 202 B 023756cf download

README current version from Hugging Face


license: apache-2.0
base_model: huihui-ai/Huihui-Qwen3.8-27B-abliterated
library_name: vllm
pipeline_tag: image-text-to-text
tags:

  • qwen3_5
  • qwen3.8
  • abliterated
  • uncensored
  • huihui
  • compressed-tensors
  • w8a16
  • int8
  • quantized
  • vllm
  • mtp
  • speculative-decoding
  • vision
  • conversational
    base_model_relation: quantized

Qwen

Qwen3.8-27B Abliterated · INT8 W8A16 · BF16 MTP

An abliterated, Ampere-optimized quantization for dual RTX 3090 inference.

Base model · Qwen3.8-27B · remove-refusals-with-transformers · vLLM · llm-compressor

Base Abliterated Format W8A16 Weights INT8 Activations FP16 or BF16 Target Ampere License Apache 2.0

[!NOTE]
This is a numerical W8A16 quantization of huihui-ai/Huihui-Qwen3.8-27B-abliterated, an abliterated (uncensored) version of Qwen/Qwen3.8-27B. All model credit belongs to huihui-ai (the abliteration) and Qwen (the base model); refer to the upstream model card for architecture, capabilities, and usage guidance. This repository changes numerics only.

[!TIP]
This build targets vLLM on Ampere-class GPUs where native FP8 tensor-core execution is unavailable. It retains the model's vision tower, thinking controls, native context, and the byte-preserved BF16 MTP head.

Relationship to the sibling quantization

This checkpoint is the abliterated counterpart of lued/Qwen3.8-27B-INT8-W8A16-MTP. The repositories share:

  • the identical W8A16 compressed-tensors recipe (400 packed GEMMs, same BF16 preservation set, group-128 symmetric data-free RTN);
  • the identical native BF16 MTP head (15 tensors, byte-identical to Qwen's original MTP — same sha256 as the sibling release);
  • the identical tokenizer, chat template, processor configs, and bounded-shard layout;
  • the same total checkpoint footprint (29.44 GiB).

The only deviation is the base language-model weights: this checkpoint's LM weights come from huihui-ai's abliterated BF16 model instead of stock Qwen3.8-27B.

Abliteration (inherited from the base)

Per the upstream card, the source is an uncensored version of Qwen/Qwen3.8-27B created with remove-refusals-with-transformers — a crude, proof-of-concept refusal-removal technique that does not use TransformerLens. The first 15 layers were retained without ablation; MTP and visual components were not modified. These are the upstream card's claims, not measurements made here. The W8A16 step touches only stored numerics of the 400 GEMMs and does not change the refusal profile.

Quantization fidelity

Measured on this exact checkpoint against its own BF16 teacher (the huihui source) on 2026-08-17 with tools/measure_kld.py (teacher-forced, full-vocabulary FP32 log-softmax, FP64 accumulation, nats/token; PyTorch 2.11.0+cu128, Transformers 5.10.1, compressed-tensors 0.17.1):

Candidate Mean KLD vs BF16 Top-1 agreement Weight files
This W8 / BF16 A16 0.000705 98.72% 29.44 GiB

The KLD result measures checkpoint-weight drift over 467 teacher-forced positions (fixed 8-prompt short suite). It is not a functional quality score; tool use, coding, multimodal quality, and long-context recall require separate behavioral evaluation.

Short-prompt suite

Prompt Category Tokens Positions Mean KLD Top-1 agreement
factual factual 13 12 0.000415 1.0000
code-fib code 87 86 0.000606 0.9884
physics-uncertainty physics 86 85 0.000826 0.9882
math-train math 58 57 0.000637 1.0000
sql-top5 technical 60 59 0.000894 0.9831
narrative-clock narrative 68 67 0.000784 1.0000
instruction-stack instruction 43 42 0.000589 0.9762
history-industrial history 60 59 0.000606 0.9661
Token-weighted mean 467 0.000705

kld_result.json in this repository is the single source of truth for these numbers.

Checkpoint profile

Property Value
Quantization Data-free symmetric RTN W8A16, group size 128
Runtime format compressed-tensors / pack-quantized
Quantized GEMMs 400: 192 MLP, 64 full-attention, 144 dense GDN projections
Preserved precision BF16 vision tower, lm_head, MTP (15 tensors, byte-identical to Qwen's original MTP), and recurrent GDN gates (96)
Structural audit STRUCTURAL AUDIT PASS (packed set, ignore list, index contract)
MTP audit MTP AUDIT PASS (15 tensors dtype-, shape-, and bit-identical to the source)
Shard layout 7 main shards (5.4 GiB max) + model_mtp.safetensors (849,400,392 bytes)
Runtime vLLM; this repository is not a GGUF checkpoint

MTP preservation

The MTP head is the exact BF16 MTP from the original Qwen3.8-27B — never touched by the abliteration and never re-quantized:

  • 15 tensors in model_mtp.safetensors (849,400,392 bytes).
  • sha256 1d8268aa85ace093a561e3e7b63b9d390dac1cd55a90cd55b5ec509c3c9da9fe — identical to the MTP bytes in the Qwen3.8-27B source, in the AEON source, and in the sibling release lued/Qwen3.8-27B-INT8-W8A16-MTP.
  • All 15 tensors verified bfloat16, shape-equal, and bit-equal to the huihui source (strict audit, not a count check).

Reproduction contract

The immutable source is huihui-ai/Huihui-Qwen3.8-27B-abliterated commit d42ca8978c5a66e92c3446d46e8adfe03ef692ff. The exact selective W8A16 configuration is recorded in recipe.yaml and config.json. A conforming rebuild must take native MTP from that source, reproduce the source-derived packed projection set, and prove all 15 MTP tensors bit-identical before resharing for publication.

Status

Built on 2026-08-17: both release audits PASS, KLD measured, published to
lued/Qwen3.8-27B-huihui-abliterated-INT8-W8A16-MTP.
Serving validated on 2026-08-18: boots through llama-swap from the lued/*
cache entry and serves a smoke request.

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-18Upload README.md with huggingface_hub1bad8297.2 KB
    Loading...
  2. 2026-08-17Upload folder using huggingface_hub6415c377.2 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration