← back to catalog · registered 2026-08-22 13:56

munekazu/Huihui-Qwen3.8-27B-abliterated-FP8

munekazu Qwen 25B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/munekazu%2FHuihui-Qwen3.8-27B-abliterated-FP8"
Response includes
  • classification m1
  • files 19
  • hub_downloads_all_time 7,377
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
7K
5K last 30d - active
Likes
3
Model age
7w ago
created 2026-08-17
Downloads over time
Now8.9K→from910↑876%
5113.6K6.6K9.7K910 on Aug 198.9K on Oct 11AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
transformers safetensors qwen3_5 image-text-to-text fp8 vllm abliterated uncensored qwen3 conversational base_model:huihui-ai/Huihui-Qwen3.8-27B-abliterated base_model:quantized:huihui-ai/Huihui-Qwen3.8-27B-abliterated

Related

Total size
28.7 GB
Files
19
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-09-08 01:08

Files by quantization

Auxiliary files 19 files 28.8 GB
model-00001-of-00006.safetensors 5.18 GB e3d0c7ad download
model-00003-of-00006.safetensors 5.06 GB dcc44c7e download
model-00002-of-00006.safetensors 5.03 GB 976d163a download
model-00004-of-00006.safetensors 5.01 GB 35bfafdd download
model-00005-of-00006.safetensors 5.00 GB e4392c70 download
model-00006-of-00006.safetensors 3.46 GB e58ab790 download
tokenizer.json 12.2 MB 0997f410 download
vocab.json 6.41 MB 0aa0ce06 download
merges.txt 3.20 MB a494e019 download
model.safetensors.index.json 152 KB 6688fe64 download
config.json 50.1 KB e66d0dcb download
tokenizer_config.json 17.5 KB 5de744b3 download
LICENSE 11.3 KB f938136e download
chat_template.jinja 8.74 KB c0c686f9 download
README.md 5.58 KB 205ab583 download
.gitattributes 1.53 KB 52373fe2 download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
generation_config.json 202 B 023756cf download

README current version from Hugging Face


license: apache-2.0
library_name: transformers
base_model:

  • huihui-ai/Huihui-Qwen3.8-27B-abliterated
    base_model_relation: quantized
    pipeline_tag: image-text-to-text
    tags:
  • fp8
  • vllm
  • abliterated
  • uncensored
  • qwen3

Huihui-Qwen3.8-27B-abliterated-FP8

Block-wise FP8 (e4m3) quantization of
huihui-ai/Huihui-Qwen3.8-27B-abliterated,
an abliterated version of Qwen/Qwen3.8-27B.

52 GiB (BF16) → 29 GiB (FP8). The vision tower, all normalization layers,
and the MTP head's own norms/fusion layer stay in BF16, so vLLM's
multimodal path and MTP speculative decoding both keep working.

Quantization details

The layout is identical to the official FP8 release of this model
(Qwen/Qwen3.8-27B-FP8):

Method quant_method: fp8, DeepSeek-style block-wise
Block size weight_block_size: [128, 128]
Scales amax / 448.0, stored as bf16 under <module>.weight_scale_inv
Activation dynamic
Quantized modules 407 (MLP gate/up/down_proj, attention q/k/v/o_proj, linear-attn in_proj_qkv / in_proj_z / out_proj, and the MTP head's self_attn/mlp projections)

Unlike the earlier Qwen3.6 FP8 releases, the official Qwen3.8 FP8 checkpoint
also quantizes part of the MTP head (mtp.layers.0.self_attn.* and
mtp.layers.0.mlp.*) — this release matches that. The quantized-module set
is taken verbatim from the reference checkpoint rather than chosen by hand,
so it tracks whatever the official release does.

Deliberately left in BF16 — quantizing these breaks the model:

  • Gated DeltaNet internals: conv1d, in_proj_a, in_proj_b, A_log, dt_bias, norm
  • All input_layernorm / post_attention_layernorm / q_norm / k_norm
  • lm_head, embed_tokens
  • The entire vision tower (model.visual.*)
  • The MTP head's own norms and mtp.fc (fusion layer)

Verification

The quantizer was validated against Qwen/Qwen3.8-27B-FP8 (the FP8 release
of this model's unablated parent). Per the base model's card, abliteration
left the first 15 layers untouched (MTP and the vision tower were not
modified at all). The verification confirms this precisely:

layer 0 / layer 3 (pre-ablation, incl. MTP)   — 10/10 probes bit-identical to reference
layer 19 / 39 / 59 (post-ablation, self_attn / mlp) — differ, as expected

Tensor names, dtypes and shapes all match the reference (1199 BF16 + 407 FP8
= 1606 tensors).

Usage (vLLM)

vllm serve mashima/Huihui-Qwen3.8-27B-abliterated-FP8 \
  --quantization fp8 \
  --dtype bfloat16 \
  --tensor-parallel-size 2 \
  --max-model-len 262144 \
  --kv-cache-dtype fp8 \
  --trust-remote-code \
  --reasoning-parser qwen3 \
  --speculative-config '{"method":"mtp","num_speculative_tokens":3}'

Tested on vLLM v0.24.0.

Chat template / thinking

Uses the model's bundled chat_template.jinja, which is reasoning_effort-based
(low / medium / high / xhigh) rather than the older enable_thinking
boolean. Pass e.g. "chat_template_kwargs": {"reasoning_effort": "low"} in
requests, or set --default-chat-template-kwargs server-side.

⚠️ Known issue observed on this checkpoint: with --reasoning-parser qwen3,
reasoning_content comes back empty in both streaming and non-streaming
responses, even though the model is actually thinking (verified via the raw
/v1/completions endpoint, which shows a normal <reasoning>...</think>
generation). Feeding the exact same generated text to the parser directly
(outside the server) extracts reasoning/content correctly, so the bug looks
like it's in how vLLM v0.24.0's decode loop drives the parser during
generation, not in the parser's extraction logic itself. Not root-caused
further — only tested on this FP8 checkpoint, so whether it also reproduces
on the unquantized model is unconfirmed. Answer quality is unaffected either
way; you just don't get the reasoning trace back over the API.

Hardware notes

Verified on 2× RTX 3090 (Ampere, sm_86) with TP=2, ~22 GB per card at
--gpu-memory-utilization 0.92 and 262K context with fp8 KV cache.

Ampere has no FP8 tensor cores, so vLLM falls back to the Marlin W8A16
kernel
(weight-only decompression). You get the memory saving but no FP8
compute speedup — expect this warning, which is normal:

Your GPU does not have native support for FP8 computation but FP8 quantization
is being used. Weight-only FP8 compression will be used leveraging the Marlin kernel.

On Ada / Hopper / Blackwell (sm_89+) FP8 runs natively.

MTP speculative decoding works: measured mean acceptance length 2.22 on 2× 3090.

Inherited warnings from the base model

This is a quantization of an abliterated (uncensored) model. The safety
filtering of the base model has been significantly reduced. The warnings from
huihui-ai
apply unchanged:

  • Risk of sensitive or controversial outputs — review generated content carefully
  • Not suitable for all audiences — may be inappropriate for public settings or underage users
  • Legal and ethical responsibility rests with the user — ensure compliance with local law
  • Intended for research, testing and controlled environments, not unsupervised production use
  • No default safety guarantees — this model has not undergone safety optimization

Quantization does not change the behaviour of the base model in this respect.
Credit for the base model and its abliteration goes to
huihui-ai; please support their work there.

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-09-08Re-quantize from updated base (739e3c5, 2026-08-24): abliteration narrowed to...badcb197.1 KB
    Loading...
  2. 2026-08-17Block-wise FP8 (e4m3) quantization, layout identical to official Qwen3.8-27B-FP8a4e43c75.6 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration