← back to catalog · registered 2026-08-25 06:02

batsclamp/Huihui-Qwen3.8-27B-abliterated-FP8-v2

batsclamp Qwen 19B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/batsclamp%2FHuihui-Qwen3.8-27B-abliterated-FP8-v2"
Response includes
  • classification m1
  • files 18
  • hub_downloads_all_time 1,506
  • author_summary 7 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
2K
Likes
0
Model age
6w ago
created 2026-08-24
Downloads over time
Now2.9K→from14↑20,300%
01K2.1K3.1K14 on Aug 262.9K on Oct 11AugSepOct
Aug 26 → Oct 11 · 47 snapshots · spans 46 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
transformers safetensors qwen3_5 image-text-to-text vllm fp8 compressed-tensors abliterated conversational base_model:huihui-ai/Huihui-Qwen3.8-27B-abliterated base_model:quantized:huihui-ai/Huihui-Qwen3.8-27B-abliterated license:apache-2.0

Related

Total size
34.3 GB
Files
18
Quantizations
1
Registered
2026-08-25 06:02
Last updated on HF
2026-08-25 05:19

Files by quantization

Auxiliary files 18 files 34.3 GB
model-00001-of-00003.safetensors 18.6 GB 858bb549 download
model-00002-of-00003.safetensors 14.9 GB 8578d115 download
model-00003-of-00003.safetensors 810 MB 90fa0e3e download
tokenizer.json 19.1 MB 06b95093 download
vocab.json 6.41 MB 0aa0ce06 download
merges.txt 3.20 MB a494e019 download
model.safetensors.index.json 135 KB dd55b506 download
config.json 29.5 KB fb868deb download
LICENSE 11.3 KB f938136e download
chat_template.jinja 8.74 KB c0c686f9 download
README.md 3.25 KB ff4d44aa download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.16 KB 33818c7f download
tokenizer_config.json 1.10 KB d1a20cc3 download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
recipe.yaml 339 B 8e0049db download
generation_config.json 214 B 0bc3addd download

README current version from Hugging Face


license: apache-2.0
base_model: huihui-ai/Huihui-Qwen3.8-27B-abliterated
base_model_relation: quantized
pipeline_tag: image-text-to-text
library_name: transformers
tags:

  • vllm
  • fp8
  • compressed-tensors
  • abliterated
  • qwen3_5

Huihui-Qwen3.8-27B-abliterated-FP8 — v2

⚠️ This is v2. It is quantized from the 2026-08-24 upstream re-release
(739e3c5b), in which huihui-ai narrowed the ablation to layers 18–51. The earlier
upstream build ablated layers 15–63; the narrower range retains more of the original
model's performance. v1 of this repo — the quantization of the older, more heavily
ablated weights — has been deleted and is no longer downloadable.
If you pulled this
repo before 2026-08-24, re-download it.

FP8 quantization of huihui-ai/Huihui-Qwen3.8-27B-abliterated,
produced with llm-compressor (FP8_DYNAMIC).

huihui-ai publish only BF16 weights and a GGUF build for this model — no FP8 — so this
repo fills that gap for vLLM users.

Version v2
Base revision 739e3c5b89849f6c238ce1e5b70008612ae42cdd (2026-08-24)
Ablated layers upstream 18–51

Scheme

Weights FP8 e4m3, per-channel (static)
Activations FP8 e4m3, dynamic per-token
Format compressed-tensors (float-quantized)
Calibration none required (data-free pipeline)

256 dense Linear modules are quantized. Everything the upstream
Qwen/Qwen3.8-27B-FP8 release leaves alone is left in BF16 here too:

  • linear_attn.* — the hybrid Mamba projections (in_proj_qkv, in_proj_a, in_proj_b, in_proj_z, out_proj), 48 layers
  • the whole vision tower (visual.blocks.*, visual.merger.*)
  • embed_tokens, lm_head, and all norms
  • the MTP drafter (mtp.*) — kept in BF16 rather than FP8, so speculative decoding still works

Note the upstream FP8 release uses per-tensor weight scales; this one uses per-channel,
a finer-grained (and therefore more accurate) scheme at the same size.

Serving with vLLM

vllm serve batsclamp/Huihui-Qwen3.8-27B-abliterated-FP8-v2 \
  --max-model-len 262144 \
  --speculative-config '{"method": "mtp", "num_speculative_tokens": 3}' \
  --reasoning-parser qwen3 \
  --tool-call-parser qwen3_coder \
  --override-generation-config '{"temperature": 1.0, "top_p": 0.95, "top_k": 20}'

Sampling defaults follow the upstream model card's thinking-mode recommendation
(temp 1.0 / top_p 0.95 / top_k 20).

Caveats

  • This is an abliterated model: refusal behaviour has been removed upstream. Safety properties are not those of the original Qwen release.
  • Quantization was verified structurally and by generation, not by a benchmark suite; no perplexity or eval numbers are claimed.
  • With reasoning_effort: xhigh this model family will spend a very large output budget inside the reasoning block — measured 16k tokens / 26 min for one hard question on a GB10. It converges and returns a full answer, but if you cap max_tokens below what it needs you get finish_reason: length and an empty content. This is not specific to this quantization or to abliteration — stock Qwen/Qwen3.8-27B-FP8 behaves identically.

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-25v2: mark version in title, document upstream 739e3c5bda587233.3 KB
    Loading...
  2. 2026-08-24FP8 quantization of huihui-ai/Huihui-Qwen3.8-27B-abliterated @ 739e3c5b00235272.8 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration