← back to catalog · registered 2026-08-22 13:56

munekazu/Huihui-ThinkingCap-Qwen3.6-27B-abliterated-FP8

munekazu Qwen 24B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/munekazu%2FHuihui-ThinkingCap-Qwen3.6-27B-abliterated-FP8"
Response includes
  • classification m1
  • files 21
  • hub_downloads_all_time 550
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
550
36 last 30d - cooling
Likes
1
Model age
2mo ago
created 2026-07-26
Downloads over time
Now568→from480↑18%
476509543577480 on Aug 19568 on Oct 11568 on Oct 9AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
transformers safetensors qwen3_5 image-text-to-text fp8 vllm abliterated uncensored qwen3_6 thinkingcap token-efficient conversational

Related

Total size
29.1 GB
Files
21
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-07-26 06:39

Files by quantization

Auxiliary files 21 files 29.1 GB
model-00004-of-00006.safetensors 5.07 GB 389ef6fa download
model-00002-of-00006.safetensors 5.06 GB 5fdf12ee download
model-00005-of-00006.safetensors 5.06 GB 84345acb download
model-00003-of-00006.safetensors 5.05 GB af622e83 download
model-00001-of-00006.safetensors 5.01 GB 56438aae download
model-00006-of-00006.safetensors 3.86 GB 72d503a6 download
tokenizer.json 19.1 MB 06b95093 download
vocab.json 6.41 MB 0aa0ce06 download
merges.txt 3.20 MB a494e019 download
model.safetensors.index.json 152 KB 6f93a77a download
config.json 37.6 KB 7d80c321 download
LICENSE 11.1 KB 1d5180a4 download
chat_template.jinja 7.58 KB a8755d82 download
README.md 4.78 KB d691688a download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.16 KB 33818c7f download
tokenizer_config.json 1.10 KB 6913705f download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
generation_config.json 213 B c20033f9 download
configuration.json 51.0 B 3a6d4256 download

README current version from Hugging Face


license: apache-2.0
library_name: transformers
base_model:

  • huihui-ai/Huihui-ThinkingCap-Qwen3.6-27B-abliterated
    base_model_relation: quantized
    pipeline_tag: image-text-to-text
    tags:
  • fp8
  • vllm
  • abliterated
  • uncensored
  • qwen3_6
  • thinkingcap
  • token-efficient

Huihui-ThinkingCap-Qwen3.6-27B-abliterated-FP8

Block-wise FP8 (e4m3) quantization of
huihui-ai/Huihui-ThinkingCap-Qwen3.6-27B-abliterated,
which is in turn an abliterated version of
bottlecapai/ThinkingCap-Qwen3.6-27B.

52 GiB (BF16) → 29 GiB (FP8), with the vision tower, the MTP head and all
normalization layers kept in BF16 so that vLLM's multimodal path and MTP
speculative decoding keep working.

Quantization details

The layout is identical to the official FP8 releases of this architecture
(Qwen/Qwen3.6-27B-FP8, bottlecapai/ThinkingCap-Qwen3.6-27B-FP8):

Method quant_method: fp8, DeepSeek-style block-wise
Block size weight_block_size: [128, 128]
Scales amax / 448.0, stored as bf16 under <module>.weight_scale_inv
Activation dynamic
Quantized modules 400 (MLP gate/up/down_proj, attention q/k/v/o_proj, linear-attn in_proj_qkv / in_proj_z / out_proj)

Deliberately left in BF16 — quantizing these breaks the model:

  • Gated DeltaNet internals: conv1d, in_proj_a, in_proj_b, A_log, dt_bias, norm
  • All input_layernorm / post_attention_layernorm / q_norm / k_norm
  • lm_head, embed_tokens
  • The entire vision tower (model.visual.*)
  • The MTP head (mtp.*) — kept unquantized so vLLM's --speculative-config works

The set of quantized modules was taken verbatim from the reference FP8
checkpoint rather than chosen by hand.

Verification

The quantizer was validated against bottlecapai/ThinkingCap-Qwen3.6-27B-FP8
(the FP8 release of this model's unablated parent). For modules that
abliteration does not touch, the output is bit-for-bit identical to the
reference checkpoint:

layers.3.mlp.gate_proj           bit-identical: weight=True  scale=True
layers.3.mlp.up_proj             bit-identical: weight=True  scale=True
layers.3.self_attn.q_proj        bit-identical: weight=True  scale=True
layers.0.linear_attn.in_proj_qkv bit-identical: weight=True  scale=True
layers.0.linear_attn.in_proj_z   bit-identical: weight=True  scale=True
layers.3.mlp.down_proj           differs  (abliteration target)
layers.3.self_attn.o_proj        differs  (abliteration target)
layers.0.linear_attn.out_proj    differs  (abliteration target)

Only the projections abliteration actually modifies differ — exactly as expected.
Tensor names, dtypes and shapes all match the reference (1199 BF16 + 400 FP8 = 1599 tensors).

Usage (vLLM)

vllm serve mashima/Huihui-ThinkingCap-Qwen3.6-27B-abliterated-FP8 \
  --quantization fp8 \
  --dtype bfloat16 \
  --tensor-parallel-size 2 \
  --max-model-len 262144 \
  --kv-cache-dtype fp8 \
  --trust-remote-code \
  --speculative-config '{"method":"mtp","num_speculative_tokens":3}'

Tested on vLLM v0.24.0.

Hardware notes

Verified on 2× RTX 3090 (Ampere, sm_86) with TP=2, ~22 GB per card at
--gpu-memory-utilization 0.92, 262K context with fp8 KV cache.

Ampere has no FP8 tensor cores, so vLLM falls back to the Marlin W8A16 kernel
(weight-only decompression). You get the memory saving but no FP8 compute
speedup — expect this warning, which is normal:

Your GPU does not have native support for FP8 computation but FP8 quantization
is being used. Weight-only FP8 compression will be used leveraging the Marlin kernel.

On Ada / Hopper / Blackwell (sm_89+) FP8 runs natively.

MTP speculative decoding works: measured mean acceptance length 2.42 on 2× 3090.

Inherited warnings from the base model

This is a quantization of an abliterated (uncensored) model. The safety
filtering of the base model has been significantly reduced. The warnings from
huihui-ai
apply unchanged:

  • Risk of sensitive or controversial outputs — review generated content carefully
  • Not suitable for all audiences — may be inappropriate for public settings or underage users
  • Legal and ethical responsibility rests with the user — ensure compliance with local law
  • Intended for research, testing and controlled environments, not unsupervised production use
  • No default safety guarantees — this model has not undergone safety optimization

Quantization does not change the behaviour of the base model in this respect.
Credit for the base model and its abliteration goes to
huihui-ai; please support their work there.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-07-26Block-wise FP8 (e4m3) quantization, layout identical to official FP8 releases7d402cb4.8 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration