← back to catalog · registered 2026-08-22 13:56

kakrotto/Qwen3.5-27B-ultra-uncensored-heretic-v1-FP8

kakrotto Qwen 24B second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/kakrotto%2FQwen3.5-27B-ultra-uncensored-heretic-v1-FP8"
Response includes
  • classification m3
  • files 16
  • hub_downloads_all_time 130
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · lifetime
130
18 last 30d - stable
Likes
0
Model age
5mo ago
created 2026-04-18
Downloads over time
Now133→from0↑0%
049981460 on Apr 15133 on Oct 11133 on Oct 7AprMayJunJulAugSepOct
Apr 15 → Oct 11 · 65 snapshots · spans 179 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
safetensors qwen3_5 fp8 quantized qwen3.5 base_model:llmfan46/Qwen3.5-27B-ultra-uncensored-heretic-v1 base_model:quantized:llmfan46/Qwen3.5-27B-ultra-uncensored-heretic-v1 license:apache-2.0 region:us

Related

Total size
28.3 GB
Files
16
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-04-20 04:59

Files by quantization

Auxiliary files 16 files 28.3 GB
model-00004-of-00007.safetensors 4.65 GB 392e5c1f download
model-00001-of-00007.safetensors 4.65 GB 4984b4bf download
model-00005-of-00007.safetensors 4.65 GB 0f43b2f3 download
model-00002-of-00007.safetensors 4.65 GB ab1aa31c download
model-00003-of-00007.safetensors 4.62 GB a740c7e5 download
model-00006-of-00007.safetensors 2.73 GB a2b3a211 download
model-00007-of-00007.safetensors 2.37 GB 4c1ebaab download
tokenizer.json 19.1 MB 639e352c download
model.safetensors.index.json 151 KB 629dae9a download
config.json 17.2 KB 0b53011e download
chat_template.jinja 7.57 KB a585dec8 download
README.md 3.39 KB 08c89167 download
.gitattributes 1.92 KB 115f12f0 download
tokenizer_config.json 1.21 KB 08e416fe download
processor_config.json 1.16 KB 33818c7f download
generation_config.json 213 B aa42cce9 download

README current version from Hugging Face


license: apache-2.0
base_model: llmfan46/Qwen3.5-27B-ultra-uncensored-heretic-v1
tags:

  • fp8
  • quantized
  • qwen3.5

Qwen3.5-27B-ultra-uncensored-heretic-v1-FP8

FP8 block-quantized version of llmfan46/Qwen3.5-27B-ultra-uncensored-heretic-v1.

Quantized to match the official Qwen/Qwen3.5-27B-FP8 format exactly.

Quantization Details

  • Method: Fine-grained FP8 quantization with block size of 128
  • Tool: Hugging Face Transformers native FineGrainedFP8Config (on-the-fly quantization during model loading)
  • Format: quant_method: "fp8" (Qwen/DeepSeek native format, NOT compressed-tensors)
  • Weight: FP8 E4M3, static, block_size=(128, 128)
  • Activation: FP8, dynamic per-token
  • Model size: ~29 GB (vs ~55 GB BF16)

Ignored Layers (modules_to_not_convert)

Copied verbatim from the official Qwen/Qwen3.5-27B-FP8 config.json, with MTP entries removed (this heretic variant has no MTP):

  • lm_head
  • model.language_model.embed_tokens
  • All linear_attn.conv1d, linear_attn.in_proj_a, linear_attn.in_proj_b (DeltaNet SSM-specific subparts)
  • All model.visual.* (entire vision tower)

Quantized layers (NOT in ignore list): linear_attn.out_proj, linear_attn.in_proj_qkv, linear_attn.in_proj_z, all self_attn Q/K/V/O projections, all MLP layers.

Quantization Script

from transformers import Qwen3_5ForConditionalGeneration, AutoProcessor, FineGrainedFP8Config
import json, torch

# Load ignore list from Qwen official FP8 config
ref = json.load(open("Qwen3.5-27B-FP8/config.json"))
ref_ignore = ref["quantization_config"]["modules_to_not_convert"]
modules_to_not_convert = [m for m in ref_ignore if not m.startswith("mtp")]

qc = FineGrainedFP8Config(
    activation_scheme="dynamic",
    weight_block_size=(128, 128),
    modules_to_not_convert=modules_to_not_convert,
    dequantize=False,
)

processor = AutoProcessor.from_pretrained(MODEL_DIR)
model = Qwen3_5ForConditionalGeneration.from_pretrained(
    MODEL_DIR,
    dtype=torch.bfloat16,
    device_map="auto",
    max_memory={0: "30GiB", 1: "30GiB"},
    quantization_config=qc,
    low_cpu_mem_usage=True,
)

model.save_pretrained(SAVE_DIR, max_shard_size="5GB", save_original_format=False)
processor.save_pretrained(SAVE_DIR)

Evaluation Results

BF16 baseline vs FP8 quantized, evaluated with lm_eval 0.4.11, vLLM backend, 2 seeds averaged.

Benchmark BF16 FP8 Recovery
GSM8k-Platinum (5-shot) 98.10% 97.89% 99.79%
IFEval inst_strict 92.15% 92.93% 100.85%
IFEval prompt_strict 89.74% 90.58% 100.93%

Generation parameters: temperature=1.0, top_p=0.95, top_k=64, max_gen_toks=16384

Usage

from vllm import LLM
model = LLM("kakrotto/Qwen3.5-27B-ultra-uncensored-heretic-v1-FP8")

Disclaimer

This is an uncensored model. The quantizer (kakrotto) is not responsible for the model's outputs or any misuse. This FP8 quantization preserves the original model's behavior. Please use responsibly.

Attribution

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-04-20Fix quantization method: FineGrainedFP8Config, not llmcompressor model_free_ptq447d1aa3.4 KB
    Loading...
  2. 2026-04-18Update README with evaluation results and attribution95520241.6 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration