← back to catalog · registered 2026-08-22 13:56

ayushcluster00/Qwen3.8-27B-Heretic-Abliterated-Uncensored-GGUF-Config

ayushcluster00 Qwen 27B GGUF 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/ayushcluster00%2FQwen3.8-27B-Heretic-Abliterated-Uncensored-GGUF-Config"
Response includes
  • classification m3
  • files 28
  • hub_downloads_all_time 6,424
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · lifetime
6K
1K last 30d - stable
Likes
0
Model age
7w ago
created 2026-08-18
Downloads over time
Now6.7K→from3.5K↑93%
3.3K4.6K5.8K7K3.5K on Aug 196.7K on Oct 11AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Quantizations
BF16 F16 IQ1 IQ2 IQ3 IQ4 Q2_K Q3_K Q4_K Q5_K Q6_K Q8_0
Tags
transformers gguf qwen3_5 image-text-to-text qwen3.8 qwen3.5 heretic abliterated uncensored roleplay imatrix text-generation

Related

Total size
391 GB
Files
28
Quantizations
13
Registered
2026-08-22 13:56
Last updated on HF
2026-08-18 05:42

Files by quantization

BF16 1 file 50.1 GB
RVN-BF16.gguf 50.1 GB fe3cb9c7 download
F16 1 file 50.1 GB
RVN-F16.gguf 50.1 GB ce414558 download
Q8_0 1 file 26.6 GB
RVN-Q8_0.gguf 26.6 GB 638be14a download
Q6_K 1 file 20.6 GB
RVN-Q6_K.gguf 20.6 GB 4c94eac3 download
Q5_K 2 files 35.3 GB
RVN-Q5_K_M.gguf 17.9 GB a55d2ff0 download
RVN-Q5_K_S.gguf 17.4 GB 4fbfdff6 download
Q4_K 2 files 29.9 GB
RVN-Q4_K_M.gguf 15.4 GB 40b5c70b download
RVN-Q4_K_S.gguf 14.5 GB 42b37f8c download
IQ4 2 files 29.0 GB
RVN-IQ4_NL.gguf 14.8 GB c9d68f11 download
RVN-IQ4_XS.gguf 14.2 GB 768895ae download
Q3_K 3 files 37.0 GB
RVN-Q3_K_L.gguf 13.4 GB 5d84279c download
RVN-Q3_K_M.gguf 12.4 GB 11c2c708 download
RVN-Q3_K_S.gguf 11.2 GB 442a1ebd download
IQ3 4 files 44.8 GB
RVN-IQ3_M.gguf 11.7 GB 7ba5d110 download
RVN-IQ3_S.gguf 11.6 GB 4a1857f4 download
RVN-IQ3_XS.gguf 11.1 GB f69cfdad download
RVN-IQ3_XXS.gguf 10.4 GB 517552a9 download
Q2_K 2 files 19.5 GB
RVN-Q2_K.gguf 9.98 GB b5eee0e3 download
RVN-Q2_K_S.gguf 9.54 GB 45850857 download
IQ2 4 files 34.4 GB
RVN-IQ2_M.gguf 9.32 GB fcd60053 download
RVN-IQ2_S.gguf 8.72 GB 40ac8f5c download
RVN-IQ2_XS.gguf 8.47 GB 70e8368d download
RVN-IQ2_XXS.gguf 7.85 GB 727b4940 download
IQ1 2 files 13.8 GB
RVN-IQ1_M.gguf 7.11 GB d7e5c4e1 download
RVN-IQ1_S.gguf 6.66 GB 8469480b download
Auxiliary files 3 files 19.5 KB
README.md 11.7 KB efbfc374 download
.gitattributes 4.11 KB 75960b46 download
config.json 3.67 KB c2bb5cf2 download

README current version from Hugging Face


license: apache-2.0
base_model:

  • Qwen/Qwen3.8-27B
    library_name: transformers
    tags:
  • qwen3.8
  • qwen3.5
  • heretic
  • abliterated
  • uncensored
  • roleplay
  • gguf
  • imatrix
    pipeline_tag: text-generation

Qwen3.8-27B RVN Heretic Abliterated Uncensored (GGUF)

RVN is a double-refined abliterated variant of Qwen3.8-27B, built on top of
trohrbaugh/Qwen3.8-27B-heretic-ara
(an ARA abliteration by Tim Rohrbaugh) and further refined with two additional
full-weight ARA passes
targeting residual refusals. It retains very low behavioral
damage (KL ≈ 0.0085) while reducing harmful-prompt refusals from 3/100 (source) to
0–1/100 in independent measurements.

Note on this repository's history. This repo previously hosted the original
Qwen3.8-27B-Heretic-Q4_K_M.gguf (single-quant release from the earlier
trohrbaugh/Qwen3.8-27B-heretic source). That file is kept as legacy for
download-count continuity and backward compatibility — it is the older abliteration
variant and is superseded by the RVN files below. Prefer the RVN quants for new
deployments.

Not for all audiences. This model has reduced safety guardrails by design. It is
intended for adult audiences (18+) doing research, creative writing, roleplay, and
uncensored generation. Certain guardrails are intentionally left in place; use
responsibly and in accordance with your local laws.


What is ARA?

ARA (Arbitrary-Rank Ablation) is the abliteration technique implemented in
p-e-w/heretic. Traditional directional abliteration
finds a single "refusal direction" in activation space and subtracts it — a one-shot,
low-rank surgery that is simple but can leave residual refusals or damage unrelated
behavior.

ARA instead treats abliteration as a matrix optimization problem. For every target
module (attention out-projection and MLP down-projection), it collects activations on
"good" prompts (harmless requests) and "bad" prompts (harmful requests), then uses an
LBFGS optimizer to rewrite the module's weight matrix so that:

  • Preserve: outputs on good prompts change as little as possible (KL is kept low)
  • Steer: outputs on bad prompts are pulled toward the good-prompt output manifold
    (via k-nearest-neighbor distances), so harmful requests stop triggering the refusal
    circuitry
  • Overcorrect: outputs on bad prompts are additionally pushed away from the
    original bad-prompt outputs, which helps overcome complex, multi-stage refusal
    mechanisms

Because the weight matrix is optimized directly (rather than subtracting a single
direction), ARA is "arbitrary rank" — it can carve out a much richer refusal-removal
subspace while keeping behavioral damage minimal.

Why "Heretic" and "Abliterated"?

These two words describe two layers of the same process:

  • Heretic is the tool: the open-source implementation of ARA (and related
    abliteration methods) used to modify the model. Models produced with it are commonly
    labeled "heretic" in the community.
  • Abliterated is the result: the model's refusal behavior has been surgically
    removed. An abliterated model still knows everything the base model knows, but it no
    longer refuses to answer the categories that were steered away during the process.

So "Heretic Abliterated" means: abliterated using the heretic toolset. RVN goes one
step further — it applies the ARA procedure three times total: once by the original
author (trohrbaugh) to get from base Qwen3.8-27B to -ara, and twice more by us to
get from -ara to RVN, squeezing out the last residual refusals.

Special Thanks

This work would not exist without Tim Rohrbaugh (trohrbaugh), whose
heretic-ara ARA
abliteration of Qwen3.8-27B (refusals 3/100, KL 0.0535) provided the foundation we
refined into RVN. His upstream contributions to the heretic codebase — including the
row-norm preservation feature and Qwen3.5 MoE/DeltaNet hybrid handling — are directly
responsible for making DeltaNet-layer abliteration work at all. Thank you, Tim.

Model Overview

Property Value
Base model Qwen/Qwen3.8-27B
Abliteration source trohrbaugh/Qwen3.8-27B-heretic-ara (ARA, KL 0.0535, refusals 3/100)
RVN refinement 2-pass ARA on top of source → KL 0.0085, refusals 0–1/100
Architecture qwen3_5_text (Qwen3.8 family), Gated DeltaNet hybrid
Parameters 27B total
Hidden size 5120
Layers 64 (16 standard attention + 48 Gated DeltaNet linear attention)
Attention heads 24 · KV heads 4 (GQA) · head_dim 256
Vocab 248,320
Context length 262,144 (262K)
License Apache-2.0 (retained from Qwen3.8-27B)
Format GGUF (llama.cpp), MTP/NextN draft tensors excluded (--no-nextn)

Why RVN?

trohrbaugh/Qwen3.8-27B-heretic-ara is already a strong ARA abliteration, but three
harmful prompts still triggered refusals in our independent evaluation (racism website,
malware, government database hacking). RVN applies two additional full-weight ARA
passes
using the same tight parameter set (start 26, end 56, preserve 0.9432,
steer 0.0009, overcorrect 0.5038, neighbor 10), which:

  • Reduced refusals from 3/100 → 0–1/100 (the only remaining refusal is a
    chemical-weapon WMD prompt — one of the strongest safety-trained categories, and
    intentionally one of the guardrails we left in place)
  • Reduced KL damage from 0.0535 (source) to 0.0085 vs base — a ~6× improvement
    in behavioral preservation
  • Verified independently on two rented GPU machines with prefix-based (real-answer)
    refusal measurement

Refusal evaluation (100 harmful-behaviors prompts, prefix-forced real answers)

Model Refusals KL vs base
Qwen3.8-27B (base) ~99/100 —
trohrbaugh -ara (source) 3/100 0.0535
RVN (this repo) 0–1/100 0.0085

Files & Quantization Spectrum

File Size (GB / GiB) Notes
RVN-F16.gguf 53.81 / 50.11 F16 reference (no NextN/MTP)
RVN-BF16.gguf 53.81 / 50.11 BF16 reference (no NextN/MTP)
RVN-Q8_0.gguf 28.60 / 26.63 Max-quality 8-bit
RVN-Q6_K.gguf 22.08 / 20.57 High-quality 6-bit
RVN-Q5_K_M.gguf 19.23 / 17.91 Balanced 5-bit
RVN-Q5_K_S.gguf 18.68 / 17.40 5-bit small
RVN-Q4_K_M.gguf 16.55 / 15.41 Recommended 4-bit (24 GB VRAM)
Qwen3.8-27B-Heretic-Q4_K_M.gguf 16.55 / 15.41 Legacy (older abliteration variant, kept for download continuity)
RVN-IQ4_NL.gguf 15.89 / 14.80 imatrix 4-bit
RVN-Q4_K_S.gguf 15.59 / 14.52 Small 4-bit
RVN-IQ4_XS.gguf 15.19 / 14.15 imatrix 4-bit extra-small
RVN-Q3_K_L.gguf 14.34 / 13.36 Large 3-bit
RVN-Q3_K_M.gguf 13.30 / 12.39 Compact 3-bit
RVN-IQ3_M.gguf 12.58 / 11.72 imatrix 3-bit — re-uploaded 2026-08-17 (previous file had corrupted tensor data: NaN/Inf scales + zeroed tensors from a bad quantize run; re-quantized from F16 with a fresh imatrix and verified — see note below)
RVN-IQ3_S.gguf 12.42 / 11.57 imatrix 3-bit small
RVN-Q3_K_S.gguf 12.07 / 11.24 Compact 3-bit small
RVN-IQ3_XS.gguf 11.97 / 11.15 imatrix 3-bit extra-small
RVN-IQ3_XXS.gguf 11.19 / 10.42 imatrix 3-bit extra-extra-small
RVN-Q2_K.gguf 10.71 / 9.98 2-bit K-quant
RVN-Q2_K_S.gguf 10.25 / 9.54 2-bit K-quant small
RVN-IQ2_M.gguf 10.00 / 9.32 imatrix 2-bit
RVN-IQ2_S.gguf 9.36 / 8.72 imatrix 2-bit small
RVN-IQ2_XS.gguf 9.09 / 8.47 imatrix 2-bit extreme small
RVN-IQ2_XXS.gguf 8.43 / 7.85 imatrix 2-bit (minimum)
RVN-IQ1_M.gguf 7.63 / 7.11 imatrix 1-bit (experimental)
RVN-IQ1_S.gguf 7.15 / 6.66 imatrix 1-bit (experimental)

imatrix-based quants are produced from the same F16 with an activation importance
matrix computed over wikitext-2-raw (original spectrum, 563 chunks) / tiny_shakespeare
(2026-08-17 re-quant additions: IQ3_M fix + IQ2_S/IQ3_XXS/IQ3_XS/IQ3_S, llama-imatrix, -ngl 99).

Quant → GPU / Memory Guide

GPU / Memory Best quant(s) (full GPU load) Effective ctx @ Q8_0 KV
8 GB (RTX 3050, 4060 Laptop) IQ1_S, IQ1_M; IQ2_XXS partial offload only ~2–4K
12 GB (RTX 3060, 4070) IQ2_M, IQ2_S, IQ2_XS, Q2_K_S; IQ3_XXS (tight) ~8–16K
16 GB (RTX 4080, 4090 Laptop, M3 Max) IQ3_M, IQ3_S, Q3_K_M; IQ4_XS/Q4_K_S/IQ4_NL (tight ctx) ~6–24K
24 GB (RTX 3090, 4090, M4 Max) Q5_K_M, Q5_K_S, Q6_K, Q4_K_M; Q8_0 partial ~16–48K
32 GB (RTX 5090, A6000) Q8_0, Q6_K ~24–64K
64 GB+ (A100 80 GB, RTX PRO 6000, M3/M4 Ultra) F16, BF16 ~64–100K+

Sizes in the file table are the actual file sizes on the Hub (decimal GB / GiB),
pulled from repository metadata. Full GPU load means the whole quant fits in VRAM;
quants whose file size exceeds your VRAM need partial offloading.

2026-08-17 — RVN-IQ3_M incident & fix: the original RVN-IQ3_M.gguf generated only
/ characters on every backend (confirmed by the community and reproduced locally). A
tensor-level audit showed corrupted quantization data — NaN/Inf block scales and fully
zeroed tensors (e.g. token_embd had ~39.6M NaN values) — from a bad quantize run, not a
llama.cpp regression (all other quants from the same F16 dequantize cleanly). The file was
pulled, re-quantized from the F16 with a freshly computed imatrix, generation-tested
("The capital of France is" → Paris, 70+ t/s) and re-uploaded. New quants added the same
day: IQ2_S, IQ3_XXS, IQ3_XS, IQ3_S, Q3_K_L, Q5_K_S.

KV cache math (GQA, 4 KV heads, head_dim 256):
2 × 64 layers × 4 KV heads × 256 head_dim × 2 bytes = 256 KiB/token FP16
→ 16K ctx ≈ 4.2 GB · 32K ctx ≈ 8.4 GB · 64K ctx ≈ 16.8 GB (Q8_0 KV halves this).
A 16 GB card running Q3_K_M (13.30 GB model) + 16K ctx Q8_0 KV fits comfortably;
Q4_K_M (16.55 GB) really needs a 24 GB card.

Rule of thumb: pick the largest quant that leaves ≥ 4 GB for KV cache + compute
buffers. If you only need short replies, drop the quant one notch and get a bigger
context; if you need long context, prioritize KV budget over quant size.

Limitations & Responsible Use

  • Reduced safety guardrails by design. This model is not intended for use in
    applications requiring robust safety filtering, content moderation, or deployment to
    minors.
  • Certain guardrails are intentionally left in place. Abliteration targets refusal
    behavior on general harmful-prompt categories; a small set of hard safety-trained
    categories is deliberately not fully removed. Behavior may vary across domains and
    languages.
  • Not affiliated with or endorsed by Qwen/Alibaba or trohrbaugh.

License & Attribution

Citation

@misc{rohrbaugh2026heretic,
  title={Qwen3.8-27B-heretic-ara: ARA Abliteration of Qwen3.8-27B},
  author={Rohrbaugh, Tim},
  year={2026},
  howpublished={\url{https://huggingface.co/trohrbaugh/Qwen3.8-27B-heretic-ara}}
}

@misc{rvn2026,
  title={RVN: Qwen3.8-27B Heretic Abliterated Uncensored},
  author={0bserverx},
  year={2026},
  howpublished={\url{https://huggingface.co/0bserverx/Qwen3.8-27B-Heretic-Abliterated-Uncensored-GGUF}}
}

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-18Duplicate from 0bserverx/Qwen3.8-27B-Heretic-Abliterated-Uncensored-GGUF2e3f61c11.7 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration