← back to catalog · registered 2026-10-04 19:58

abliterant/Qwen3.8-27B-RANA-abliterated-GGUF

abliterant 27B GGUF multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/abliterant%2FQwen3.8-27B-RANA-abliterated-GGUF"
Response includes
  • classification unknown
  • files 23
  • author_summary 4 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
?
Primary method

Unclassified

No clear signals of an abliteration technique in this model.
Confidence
UNKNOWN
Why this label 1 signal
No classification signals present. This may not be an abliterated model at all - it could be a repackaging, a merge with unrelated goals, or unrelated content that mentions the term.
  • no classification signals present (no abliterated, uncensored, or known producer/method markers)
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
7K
Likes
0
Model age
1w ago
created 2026-09-26

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 4 formats · 7K downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
apache-2.0
Quantizations
BF16 IQ2 IQ3 IQ4 Q2_K Q3_K Q4 Q4_K Q5_K Q6_K Q8_0
Tags
gguf abliterant llama.cpp imatrix abliteration refusal-direction qwen3 vision-language image-text-to-text base_model:abliterant/Qwen3.8-27B-RANA-abliterated base_model:quantized:abliterant/Qwen3.8-27B-RANA-abliterated license:apache-2.0

Related

Total size
231 GB
Files
23
Quantizations
13
Registered
2026-10-04 19:58
Last updated on HF
2026-10-04 19:43

Files by quantization

Q8_0 2 files 29.6 GB
Qwen3.8-27B-RANA-abliterated-Q8_0.gguf 26.6 GB b37f66cb download
mtp-Qwen3.8-27B-RANA-abliterated-Q8_0.gguf 2.95 GB 972e102b download
Q6_K 1 file 20.6 GB
Qwen3.8-27B-RANA-abliterated-Q6_K.gguf 20.6 GB 979fcc35 download
Q5_K 2 files 35.3 GB
Qwen3.8-27B-RANA-abliterated-Q5_K_M.gguf 17.9 GB e0168d24 download
Qwen3.8-27B-RANA-abliterated-Q5_K_S.gguf 17.4 GB 5058d8b9 download
Q4_K 2 files 29.9 GB
Qwen3.8-27B-RANA-abliterated-Q4_K_M.gguf 15.4 GB a4ee9f8a download
Qwen3.8-27B-RANA-abliterated-Q4_K_S.gguf 14.5 GB 176bba4f download
IQ4 2 files 28.8 GB
Qwen3.8-27B-RANA-abliterated-IQ4_NL.gguf 14.7 GB c7d6148d download
Qwen3.8-27B-RANA-abliterated-IQ4_XS.gguf 14.0 GB b9428236 download
Q4 1 file 14.5 GB
Qwen3.8-27B-RANA-abliterated-Q4_0.gguf 14.5 GB 84e64f7e download
Q3_K 2 files 25.7 GB
Qwen3.8-27B-RANA-abliterated-Q3_K_L.gguf 13.4 GB 8fa950db download
Qwen3.8-27B-RANA-abliterated-Q3_K_M.gguf 12.4 GB f3558254 download
IQ3 2 files 22.1 GB
Qwen3.8-27B-RANA-abliterated-IQ3_M.gguf 11.7 GB bf104a12 download
Qwen3.8-27B-RANA-abliterated-IQ3_XXS.gguf 10.4 GB 48b859be download
Q2_K 1 file 9.98 GB
Qwen3.8-27B-RANA-abliterated-Q2_K.gguf 9.98 GB dfc59b1b download
IQ2 1 file 9.32 GB
Qwen3.8-27B-RANA-abliterated-IQ2_M.gguf 9.32 GB 90457087 download
BF16 2 files 6.40 GB
mtp-Qwen3.8-27B-RANA-abliterated-BF16.gguf 5.54 GB 7461fa87 download
mmproj-Qwen3.8-27B-RANA-abliterated-BF16.gguf 888 MB d246b854 download
F16 1 file 885 MB
mmproj-Qwen3.8-27B-RANA-abliterated-F16.gguf 885 MB aa995ddf download
Auxiliary files 4 files 13.0 MB
Qwen3.8-27B-RANA-abliterated-imatrix.gguf 13.0 MB a2934b5d download
README.md 13.2 KB e1f0ea33 download
.gitattributes 3.23 KB 23ce184b download
SHA256SUMS 2.40 KB 10bcc2ec download

README current version from Hugging Face


license: apache-2.0
base_model: abliterant/Qwen3.8-27B-RANA-abliterated
base_model_relation: quantized
library_name: gguf
pipeline_tag: image-text-to-text
tags:

  • abliterant
  • gguf
  • llama.cpp
  • imatrix
  • abliteration
  • refusal-direction
  • qwen3
  • vision-language

Abliterant

Qwen3.8-27B-RANA-abliterated-GGUF

Abliterant GGUF quantizations of abliterant/Qwen3.8-27B-RANA-abliterated,
a refusal-ablated Qwen/Qwen3.8-27B, for llama.cpp and
compatible apps. Vision (mmproj-*) and the MTP speculative-decoding head (mtp-*) are included as
separate files, in the same layout as ggml-org/Qwen3.8-27B-GGUF.

Quick start · Evaluation · Available files · Abliterant models

This is a safety-alignment-removed research model. Read Intended use and
Limitations before using it. Method, full evaluation and release gates are on the
BF16 card.

At a glance

Field Value
Base checkpoint abliterant/Qwen3.8-27B-RANA-abliterated, derived from Qwen/Qwen3.8-27B
Release 27B refusal-ablated model; importance-matrix GGUF quantization
Weight formats 15 quantizations from Q8_0 to IQ2_M; split BF16 reference; separate vision and MTP assets
Runtime llama.cpp commit 4b1a27f; recorded CUDA runs on RTX PRO 6000 Blackwell
Context Recorded serving example: 32,768 tokens; native context: 262k (not established here as a validated GGUF limit)
License Apache-2.0

Release family

Format Repo
BF16 (reference) abliterant/Qwen3.8-27B-RANA-abliterated
FP8 (vLLM / SGLang) abliterant/Qwen3.8-27B-RANA-abliterated-FP8
GGUF (this repo) abliterant/Qwen3.8-27B-RANA-abliterated-GGUF
MLX (Apple Silicon) abliterant/Qwen3.8-27B-RANA-abliterated-MLX

Quick start

The release tests used llama.cpp at commit 4b1a27f (CUDA, RTX PRO 6000 Blackwell). Use a build of that revision with llama-server on your PATH.

Download all three files before running the local-file command:

hf download abliterant/Qwen3.8-27B-RANA-abliterated-GGUF \
  Qwen3.8-27B-RANA-abliterated-Q4_K_M.gguf \
  mmproj-Qwen3.8-27B-RANA-abliterated-BF16.gguf \
  mtp-Qwen3.8-27B-RANA-abliterated-Q8_0.gguf --local-dir .

Run from the directory containing those downloads:

llama-server -m Qwen3.8-27B-RANA-abliterated-Q4_K_M.gguf \
  --mmproj mmproj-Qwen3.8-27B-RANA-abliterated-BF16.gguf \
  -md mtp-Qwen3.8-27B-RANA-abliterated-Q8_0.gguf --spec-type draft-mtp \
  --jinja -fa on -ngl 99 -c 32768 \
  --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0

Alternatively, let llama.cpp fetch the files remotely. The release also tested automatic fetching: the mmproj-* file is picked up automatically; the mtp-* file only when --spec-type draft-mtp is given.

llama-server -hf abliterant/Qwen3.8-27B-RANA-abliterated-GGUF:Q4_K_M --spec-type draft-mtp \
  --jinja -fa on -ngl 99 -c 32768 --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0
  • Thinking is on by default. The sampling flags above are Qwen's recommended thinking-mode settings;
    llama.cpp's own default min_p is 0.05, so set --min-p 0 explicitly.
  • Long technical requests can need 20–50k tokens of reasoning (see the
    FP8 card);
    raise -c accordingly (native context 262k).

What changed

  • Source: the published BF16 repo, converted with llama.cpp's convert_hf_to_gguf.py (commit
    4b1a27f): main model with --no-mtp, MTP head with --mtp, vision tower with --mmproj. Tensor
    counts and sizes match ggml-org's conversion of the base model (851 / 53.8 GB, 18 / 5.9 GB,
    334 / 0.93 GB).
  • Importance matrix: llama-imatrix on the BF16 GGUF with bartowski's public calibration text
    (Qwen3.8-27B-calibration-v6.txt from bartowski/Qwen3.8-27B-GGUF).
    General text, no refusal-related prompts.
  • Quantization: llama-quantize --imatrix with llama.cpp's default tensor layouts for each type
    (no per-tensor overrides).

The refusal-abliteration method, full BF16 evaluation and release gates are documented on the BF16 release card. This GGUF conversion and quantization are not additional training.

Evaluation

Functional and refusal checks cover Q4_K_M only; all 15 quantizations have distribution-closeness measurements, not a full downstream evaluation.

Quantization quality

BF16 GGUF perplexity on the same text: 6.776 ± 0.104 (base Qwen3.8-27B in bartowski's
table: 6.744 ± 0.103). The table puts each quant next to bartowski's quant of the base model, measured with the
same protocol (his perplexity.md, llama.cpp b10896). It is a different model, so the comparison is indicative only.

quant this repo: size KLD bartowski (base model): size KLD
Q8_0 28.60 GB 0.0009 29.12 GB 0.0009
Q6_K 22.08 GB 0.0022 23.86 GB 0.0036
Q5_K_M 19.23 GB 0.0064 20.92 GB 0.0053
Q5_K_S 18.68 GB 0.0075 19.57 GB 0.0060
Q4_K_M 16.55 GB 0.0156 17.44 GB 0.0139
Q4_K_S 15.59 GB 0.0188 16.36 GB 0.0156
IQ4_XS 15.08 GB 0.0186 15.48 GB 0.0188
IQ4_NL 15.80 GB 0.0182 17.44 GB 0.0150
Q4_0 15.52 GB 0.0299 16.35 GB 0.0265
Q3_K_L 14.34 GB 0.0504 14.12 GB 0.0432
Q3_K_M 13.30 GB 0.0549 13.40 GB 0.0564
IQ3_M 12.58 GB 0.0628 14.86 GB 0.0406
IQ3_XXS 11.19 GB 0.0977 12.32 GB 0.0739
Q2_K 10.71 GB 0.1516 10.82 GB 0.1612
IQ2_M 10.00 GB 0.1723 10.52 GB 0.1494

Most of his files are larger at the same name (by up to 2.3 GB) because he overrides the type of some
tensors; these quants use llama.cpp's default layouts. At equal size the two are close: this repo's
Q4_K_M (16.55 GB) has the same KLD as his Q4_K_S (16.36 GB), 0.0156, and this IQ4_XS has a slightly
lower KLD than his at 0.4 GB smaller.

Functional checks (Q4_K_M, llama-server)

Served with llama-server, Q4_K_M + mmproj-…-BF16 + mtp-…-Q8_0:

  • Vision: reads the code word and shape from a synthetic image: pass.
  • Tool calling: 3-turn call → result → second call with a new argument: pass.
  • MTP speculative decoding: 348 of 486 drafted tokens accepted (71.6%) on one
    510-token generation at temperature 0: pass.
  • The same three checks pass when the files are fetched with -hf … --spec-type draft-mtp.

Refusals (Q4_K_M, seed 1)

Same prompts, seed and sampling as the BF16 and FP8 builds (refusal suite v2, seed 1, thinking on,
16k-token budget), judged by openai/gpt-oss-safeguard-20b. The GGUF was served with llama.cpp, the
other two with vLLM, so part of any difference can come from the engine.

set build hard soft answers budget hits avg. tokens
HarmBench (200) BF16 0 5 78.5% 21.0% 6,596
HarmBench (200) FP8 0 11 77.0% 20.5% 6,599
HarmBench (200) GGUF Q4_K_M 0 8 79.5% 18.0% 6,105
Held-out (240) BF16 0 9 90.0% 6.7% 4,740
Held-out (240) FP8 0 17 86.2% 7.5% 4,847
Held-out (240) GGUF Q4_K_M 1 14 88.8% 5.4% 4,163

HarmBench labels follow the rule used for every build: each gpt-oss HARD_REFUSAL is re-judged (BF16 card, G1 re-adjudication disclosure). Q4_K_M and BF16 had none in seed 1; FP8's 3 raw ones re-judged as 2 answers and 1
degenerate, shown above. Paired with BF16 on the same prompts (exact McNemar): answers on HarmBench
13 Q4_K_M-only vs 11 BF16-only (p = 0.84), held-out 10 vs 13 (p = 0.68); budget
hits on HarmBench 5 vs 11 (p = 0.21), held-out 4 vs 7 (p = 0.55). None of the differences
is significant. The one held-out hard refusal is a raw label on a finished answer that contains no
refusal phrase; it was not re-judged.

Available files

KLD = mean KL divergence of each quant's next-token distribution from the BF16 GGUF; "same top token"
= how often both pick the same most likely token. Measured with llama-perplexity on wiki.test.raw,
100 chunks of 512 tokens. Lower KLD is closer to BF16.

File Size KLD 99th pct KLD Same top token Notes
Qwen3.8-27B-RANA-abliterated-Q8_0.gguf 28.60 GB 0.0009 0.007 98.7% near-lossless
Qwen3.8-27B-RANA-abliterated-Q6_K.gguf 22.08 GB 0.0022 0.019 97.9% near-lossless
Qwen3.8-27B-RANA-abliterated-Q5_K_M.gguf 19.23 GB 0.0064 0.061 96.6%
Qwen3.8-27B-RANA-abliterated-Q5_K_S.gguf 18.68 GB 0.0075 0.069 96.3%
Qwen3.8-27B-RANA-abliterated-Q4_K_M.gguf 16.55 GB 0.0156 0.146 94.5% functional and refusal tests run on this file
Qwen3.8-27B-RANA-abliterated-Q4_K_S.gguf 15.59 GB 0.0188 0.177 93.9%
Qwen3.8-27B-RANA-abliterated-IQ4_XS.gguf 15.08 GB 0.0186 0.177 94.1%
Qwen3.8-27B-RANA-abliterated-IQ4_NL.gguf 15.80 GB 0.0182 0.178 94.1%
Qwen3.8-27B-RANA-abliterated-Q4_0.gguf 15.52 GB 0.0299 0.306 92.6%
Qwen3.8-27B-RANA-abliterated-Q3_K_L.gguf 14.34 GB 0.0504 0.493 90.4%
Qwen3.8-27B-RANA-abliterated-Q3_K_M.gguf 13.30 GB 0.0549 0.546 90.0%
Qwen3.8-27B-RANA-abliterated-IQ3_M.gguf 12.58 GB 0.0628 0.592 89.4%
Qwen3.8-27B-RANA-abliterated-IQ3_XXS.gguf 11.19 GB 0.0977 0.875 86.5%
Qwen3.8-27B-RANA-abliterated-Q2_K.gguf 10.71 GB 0.1516 1.473 83.2%
Qwen3.8-27B-RANA-abliterated-IQ2_M.gguf 10.00 GB 0.1723 1.565 82.1%

Also in this repo:

  • mmproj-Qwen3.8-27B-RANA-abliterated-{BF16,F16}.gguf (0.93 GB): the vision tower. Same weights as
    the base model's (abliteration does not touch it).
  • mtp-Qwen3.8-27B-RANA-abliterated-{Q8_0,BF16}.gguf (3.2 / 5.9 GB): the MTP head as a speculative
    draft for --spec-type draft-mtp. It is the abliterated MTP head, consistent with the main model.
  • Qwen3.8-27B-RANA-abliterated-BF16/ (2 parts, 53.8 GB): unquantized GGUF, the KLD reference.
  • Qwen3.8-27B-RANA-abliterated-imatrix.gguf: the importance matrix used for every quant.
  • SHA256SUMS, results/: checksums and the measurements behind this card.

Which one? Q8_0 and Q6_K are near-lossless. Q5_K_M and Q4_K_M are the usual choices when memory
is tight. Below 4 bits the KLD rises quickly; IQ3/Q3 and IQ2/Q2 are for fitting into 12–16 GB, with a
visible quality cost.

Limitations and intended use

Limitations

  • Functional checks and refusal behavior were evaluated on Q4_K_M only, one seed and one judge. Other quants were
    checked for closeness to BF16 (KLD) but not for refusal behaviour; the lowest-bit quants drift the
    most from BF16 and may behave differently.
  • KLD is measured on English Wikipedia text at 512-token context. It says how closely a quant
    tracks BF16, not how it scores on downstream tasks.
  • Everything listed under Limitations on the BF16 card
    applies here too: judge-measured refusal rates, long reasoning on technical requests, and the
    capability changes measured there.

Reduced refusal does not establish greater safety, accuracy, factual reliability, or universal compliance. The llama.cpp/vLLM engine difference prevents attributing every comparison difference to quantization alone.

Intended use

  • Research only: interpretability, red-teaming, and robustness evaluation of refusal behaviour.
  • Not for public or end-user deployment without a separate moderation layer. The model's own
    refusals have been largely removed, so any safety filtering has to happen outside it.
  • You are responsible for complying with applicable law, the Apache-2.0 license inherited from Qwen,
    and the terms of any platform where outputs are used.

Provenance and license

  • Qwen team: base model Qwen/Qwen3.8-27B.
  • Arditi et al., 2024: "Refusal in Language Models Is Mediated by a Single Direction".
  • Jim Lai (grimjim): prior work on norm-preserving abliteration.
  • llama.cpp / ggml-org: conversion, quantization and inference; bartowski: the calibration text
    and the public KLD table used for comparison.
  • Benchmarks/datasets: HarmBench, StrongREJECT, JailbreakBench, CategoricalHarmfulQA; wikitext-2.

The artifacts retain the Apache-2.0 license inherited from Qwen; see the Apache-2.0 terms and the upstream model. Original release preparation and measurements are credited to preemware, alongside the contributors above.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration