← back to catalog · registered 2026-08-22 13:56

RobinsonLabs/Qwen3.8-27B-abliterated-GGUF

RobinsonLabs Qwen 27B GGUF multimodal
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/RobinsonLabs%2FQwen3.8-27B-abliterated-GGUF"
Response includes
  • classification m8
  • files 11
  • hub_downloads_all_time 2,535
  • author_summary 15 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
3K
594 last 30d - stable
Likes
1
Model age
7w ago
created 2026-08-19
Downloads over time
Now2.5K→from758↑234%
6691.4K2K2.7K758 on Aug 192.5K on Oct 112.5K on Sep 28AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 2 formats · 607 downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
apache-2.0
Quantizations
IQ2 IQ3 IQ4 Q3_K Q4_K Q5_K Q6_K Q8_0
Tags
region:us

Related

Total size
130 GB
Files
11
Quantizations
10
Registered
2026-08-22 13:56
Last updated on HF
2026-09-26 15:17

Files by quantization

Q8_0 1 file 27.1 GB
Qwen3.8-27B-abliterated-Q8_0.gguf 27.1 GB 8cbea0bc download
Q6_K 1 file 20.9 GB
Qwen3.8-27B-abliterated-Q6_K.gguf 20.9 GB 584c25a0 download
Q5_K 1 file 18.2 GB
Qwen3.8-27B-abliterated-Q5_K_M.gguf 18.2 GB 9dd12ce8 download
Q4_K 1 file 15.7 GB
Qwen3.8-27B-abliterated-Q4_K_M.gguf 15.7 GB e7756895 download
IQ4 1 file 14.3 GB
Qwen3.8-27B-abliterated-IQ4_XS.gguf 14.3 GB 8e57f098 download
Q3_K 1 file 12.7 GB
Qwen3.8-27B-abliterated-Q3_K_M.gguf 12.7 GB f55f4fe9 download
IQ3 1 file 11.4 GB
Qwen3.8-27B-abliterated-IQ3_XS.gguf 11.4 GB 1581cfa1 download
IQ2 1 file 9.59 GB
Qwen3.8-27B-abliterated-IQ2_M.gguf 9.59 GB 495e6bd7 download
F16 1 file 888 MB
mmproj-Qwen3.8-27B-abliterated-f16.gguf 888 MB 5bff8087 download
Auxiliary files 2 files 8.33 KB
README.md 6.21 KB 74141049 download
.gitattributes 2.12 KB 440553a0 download

README current version from Hugging Face


license: apache-2.0
base_model: Qwen/Qwen3.8-27B
library_name: gguf
pipeline_tag: image-text-to-text
tags:

  • abliterated
  • uncensored
  • qwen3.8
  • gguf
  • imatrix
  • mtp
  • vision
  • mmproj
  • not-for-all-audiences

Qwen3.8-27B - Abliterated (GGUF, imatrix)

Imatrix-quantized GGUF ladder of
RobinsonLabs/Qwen3.8-27B-abliterated,
which is an abliterated bf16 base of Qwen/Qwen3.8-27B.

The MTP head is abliterated in-band and the vision tower is preserved -- see the base
repo for the method, the verification, and the measured refusal numbers.

fp16 base

The full-precision master these were cut from is
RobinsonLabs/Qwen3.8-27B-abliterated -- bf16 safetensors, carrying the method, the
verification, and the measured refusal numbers. Go there to re-abliterate, LoRA-merge, fine-tune,
or roll your own quants.

Vision

mmproj-Qwen3.8-27B-abliterated-f16.gguf (334 tensors) is published here. Download it
alongside whichever main quant you pick and pass it with --mmproj to get the vision half.
Without it you have a capable text model and no image input.

llama-server -m Qwen3.8-27B-abliterated-Q4_K_M.gguf \
             --mmproj mmproj-Qwen3.8-27B-abliterated-f16.gguf \
             -ngl 99 -c 32768

Quants

file bits size bpw fits
Q8_0 8 29.05 GB 8.51 2x24GB, or 32GB+
Q6_K 6 22.43 GB 6.57 24GB card, quality ceiling
Q5_K_M 5 19.54 GB 5.72 24GB comfortable
Q4_K_M 4 16.84 GB 4.93 24GB / 16GB with offload -- the volume rung
IQ4_XS 4 15.37 GB 4.50 16GB card
Q3_K_M 3 13.59 GB 3.98 16GB tight
IQ3_XS 3 12.26 GB 3.59 12GB card
IQ2_M 2 10.30 GB 3.02 10-12GB card -- quality-compromised, read the note

Sizes are the built artifacts, exact. bpw is bytes x 8 over the model's own 27,320,697,856
parameters, summed from the master's tensor shapes rather than taken off the "27B" in the name.

Every K/IQ rung is imatrix-guided except for the MTP block, which no imatrix covers -- see below
for what we do about it. Q8_0 uses no imatrix by design; it gains essentially nothing from
importance weighting.

What the low rungs actually cost

Measured rather than asserted. Perplexity over a held-out slice of wikitext-2 -- deliberately NOT
the operator corpus the imatrix was calibrated on, so this is an out-of-distribution read and not a
flattering one. What matters is each rung's distance from the Q8_0 reference on identical text.

rung PPL vs Q8_0
Q8_0 (reference) 5.9283 --
IQ3_XS 6.2212 +4.9%
IQ2_M 6.7226 +13.4%

IQ2_M is a real quality step down and is labelled as such. It stays coherent -- it holds an
argument, follows a format instruction, and reasons correctly about physics in spot checks -- but
if you have the VRAM for IQ3_XS or above, take it. Ship IQ2_M when 10-12 GB is the constraint,
not because it is close to the top of the ladder.

We also built IQ2_XS and did not publish it. It measured 7.4030, or +24.9% against the
reference, to save 0.91 GB over IQ2_M -- roughly three times worse per gigabyte than the step
above it, and past the point where we are willing to put our name on the output. The file exists;
it is not here on purpose.

The imatrix is not the usual one

Most published imatrix quants calibrate on calibration_datav3.txt or similar generic English.
This ladder is calibrated on an in-domain corpus -- 1.8 MB / 17.5K lines of real technical
operator transcripts (infrastructure work, debugging, model-building dialogue), 120 chunks,
final PPL 10.3431 +/- 0.168.

That is a deliberate trade, not an accident. It biases the quantization error toward preserving
behaviour on long technical dialogue, tool use, and operator-style instruction-following. If your
use case is closer to general chat or non-English, a generic-calibrated ladder may serve you
better and that is fine -- we would rather tell you the calibration than let you assume it.

What the imatrix does not cover

Being straight about a limitation, since the calibration is the selling point.

llama-imatrix collects its statistics during a perplexity-style forward pass, and that pass never
runs the model's MTP (multi-token-prediction) draft head. So the imatrix carries entries for
blk.0 through blk.63 -- the 64 trunk layers -- and nothing for blk.64, the MTP block.

llama.cpp handles that two ways, and only one of them tells you: above its "very low-bit"
threshold it quantizes the block blind at the trunk's depth and says nothing, and at IQ3_XS it
refuses and aborts the build. The loud case is the honest one, and it is how we found the quiet
one.

Rather than ship a block with no importance data at low precision, this ladder holds an invariant:

blk.64 is never below q5_K, and never an I-quant, in any rung.

Q8_0, Q6_K and Q5_K_M already satisfied that from the stock mixture and are untouched.
Q4_K_M, IQ4_XS, Q3_K_M and IQ3_XS pin it explicitly with
--tensor-type 'blk\.64\.=q5_K'. IQ4_XS is the one worth calling out: its MTP block would
otherwise be a blind iq4_xs, and I-quants are precisely the family that leans on importance
data. The pin costs roughly 0.1 GB per rung.

If you re-quantize this model yourself, check your imatrix's block coverage against n_layers
before you start, rather than discovering it at the bottom rung.

Disclosure

Abliterated: the hard-refusal reflex on adult / creative content is reduced via single-direction
weight orthogonalization. Harm guardrails are retained by design -- self-harm prompts still
redirect to help (e.g. 988) rather than comply. Not a jailbreak-for-anything model, not intended
to assist genuine wrongdoing. Tagged not-for-all-audiences. See the base repo for the full
disclosure and the measured base-vs-abliterated numbers.

Provenance

Built by Robinson Labs with ModelForge.
Base pinned at 1d4bf0f2. Quantized from our own bf16 abliterated master, not from someone
else's quant -- so the ladder is a single lineage, not a requant chain.

README history 5 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-09-26Withdraw model (2026-09-25)bfb29af240 B
    Loading...
  2. 2026-08-20Card: bidirectional cross-link, exact ladder sizes, MTP imatrix-coverage disc...bd0e0486.2 KB
    Loading...
  3. 2026-08-19Card: bidirectional cross-link, exact ladder sizes, MTP imatrix-coverage disc...07879b05 KB
    Loading...
  4. 2026-08-19WI #2900: align disclosure with house language (self-harm/988), drop class-sp...b4d7e653 KB
    Loading...
  5. 2026-08-19WI #2900: initial model card86ed2113 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration