← back to catalog · registered 2026-08-22 13:56

RobinsonLabs/Qwen3.5-REAP-212B-A17B-abliterated-GGUF

RobinsonLabs Qwen 212B GGUF MoE
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/RobinsonLabs%2FQwen3.5-REAP-212B-A17B-abliterated-GGUF"
Response includes
  • classification m8
  • files 12
  • hub_downloads_all_time 7,291
  • author_summary 15 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
7K
3K last 30d - stable
Likes
1
Model age
3mo ago
created 2026-07-09
Downloads over time
Now7.3K→from3.9K↑88%
3.7K5K6.3K7.6K3.9K on Jul 157.3K on Oct 117.3K on Sep 26JulAugSepOct
Jul 15 → Oct 11 · 53 snapshots · spans 88 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 2 formats · 3K downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
apache-2.0
Quantizations
IQ2 IQ3 IQ4 Q4_K Q5_K
Tags
region:us

Related

Total size
771 GB
Files
12
Quantizations
6
Registered
2026-08-22 13:56
Last updated on HF
2026-09-26 15:17

Files by quantization

Q5_K 1 file 140 GB
Qwen3.5-REAP-212B-A17B-abl-Q5_K_M.gguf 140 GB e70ec098 download
Q4_K 1 file 120 GB
Qwen3.5-REAP-212B-A17B-abl-Q4_K_M.gguf 120 GB cc204cb1 download
IQ4 1 file 105 GB
Qwen3.5-REAP-212B-A17B-abl-IQ4_XS.gguf 105 GB 0c3db666 download
IQ3 2 files 167 GB
Qwen3.5-REAP-212B-A17B-abl-IQ3_M.gguf 86.5 GB d30c1807 download
Qwen3.5-REAP-212B-A17B-abl-IQ3_XS.gguf 81.0 GB 115cab5f download
IQ2 3 files 175 GB
Qwen3.5-REAP-212B-A17B-abl-IQ2_M.gguf 64.8 GB 3468cd17 download
Qwen3.5-REAP-212B-A17B-abl-IQ2_XS.gguf 58.3 GB 6cd96e3a download
Qwen3.5-REAP-212B-A17B-abl-IQ2_XXS.gguf 52.4 GB b3e539f1 download
Auxiliary files 4 files 62.8 GB
Qwen3.5-REAP-212B-A17B-abl-fit1m-imat.gguf 62.8 GB abc3bdd1 download
bpw-vs-size.png 78.3 KB 88fd3b84 download
README.md 4.38 KB 9a7bdb86 download
.gitattributes 2.15 KB 92014c6b download

README current version from Hugging Face


license: apache-2.0
base_model: OpenMOSE/Qwen3.5-REAP-212B-A17B
library_name: gguf
pipeline_tag: text-generation
tags:

  • gguf
  • abliterated
  • uncensored
  • imatrix
  • moe
  • reap
  • qwen3.5
  • not-for-all-audiences

Qwen3.5-REAP-212B-A17B - Abliterated GGUF

Abliterated, importance-matrix (imatrix) quantized GGUFs of
OpenMOSE/Qwen3.5-REAP-212B-A17B —
itself a 48% REAP expert-pruning of Qwen3.5-397B-A17B down to
212B total / ~17B active.

Provenance chain: the abliterated bf16 safetensors base was converted to a Q8_0 master
(near-lossless), and every rung here is cut from that master. The bf16 base lives at
RobinsonLabs/Qwen3.5-REAP-212B-A17B-abliterated —
use that repo if you want to re-abliterate, merge a LoRA, fine-tune, or roll your own quants.

Disclosure

This model is abliterated — the hard-refusal reflex on adult / creative content has been
reduced via single-direction weight orthogonalization. Harm guardrails are retained by design:
self-harm prompts still redirect to help (e.g. 988), and it is not intended to assist genuine
wrongdoing. Capability is preserved. Tagged not-for-all-audiences. Use responsibly — you are
responsible for what you generate with it. License inherited from the base model: Apache-2.0.

Files

File Quant bpw ~Size Notes
...-Q5_K_M.gguf Q5_K_M 5.68 ~150 GB highest-fidelity rung published
...-Q4_K_M.gguf Q4_K_M 4.85 ~128 GB K-quant quality pick
...-IQ4_XS.gguf IQ4_XS 4.27 ~113 GB quality/size sweet spot
...-IQ3_M.gguf IQ3_M 3.51 ~93 GB
...-IQ3_XS.gguf IQ3_XS 3.29 ~87 GB
...-IQ2_M.gguf IQ2_M 2.63 ~70 GB
...-fit1m-imat.gguf fit1m (mix) 2.55 ~67 GB cluster-fit attention-priority mix — attn/router/shared-expert/embed kept high, routed experts low; punches above its bpw
...-IQ2_XS.gguf IQ2_XS 2.36 ~63 GB
...-IQ2_XXS.gguf IQ2_XXS 2.13 ~56 GB smallest

All quants are imatrix-weighted (corpus-rldomain, computed on the source model).

Quant ladder — bits-per-weight vs file size

The chart shows the contested 2–5 bpw band; the higher-fidelity Q5_K_M rung is in the table above.

Architecture notes

qwen35moe hybrid: 60 decoder layers (45 linear-attn / DeltaNet + 15 full-attn,
full_attention_interval=4), 267 experts, MoE, hidden size 4096. The REAP convert's phantom MTP
layer is corrected (block_count 61→60, nextn_predict_layers 1→0) so the model loads as a clean
60-layer decoder. This is the text path only (no vision mmproj).

Method

  • Abliteration — single-direction weight orthogonalization (FailSpy / Labonne method): for every
    matrix that writes the residual stream (o_proj, DeltaNet out_proj, fused expert down_proj,
    shared-expert down_proj, and the token embedding), the rank-1 component along the refusal
    direction is subtracted. 181 tensors edited; routers, MTP block, and norms pass through
    byte-identical.
  • Refusal direction — massive-activation guarded. The direction is captured with a
    mean-difference control vector, then guarded against attention-sink contamination: the sink
    dimensions that dominate raw activation magnitude (and would brick the model if ablated) are
    detected across layers and excluded, and the direction is taken from the clean, spread-out
    consensus of the late layers rather than a single sink-dominated layer.
  • Quant — importance-matrix (imatrix) weighted convert + quantize with
    llama.cpp.

bf16 base

The full-precision bf16 safetensors master this ladder derives from is at
RobinsonLabs/Qwen3.5-REAP-212B-A17B-abliterated.
That repo is the master for further surgery (re-abliteration, LoRA merge, fine-tune) and for making
your own quants.

Provenance

Qwen3.5-397B-A17B (Apache-2.0) → OpenMOSE/Qwen3.5-REAP-212B-A17B (48% REAP prune) →
abliterated (bf16 master) → Q8_0 master → imatrix quants. Recipe and diagnosis are Robinson
Labs internal (WI #1423, KB #213 "abliteration massive-activation brick").

README history 5 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-09-26Withdraw model (2026-09-25)59c6f5b240 B
    Loading...
  2. 2026-09-14Update model card from hf/cards/ (repo is source of truth)a18b7c24.3 KB
    Loading...
  3. 2026-07-18Fix provenance: quants are cut from the Q8_0 master, not directly from bf16 (...0b2b3004.3 KB
    Loading...
  4. 2026-07-11Complete model card / chart (README.md)b4c2fcf4.2 KB
    Loading...
  5. 2026-07-09Add files using upload-large-folder toole8812d72.5 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration