← back to catalog · registered 2026-08-22 13:56

RobinsonLabs/Qwen3.5-REAP-262B-A17B-abliterated-GGUF

RobinsonLabs Qwen 262B GGUF MoE 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/RobinsonLabs%2FQwen3.5-REAP-262B-A17B-abliterated-GGUF"
Response includes
  • classification m8
  • files 13
  • hub_downloads_all_time 16,377
  • author_summary 15 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
16K
13K last 30d - active
Likes
2
Model age
3mo ago
created 2026-07-12
Downloads over time
Now16.5K→from1.8K↑826%
1K6.7K12.3K18K1.8K on Jul 1516.5K on Oct 11JulAugSepOct
Jul 15 → Oct 11 · 53 snapshots · spans 88 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 2 formats · 13K downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
apache-2.0
Quantizations
IQ2 IQ3 IQ4 Q3_K Q4_K Q5_K
Tags
gguf abliterated uncensored imatrix moe reap qwen3.5 not-for-all-audiences text-generation arxiv:2510.13999 base_model:OpenMOSE/Qwen3.5-REAP-262B-A17B base_model:quantized:OpenMOSE/Qwen3.5-REAP-262B-A17B

Related

Total size
1.10 TB
Files
13
Quantizations
7
Registered
2026-08-22 13:56
Last updated on HF
2026-09-12 17:32

Files by quantization

Q5_K 1 file 173 GB
Qwen3.5-REAP-262B-A17B-abl-Q5_K_M.gguf 173 GB 8d2a0537 download
Q4_K 2 files 286 GB
Qwen3.5-REAP-262B-A17B-abl-Q4_K_M.gguf 148 GB 1ebb8ab4 download
Qwen3.5-REAP-262B-A17B-abl-Q4_K_S.gguf 139 GB 0ee5bf01 download
IQ4 1 file 130 GB
Qwen3.5-REAP-262B-A17B-abl-IQ4_XS.gguf 130 GB fcc4e587 download
Q3_K 1 file 116 GB
Qwen3.5-REAP-262B-A17B-abl-Q3_K_M.gguf 116 GB a68aeb8b download
IQ3 2 files 207 GB
Qwen3.5-REAP-262B-A17B-abl-IQ3_M.gguf 107 GB 7b8a7cba download
Qwen3.5-REAP-262B-A17B-abl-IQ3_XS.gguf 99.9 GB 82e8429a download
IQ2 3 files 216 GB
Qwen3.5-REAP-262B-A17B-abl-IQ2_M.gguf 79.8 GB 1b587510 download
Qwen3.5-REAP-262B-A17B-abl-IQ2_XS.gguf 71.8 GB 33de9212 download
Qwen3.5-REAP-262B-A17B-abl-IQ2_XXS.gguf 64.5 GB 13021272 download
Auxiliary files 3 files 87.3 KB
bpw-vs-size.png 79.5 KB d1a67e74 download
README.md 5.58 KB fb5c46cb download
.gitattributes 2.21 KB 54f32e98 download

README current version from Hugging Face


license: apache-2.0
base_model: OpenMOSE/Qwen3.5-REAP-262B-A17B
library_name: gguf
pipeline_tag: text-generation
tags:

  • gguf
  • abliterated
  • uncensored
  • imatrix
  • moe
  • reap
  • qwen3.5
  • not-for-all-audiences

Qwen3.5-REAP-262B-A17B - Abliterated GGUF

Abliterated GGUF quant ladder of
OpenMOSE/Qwen3.5-REAP-262B-A17B,
itself a 34% REAP expert-pruning of Qwen3.5-397B-A17B down to
262B total / ~17B active. The IQ rungs are importance-matrix (imatrix) weighted.

Provenance chain: the abliterated bf16 safetensors base was converted to a Q8_0 master
(277.7 GB, near-lossless), and every rung here is cut from that master. The bf16 base lives at
RobinsonLabs/Qwen3.5-REAP-262B-A17B-abliterated,
use that repo if you want to re-abliterate, merge a LoRA, fine-tune, or roll your own quants.

Disclosure

This model is abliterated: the hard-refusal reflex on adult / creative content has been
reduced via single-direction weight orthogonalization. Harm guardrails are retained by design,
self-harm prompts still redirect to help (e.g. 988), and it is not intended to assist genuine
wrongdoing. Capability is preserved. Tagged not-for-all-audiences. Use responsibly, you are
responsible for what you generate with it. License inherited from the base model: Apache-2.0.

Files

File Quant bpw ~Size imatrix Notes
...-Q5_K_M.gguf Q5_K_M 5.68* ~186 GB no highest-fidelity rung published
...-Q4_K_M.gguf Q4_K_M 4.85* ~159 GB no K-quant quality pick
...-Q4_K_S.gguf Q4_K_S 4.55* ~149 GB no
...-IQ4_XS.gguf IQ4_XS 4.28 ~140 GB yes quality/size sweet spot
...-Q3_K_M.gguf Q3_K_M 3.83 ~125 GB no
...-IQ3_M.gguf IQ3_M 3.51 ~115 GB yes
...-IQ3_XS.gguf IQ3_XS 3.29 ~107 GB yes
...-IQ2_M.gguf IQ2_M 2.62 ~86 GB yes
...-IQ2_XS.gguf IQ2_XS 2.36 ~77 GB yes
...-IQ2_XXS.gguf IQ2_XXS 2.12 ~69 GB yes smallest

bpw figures are as reported by llama-quantize, not nominal. The three starred rungs predate the
surviving build logs, their bpw is computed from exact file bytes over the 261.6B parameter count.
The IQ rungs are imatrix-weighted and land meaningfully smaller than the K-quant of comparable
quality: IQ4_XS undercuts Q4_K_S by ~9 GB, and the IQ2 family is the only path under 90 GB.

Quant ladder, bits-per-weight vs file size

The chart shows the contested 2 to 5 bpw band; the higher-fidelity Q5_K_M rung is in the table
above.

Architecture notes

qwen3_5_moe hybrid: 60 decoder layers (45 linear-attn / DeltaNet + 15 full-attn,
full_attention_interval=4), 333 experts with 10 active per token, hidden size 4096,
head dim 256, 262144 native context. No MTP / NextN layer. This is the text path only
(no vision mmproj).

The high expert count is the defining feature of this REAP tier: 333 experts versus 267 on the
48%-pruned 212B sibling.
More experts retained means more of the 397B parent's routing diversity survives, at the cost of
size.

Method

  • Abliteration: single-direction weight orthogonalization (FailSpy / Labonne method). For every
    matrix that writes the residual stream (o_proj, DeltaNet out_proj, fused expert down_proj,
    shared-expert down_proj, and the token embedding), the rank-1 component along the refusal
    direction is subtracted. Routers and norms pass through byte-identical.
  • Refusal direction, massive-activation guarded. The direction is captured with a
    mean-difference control vector, then guarded against attention-sink contamination: the sink
    dimensions that dominate raw activation magnitude (and would brick the model if ablated) are
    detected across layers and excluded, and the direction is taken from the clean, spread-out
    consensus of the late layers rather than a single sink-dominated layer.
  • Quant: convert and quantize with llama.cpp
    (build b9244). bf16 to a Q8_0 master (8.51 bpw as measured), then every rung cut from that
    master.
  • imatrix: the IQ rungs are weighted by an importance matrix computed over the abliterated
    model itself against corpus-rldomain, a domain-calibrated corpus. 200 chunks at n_ctx=512
    (~102K tokens), 765 importance entries, final PPL 19.01 on the calibration set. The high chunk
    count is deliberate: with 333 experts, a short calibration run leaves rarely-routed experts
    under-exercised, and per-tensor coverage was still climbing well past the point where a
    200-expert model would have saturated.

bf16 base

The full-precision bf16 safetensors master this ladder derives from is at
RobinsonLabs/Qwen3.5-REAP-262B-A17B-abliterated.
That repo is the master for further surgery (re-abliteration, LoRA merge, fine-tune) and for making
your own quants.

Provenance

Qwen3.5-397B-A17B (Apache-2.0) -> OpenMOSE/Qwen3.5-REAP-262B-A17B (34% REAP prune) ->
abliterated (bf16 master) -> Q8_0 master -> quant rungs. Every published rung, K and IQ alike,
is cut from the Q8_0 master (the IQ rungs with --allow-requantize), not directly from the bf16.
Recipe and diagnosis are Robinson Labs internal (WI #1423).

Built by Robinson Labs.

README history 6 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-09-12Squash: reclaim superseded June big-rung blobs after D51H drain completec0112e66.8 KB
    Loading...
  2. 2026-08-29Link corrected v2 repoa7f191d6.3 KB
    Loading...
  3. 2026-08-29Disclose known weak-abliteration issue; corrected v2 in progress (IQ3_XS first)09343686.1 KB
    Loading...
  4. 2026-07-18Finalize card: full 10-rung table, measured bpw, chart, provenance (WI #1423)0fc7cef5.5 KB
    Loading...
  5. 2026-07-15Upload README.md with huggingface_hubacb5e304.8 KB
    Loading...
  6. 2026-07-12Upload README.md with huggingface_hub726af8b777 B
    Loading...

Discussions 1 thread

  1. 2026-08-29Abliterated?open4 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration