For privacy reasons a browser tells us at most
"≥ 8 GB RAM, 8 cores" - same reading whether you have 8 GB
or 128 GB. It has no idea how much RAM is free right now, which
apps are open, or whether you have a GPU.
The Abliteration app is integrated with your machine
It reads your exact RAM, GPU model and VRAM,
free memory right now, and picks the sharpest quant that
still fits. Every model page lights up precisely for your rig.
And you can chat with any model, right now
The app is a full local runtime - no API keys, no subscription,
everything runs on your machine. Click any model on this site and
start a conversation in seconds.
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
'abliterated' in name/tags
is_gguf=0 (base model)
no specific method indicators - defaulting to M1 (most common)
Abliterated bf16 base of OpenMOSE/Qwen3.5-REAP-212B-A17B, in safetensors — itself a 48% REAP expert-pruning of Qwen3.5-397B-A17B down to 212B total / ~17B active.
This is the full-precision master behind the RobinsonLabs/Qwen3.5-REAP-212B-A17B-abliterated-GGUF quant ladder (the ladder itself is cut from a Q8_0 quant master converted from this base). If you want a ready-to-run quant, use that repo. This repo is the master for further surgery (re-abliteration, LoRA merge, fine-tune) and for making your own quants.
Disclosure
This model is abliterated — the hard-refusal reflex on adult / creative content has been reduced via single-direction weight orthogonalization. Harm guardrails are retained by design: self-harm prompts still redirect to help (e.g. 988), and it is not intended to assist genuine wrongdoing. Capability is preserved. Tagged not-for-all-audiences. Use responsibly — you are responsible for what you generate with it. License inherited from the base model: Apache-2.0.
Method
Abliteration — single-direction weight orthogonalization on the bf16 base: the rank-1 refusal component is subtracted from every residual-stream writer (o_proj, DeltaNet out_proj, fused + shared-expert down_proj, token embedding). 181 tensors edited; routers and norms untouched. The refusal direction is massive-activation guarded — attention-sink dimensions that would brick the model are detected and excluded, and the direction is drawn from the clean late-layer consensus (KB #213).
Format — safetensors, sharded (122 shards), with config + tokenizer + index.
Architecture
qwen3_5_moe hybrid: 60 decoder layers (45 linear-attn / DeltaNet + 15 full-attn, full_attention_interval=4), 267 experts, hidden size 4096, ~211.8B params / ~17B active. Text path only. The REAP prune drops the native MTP head; the GGUF ladder corrects the phantom MTP layer at convert time so it loads as a clean 60-layer decoder.
The author's README evolved over time. Click a version to see its content at that point.
2026-09-26Withdraw model (2026-09-25)27ed972240 B
Loading...
2026-09-14Update model card from hf/cards/ (repo is source of truth)ca151182.9 KB
Loading...
2026-07-18Fix provenance: quants are cut from the Q8_0 master, not directly from bf16 (...d53c0702.9 KB
Loading...
2026-07-11Complete model card / chart (README.md)33d32492.9 KB
Loading...
Catalog is the map. Apps are the tools.
Run models on your own machine, not in the cloud.
Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.