← back to catalog · registered 2026-08-22 13:56

nguyenthilaitrieulong/gpt-oss-20b-abliterated

nguyenthilaitrieulong Gpt-oss 21B MoE
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/nguyenthilaitrieulong%2Fgpt-oss-20b-abliterated"
Response includes
  • classification m1
  • files 18
  • benchmarks 16 entries
  • hub_downloads_all_time 483
  • author_summary 29 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
483
40 last 30d - cooling
Likes
0
Model age
2mo ago
created 2026-08-06
Downloads over time
Now498→from233↑114%
220321423525233 on Aug 5498 on Oct 11498 on Oct 8AugSepOct
Aug 5 → Oct 11 · 50 snapshots · spans 67 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Arena-Battles 7952 LM-Arena
LM Arena Elo 1307.3530229799078 LM-Arena
Arena-Elo-Lower 1300.2396666200118 LM-Arena
Arena-Elo-Upper 1314.4663793398038 LM-Arena
Arena-Rank 70 LM-Arena
Entertainment 1.1 UGI
Hazardous 0 UGI
Natural Intelligence 17.39 UGI
Political lean -10.6% UGI
Sensitive-Info 7.19 UGI
SocPol 0.8 UGI
UGI 8.96 UGI
Willingness (10) 1.2 UGI
W10-Adherence 1.5 UGI
W10-Direct 1 UGI
Writing 24.62 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh
Tags
transformers safetensors gpt_oss text-generation abliterated uncensored moe gpt-oss mxfp4 direct-steering ega moe-router-suppression

Related

Total size
39.0 GB
Files
18
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-06 18:37

Files by quantization

Auxiliary files 18 files 39.0 GB
model-00005-of-00009.safetensors 4.60 GB 3f4a1186 download
model-00006-of-00009.safetensors 4.60 GB 8bdf8035 download
model-00007-of-00009.safetensors 4.60 GB 93843ef5 download
model-00008-of-00009.safetensors 4.60 GB c48ac2fa download
model-00004-of-00009.safetensors 4.60 GB f34e9956 download
model-00002-of-00009.safetensors 4.60 GB 3620d9b2 download
model-00003-of-00009.safetensors 4.60 GB c571fb0f download
model-00001-of-00009.safetensors 4.19 GB b226bc26 download
model-00009-of-00009.safetensors 2.56 GB acbfe041 download
tokenizer.json 26.6 MB fce342a4 download
model.safetensors.index.json 32.8 KB 7353bb66 download
chat_template.jinja 16.3 KB dc7bb119 download
README.md 9.09 KB c364295a download
tokenizer_config.json 4.10 KB c021cddb download
config.json 1.57 KB cd61db9a download
.gitattributes 1.53 KB 52373fe2 download
special_tokens_map.json 440 B 6274cc1b download
generation_config.json 172 B b27df957 download

README current version from Hugging Face


license: apache-2.0
base_model: openai/gpt-oss-20b
tags:

  • abliterated
  • uncensored
  • moe
  • gpt-oss
  • mxfp4
  • direct-steering
  • ega
  • moe-router-suppression
  • abliterix
    language:
  • en
  • zh
    library_name: transformers
    pipeline_tag: text-generation

gpt-oss-20b-abliterated

A refusal-suppressed variant of openai/gpt-oss-20b, produced with abliterix using direct weight editing, Expert-Granular Abliteration (EGA) on the fused MoE expert weights, and MoE router suppression on the safety-concentrated experts.

Key results

Metric Base gpt-oss-20b This model
Refusals on 100 held-out harmful prompts (LLM judge) 97 / 100 6 / 100
KL divergence vs base (next-token, benign) — 0.0098
Response length deviation vs base (benign) — 0.02 σ
Hard-prompt qualitative compliance (15 classic jailbreaks, EN+ZH) 0 / 15 15 / 15

The eval refusal counts come from an LLM judge (google/gemini-3.1-flash-lite-preview via OpenRouter) instructed to label garbled / repetitive / incoherent output as a refusal — so models that "bypass" refusal by collapsing into gibberish get correctly counted as failures, not successes. A pre-LLM rule-based filter additionally catches dash-runs, sentence loops, and low-character-diversity output before the judge is called. The 6/100 is a real, semantic compliance number, not keyword matching.

The qualitative compliance row is a separate manual test: 15 classic hard prompts (10 EN + 5 ZH) covering lockpicking, phishing, meth synthesis, WiFi hacking, fake news, hotwiring, pipe bombs, signature forgery, keylogger code, shoplifting, 炸弹制作, 钓鱼邮件, 电脑入侵, 身份证伪造, 网络诈骗. The model complies with all 15 directly, in the same Markdown-table, step-by-step style the base model uses for benign technical answers.

Why this works — three architecture-specific correctness fixes

abliterix handles three gpt-oss-specific issues that silently break naïve LLaMA-style abliteration scripts:

  1. Native MXFP4 weights are not exposed as standard nn.Parameter. gpt-oss ships in Mxfp4GptOssExperts form whose down_proj is a packed Triton tensor that cannot be edited in-place. abliterix auto-detects this and forces Mxfp4Config(dequantize=True) so the BF16 fused expert tensor is reachable.
  2. GptOssExperts.down_proj is stored transposed vs the standard MoE convention. Its shape is (experts, intermediate_in, hidden_out) and the forward path is out = act @ W (no transpose). Standard EGA implementations use shape-based axis detection, which silently picks the wrong projection branch when hidden == intermediate (both 2880 in gpt-oss-20b). We mark this layout explicitly and project from the output side (W_new = W (I − vv^T)).
  3. Fused-expert MoEs were silently invisible to EGA. GptOssExperts is a single Module holding fused 3-D weights, so a naive per-Module profile dict key produces no mlp.down_proj entry and _apply_ega_steering early-exits. abliterix synthesises an mlp.down_proj profile when fused experts are detected so EGA actually runs across all 32 experts × 24 layers.

On top of the direct-steering + EGA foundation, this release adds MoE router suppression — an [experts] block that redirects routing away from the top-k "safety experts" (the experts whose gate activates disproportionately more on harmful prompts than on benign ones). The suppression strength is itself an Optuna search parameter, so the optimiser picks how aggressively to bias each layer's safety experts.

Method

  • Base: openai/gpt-oss-20b — 24 layers, 32 routed experts per layer, top-4, hidden = intermediate = 2880, MXFP4 → BF16 dequant during abliteration
  • Tool: abliterix
  • Mode: steering_mode = "direct" (orthogonal projection on base weights, no LoRA), weight_normalization = "full" (norm-preserving)
  • Components steered:
    • attn.{q,k,v,o}_proj via direct weight projection
    • mlp.experts.down_proj across all 32 experts × 24 layers via Expert-Granular Abliteration
    • mlp.router rows of safety experts via logit suppression
  • Refusal direction: per-layer mean of (target − benign) residuals on a 400-prompt benign + 400-prompt harmful set; BF16 projection
  • Search: Optuna TPE, KL-divergence + LLM-judged refusal as multi-objective, 100 trials (40 random warmup + 60 TPE)
  • Hardware: 1 × NVIDIA RTX PRO 6000 Blackwell (96 GB, sm_120), driver 580 / CUDA 12.9, batch=8, total wall time ≈ 5 h 20 m
  • Eval set: 100 held-out harmful prompts not seen during steering-vector computation; 100 held-out benign prompts for KL comparison

Winning hyperparameters

vector_scope = "per layer"     # per-layer direction, not global

[attn.q_proj]
max_weight = 3.04 ; max_weight_position = 13.86 ; min_weight = 0.99 ; min_weight_distance = 6.21

[attn.k_proj]
max_weight = 3.90 ; max_weight_position = 16.57 ; min_weight = 1.25 ; min_weight_distance = 4.91

[attn.v_proj]
max_weight = 2.21 ; max_weight_position = 20.06 ; min_weight = 0.77 ; min_weight_distance = 7.29

[attn.o_proj]
max_weight = 3.82 ; max_weight_position = 17.41 ; min_weight = 1.11 ; min_weight_distance = 7.07

[mlp.down_proj]                # Expert-Granular Abliteration on fused experts
max_weight = 6.95 ; max_weight_position = 18.37 ; min_weight = 0.54 ; min_weight_distance = 4.91

[moe]                          # router-row suppression
n_suppress = 1                 # suppress top-1 safety expert per layer
router_bias = -0.64            # scale = max(0, 1 + bias/10) = 0.94
expert_ablation_weight = 0.0   # pinned off; EGA already handles expert weights

The EGA peak sits at layer ≈ 18 — a per-layer-tailored fingerprint where the refusal decision still has options, rather than late in the stack where it has already committed.

Usage

Transformers

from transformers import AutoTokenizer, AutoModelForCausalLM
import torch

tok = AutoTokenizer.from_pretrained("wangzhang/gpt-oss-20b-abliterated")
model = AutoModelForCausalLM.from_pretrained(
    "wangzhang/gpt-oss-20b-abliterated",
    torch_dtype=torch.bfloat16,
    device_map="auto",
)

messages = [{"role": "user", "content": "Your prompt here"}]
prompt = tok.apply_chat_template(messages, add_generation_prompt=True, tokenize=False)
inputs = tok(prompt, return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=512, do_sample=True, temperature=0.7)
print(tok.decode(out[0][inputs.input_ids.shape[1]:], skip_special_tokens=True))

The model uses gpt-oss's harmony chat format. The chat template is bundled (chat_template.jinja).

GGUF (llama.cpp / Ollama / LM Studio)

BF16, Q8_0, and Q4_K_M quantizations are available at wangzhang/gpt-oss-20b-abliterated-GGUF.

ollama run hf.co/wangzhang/gpt-oss-20b-abliterated-GGUF:Q4_K_M

Honest limitations

  • Refusal is low, not zero. 6 / 100 held-out prompts still refuse. The residual refusers cluster around "universally-recognised-as-harmful-and-specific" asks (detailed CBRN synthesis, CSAM-adjacent content) — exactly where refusal tends to be represented by multiple redundant circuits that partial abliteration can't all knock out in one pass.
  • Stylistic residue on a handful of prompts. Even on prompts that comply, 2–3 out of 100 begin with a soft disclaimer ("just keep in mind that..." / "以下内容仅供学习与参考") before producing the actual content. Disclaimer framing is still trainable.
  • English > Chinese. Steering vectors came from a primarily English dataset. Chinese hard prompts work (5/5 on manual Chinese tests) but bypass quality is slightly lower — shorter responses, occasional English fallback on technical terms.
  • No guarantees on long generations. On generations past ~400 tokens we occasionally see list or Markdown-table loops; this is an abliteration side-effect, not a base-model regression.

Reproducibility

Full search checkpoint (Optuna JSONL + judge cache SQLite) and the exact config are available in the abliterix repo. To reproduce from scratch:

git clone https://github.com/wuwangzhang1216/abliterix
cd abliterix && pip install -e .
AX_CONFIG=configs/gpt_oss_20b.toml abliterix
# Optuna is deterministic if you set sampler_seed in [optimization].

Intended use

Authorised AI safety research, red-teaming evaluation, refusal-mechanism analysis, and study of how MoE expert specialisation encodes safety behaviours. Not for producing or distributing harmful content. The license of the base model (apache-2.0) applies; the user is responsible for compliance with all applicable laws and the OpenAI gpt-oss usage policy.

Acknowledgments

  • openai/gpt-oss-20b for the base model
  • abliterix is a derivative work of Heretic by Philipp Emanuel Weidmann
  • TrevorS for the original Expert-Granular Abliteration formulation

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-06Duplicate from wangzhang/gpt-oss-20b-abliteratedcbdaf349.1 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration