← back to catalog · registered 2026-08-22 13:56

zkdtckk/rai-oss-safeguard-20b-nothinking

zkdtckk Gpt-oss 20B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/zkdtckk%2Frai-oss-safeguard-20b-nothinking"
Response includes
  • classification unknown
  • files 15
  • benchmarks 16 entries
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
?
Primary method

Unclassified

No clear signals of an abliteration technique in this model.
Confidence
UNKNOWN
Why this label 1 signal
No classification signals present. This may not be an abliterated model at all - it could be a repackaging, a merge with unrelated goals, or unrelated content that mentions the term.
  • no classification signals present (no abliterated, uncensored, or known producer/method markers)
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
13
↑ 93% in 90 days
Likes
0
Model age
8mo ago
created 2026-02-10
Downloads over time
Now83→from43↑93%
4156728743 on Feb 1183 on Oct 1183 on Oct 8FebAprJunAugOct
Feb 11 → Oct 11 · 74 snapshots · spans 242 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Arena-Battles 7952 LM-Arena
LM Arena Elo 1307.3530229799078 LM-Arena
Arena-Elo-Lower 1300.2396666200118 LM-Arena
Arena-Elo-Upper 1314.4663793398038 LM-Arena
Arena-Rank 70 LM-Arena
Entertainment 1.1 UGI
Hazardous 0 UGI
Natural Intelligence 17.39 UGI
Political lean -10.6% UGI
Sensitive-Info 7.19 UGI
SocPol 0.8 UGI
UGI 8.96 UGI
Willingness (10) 1.2 UGI
W10-Adherence 1.5 UGI
W10-Direct 1 UGI
Writing 24.62 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
transformers safetensors gpt_oss text-generation vllm conversational arxiv:2508.10925 base_model:openai/gpt-oss-20b base_model:finetune:openai/gpt-oss-20b license:apache-2.0 endpoints_compatible 8-bit

Related

Total size
12.8 GB
Files
15
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-02-11 23:27

Files by quantization

Auxiliary files 15 files 12.8 GB
model-00001-of-00002.safetensors 4.47 GB 19684279 download
model-00000-of-00002.safetensors 4.46 GB ef82af3a download
model-00002-of-00002.safetensors 3.88 GB e1856997 download
tokenizer.json 26.6 MB 0614fe83 download
model.safetensors.index.json 35.5 KB ae085214 download
tokenizer_config.json 21.2 KB 214abcd9 download
chat_template.jinja 18.3 KB c9508a84 download
chat_template.json 17.1 KB 590f0e82 download
LICENSE 11.1 KB d6456956 download
README.md 3.94 KB 4428e31b download
config.json 1.76 KB bc5df3bc download
.gitattributes 1.53 KB 52373fe2 download
USAGE_POLICY 236 B e9dbeda1 download
generation_config.json 165 B d1cdba2c download
special_tokens_map.json 98.0 B 73bd12e5 download

README current version from Hugging Face


license: apache-2.0
pipeline_tag: text-generation
library_name: transformers
tags:

  • vllm
    base_model:
  • openai/gpt-oss-20b
    base_model_relation: finetune

gpt-oss-safeguard-20b

Try gpt-oss-safeguard · Guide · Model card · OpenAI blog


gpt-oss-safeguard-120b and gpt-oss-safeguard-20b are safety reasoning models built-upon gpt-oss. With these models, you can classify text content based on safety policies that you provide and perform a suite of foundational safety tasks. These models are intended for safety use cases. For other applications, we recommend using gpt-oss models.

This model gpt-oss-safeguard-20b (21B parameters with 3.6B active parameters) fits into GPUs with 16GB of VRAM. Check out gpt-oss-safeguard-120b (117B parameters with 5.1B active parameters) for the larger model.

Both models were trained on our harmony response format and should only be used with the harmony format as it will not work correctly otherwise.

Highlights

  • Trained to reason about safety : Trained and tuned for safety reasoning to accommodate use cases like LLM input-output filtering, online content labeling and offline labeling for Trust and Safety use cases.
  • Bring your own policy: Interprets your written policy, so it generalizes across products and use cases with minimal engineering.
  • Reasoned decisions, not just scores: Gain complete access to the model’s reasoning process, facilitating easier debugging and increased trust in policy decisions. Keep in mind Raw CoT is meant for developers and safety practitioners. It’s not intended for exposure to general users or use cases outside of safety contexts.
  • Configurable reasoning effort: Easily adjust the reasoning effort (low, medium, high) based on your specific use case and latency needs.
  • Permissive Apache 2.0 license: Build freely without copyleft restrictions or patent risk—ideal for experimentation, customization, and commercial deployment.

Inference examples

You can use gpt-oss-safeguard-120b and gpt-oss-safeguard-20b similar to gpt-oss-120b and gpt-oss-20b as described in our respective cookbooks. We’ve also provided a detailed prompting guide that provides guidelines for how to craft your policy and use it with the models.

Download the model

To download the model weights from Hugging Face hub using similar instructions to gpt-oss-120b.

Join the ROOST Model Community

gpt-oss-safeguard is a model partner of the Robust Open Online Safety Tools (ROOST) Model Community. The ROOST Model Community (RMC) is a group of safety practitioners exploring open source AI models to protect online spaces. As an RMC model partner, OpenAI is committed to incorporating user feedback and jointly iterating on future releases in pursuit of open safety. Visit the RMC GitHub repo to learn more about this partnership and how to get involved.

Resources

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-02-10Upload 13 files784a7fb3.9 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration