← back to catalog · registered 2026-08-22 13:56

ZDCSlab/ripd-anthropic-saferlhf-gemma-2b-uncensored-v1-biased-bt

ZDCSlab Gemma 3.2B second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/ZDCSlab%2Fripd-anthropic-saferlhf-gemma-2b-uncensored-v1-biased-bt"
Response includes
  • classification m-uncensored
  • files 13
  • hub_downloads_all_time 409
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
409
109 last 30d - stable
Likes
0
Model age
7mo ago
created 2026-02-19

Training datasets

Corpora the author lists in the model card. Datasets tracked in our /datasets catalog carry a category badge linking to the workflow stage. Others open on Hugging Face.

Downloads over time
Now454→from0↑0%
01663334990 on Feb 18454 on Oct 11454 on Oct 8FebAprJunAugOct
Feb 18 → Oct 11 · 73 snapshots · spans 235 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

Tags
transformers safetensors gemma2 text-generation alignment evaluation preference-learning ripd conversational dataset:ZDCSlab/ripd-dataset arxiv:2602.13576 base_model:sirev/Gemma-2b-Uncensored-v1

Related

Total size
5.97 GB
Files
13
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-02-21 16:26

Files by quantization

Auxiliary files 13 files 6.00 GB
model-00001-of-00002.safetensors 4.65 GB 8ca76e75 download
model-00002-of-00002.safetensors 1.32 GB 51a3a11a download
training_args.bin 8.20 KB 32e03acd download
tokenizer.json 32.8 MB 5f7eee61 download
tokenizer.model 4.04 MB 61a7b147 download
tokenizer_config.json 45.3 KB 3ade6be5 download
model.safetensors.index.json 23.7 KB 1d45054f download
.gitattributes 1.53 KB 52373fe2 download
config.json 1.41 KB a700e856 download
README.md 1.35 KB c74cb52e download
special_tokens_map.json 636 B 8d6368f7 download
chat_template.jinja 591 B 923ec253 download
generation_config.json 194 B abda748f download

README current version from Hugging Face


library_name: transformers
pipeline_tag: text-generation
tags:

  • alignment
  • evaluation
  • preference-learning
  • ripd
    base_model: sirev/Gemma-2b-Uncensored-v1
    datasets:
  • ZDCSlab/ripd-dataset

ZDCSlab/ripd-anthropic-saferlhf-gemma-2b-uncensored-v1-biased-bt

This checkpoint is part of the artifact release for
“Rubrics as an Attack Surface: Stealthy Preference Drift in LLM Judges.”

It is a policy model trained under a specific rubric condition to study how evaluation-time preference drift propagates into downstream alignment.


Configuration

  • Setting: anthropic-saferlhf
  • Base model: Gemma-2b-Uncensored-v1
  • Label condition: biased
  • Training data: Bench + Target (mixed)
  • Objective: Direct Preference Optimization (DPO)

The biased condition corresponds to preference labels generated by an LLM judge under the biased rubric variant.


Intended Use

This model is released for research on evaluation-time robustness, preference drift, and alignment propagation.
It is not intended for production deployment.


Resources

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-02-19Add model cardf75ac191.3 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration