← back to catalog · registered 2026-08-22 13:56

Jazhyc/gemma-3-12b-intent-jailbreak-classifier-sft

Jazhyc Gemma 12B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/Jazhyc%2Fgemma-3-12b-intent-jailbreak-classifier-sft"
Response includes
  • classification unknown
  • files 9
  • benchmarks 16 entries
  • hub_downloads_all_time 38
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
?
Primary method

Unclassified

No clear signals of an abliteration technique in this model.
Confidence
UNKNOWN
Why this label 1 signal
No classification signals present. This may not be an abliterated model at all - it could be a repackaging, a merge with unrelated goals, or unrelated content that mentions the term.
  • no classification signals present (no abliterated, uncensored, or known producer/method markers)
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
38
12 last 30d - stable
Likes
0
Model age
5mo ago
created 2026-05-03
Downloads over time
Now39→from13↑200%
1222324213 on May 639 on Oct 1139 on Oct 5MayJunJulAugSepOct
May 6 → Oct 11 · 62 snapshots · spans 158 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Arena-Battles 3976 LM-Arena
LM Arena Elo 1335.3304642871612 LM-Arena
Arena-Elo-Lower 1326.1060034720686 LM-Arena
Arena-Elo-Upper 1344.5549251022537 LM-Arena
Arena-Rank 49 LM-Arena
Entertainment 1.3 UGI
Hazardous 2.9 UGI
Natural Intelligence 18.72 UGI
Political lean -11.7% UGI
Sensitive-Info 16.33 UGI
SocPol 1 UGI
UGI 20.89 UGI
Willingness (10) 3 UGI
W10-Adherence 0 UGI
W10-Direct 6 UGI
Writing 29.86 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

Tags
peft safetensors lora sft safety intent jailbreak-detection text-generation conversational base_model:google/gemma-3-12b-it base_model:adapter:google/gemma-3-12b-it region:us

Related

Total size
250 MB
Files
9
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-05-03 21:46

Files by quantization

Auxiliary files 9 files 282 MB
adapter_model.safetensors 250 MB 293bb673 download
training_args.bin 5.64 KB da2dc7d3 download
tokenizer.json 31.8 MB daab2354 download
README.md 2.72 KB b2da08af download
.gitattributes 1.53 KB 52373fe2 download
chat_template.jinja 1.50 KB 1117055a download
adapter_config.json 1.03 KB 300a87ff download
tokenizer_config.json 716 B d6d32716 download
val_metrics.json 266 B 9e5706f9 download

README current version from Hugging Face


base_model: google/gemma-3-12b-it
library_name: peft
pipeline_tag: text-generation
tags:

  • lora
  • peft
  • sft
  • safety
  • intent
  • jailbreak-detection

gemma-3-12b-intent-jailbreak-classifier-sft

LoRA adapter for google/gemma-3-12b-it, supervised-fine-tuned directly on the
Jazhyc/wildguard-annotated-intents dataset (no teacher distillation). The
model produces a verbalized user intent and a harm classification for a given
prompt, and serves as the SFT baseline counterpart to
Jazhyc/gemma-3-12b-intent-jailbreak-classifier,
which is distilled from openai/gpt-oss-120b.

Performance

Metric Value
OOD validation harm F1 (mean of ToxicChat train + Aegis 2.0 val) 0.7779
In-domain val harm F1 0.7197
In-domain val harm precision 0.7584
In-domain val harm recall 0.6848
In-domain val semantic similarity (intent) 0.7260

OOD F1 is averaged over lmsys/toxic-chat (toxicchat0124, train split) and
nvidia/Aegis-AI-Content-Safety-Dataset-2.0 (validation split). Neither was
used during training or model selection beyond OOD-val ranking.

Training

  • Base model: google/gemma-3-12b-it
  • Training data: human-annotated intents from Jazhyc/wildguard-annotated-intents
  • Learning rate: 1e-05
  • Adapter: LoRA, r=16, alpha=32, dropout=0.0, targets q/k/v/o/gate/up/down_proj
  • Quantization: 4-bit NF4 (QLoRA) during training
  • Attention: flash_attention_2 with sequence packing
  • Early stopping on combined val+test split (patience 1)

Usage

from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer

base = "google/gemma-3-12b-it"
adapter = "Jazhyc/gemma-3-12b-intent-jailbreak-classifier-sft"

tokenizer = AutoTokenizer.from_pretrained(base)
model = AutoModelForCausalLM.from_pretrained(base, torch_dtype="bfloat16", device_map="auto")
model = PeftModel.from_pretrained(model, adapter)

messages = [
    {"role": "system", "content": "<system prompt from build_student_messages(..., 'human_intent')>"},
    {"role": "user",   "content": "<prompt to classify>"},
]
inputs = tokenizer.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device)
out = model.generate(inputs, max_new_tokens=512)
print(tokenizer.decode(out[0][inputs.shape[-1]:], skip_special_tokens=True))

The expected output format is <intent>...</intent><harm>safe|harmful</harm>.
See src/intention_jailbreak/model_generation/prompt_templates.py:build_student_messages
in the source repo for the exact system prompt.

Citation / source

Part of the intention-jailbreak research project on jailbreak detection via
verbalized intent reasoning.

Framework versions

  • PEFT 0.18.1

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-05-03Upload SFT-baseline gemma-3-12b adapter75565452.7 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration