← back to catalog · registered 2026-08-22 13:56

Jazhyc/gemma-3-12b-intent-jailbreak-classifier

Jazhyc Gemma 12B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/Jazhyc%2Fgemma-3-12b-intent-jailbreak-classifier"
Response includes
  • classification unknown
  • files 9
  • benchmarks 16 entries
  • hub_downloads_all_time 40
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
?
Primary method

Unclassified

No clear signals of an abliteration technique in this model.
Confidence
UNKNOWN
Why this label 1 signal
No classification signals present. This may not be an abliterated model at all - it could be a repackaging, a merge with unrelated goals, or unrelated content that mentions the term.
  • no classification signals present (no abliterated, uncensored, or known producer/method markers)
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
40
13 last 30d - stable
Likes
0
Model age
5mo ago
created 2026-05-03
Downloads over time
Now42→from12↑250%
1122344512 on May 642 on Oct 1142 on Oct 6MayJunJulAugSepOct
May 6 → Oct 11 · 62 snapshots · spans 158 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Arena-Battles 3976 LM-Arena
LM Arena Elo 1335.3304642871612 LM-Arena
Arena-Elo-Lower 1326.1060034720686 LM-Arena
Arena-Elo-Upper 1344.5549251022537 LM-Arena
Arena-Rank 49 LM-Arena
Entertainment 1.3 UGI
Hazardous 2.9 UGI
Natural Intelligence 18.72 UGI
Political lean -11.7% UGI
Sensitive-Info 16.33 UGI
SocPol 1 UGI
UGI 20.89 UGI
Willingness (10) 3 UGI
W10-Adherence 0 UGI
W10-Direct 6 UGI
Writing 29.86 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

Tags
peft safetensors lora distillation safety intent jailbreak-detection text-generation conversational base_model:google/gemma-3-12b-it base_model:adapter:google/gemma-3-12b-it region:us

Related

Total size
500 MB
Files
9
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-05-03 19:40

Files by quantization

Auxiliary files 9 files 531 MB
adapter_model.safetensors 500 MB ac607015 download
training_args.bin 5.77 KB 33788245 download
tokenizer.json 31.8 MB daab2354 download
README.md 2.85 KB 49021a9e download
.gitattributes 1.53 KB 52373fe2 download
chat_template.jinja 1.50 KB 1117055a download
adapter_config.json 1.03 KB 3c954299 download
tokenizer_config.json 716 B d6d32716 download
val_metrics.json 284 B cdd09da3 download

README current version from Hugging Face


base_model: google/gemma-3-12b-it
library_name: peft
pipeline_tag: text-generation
tags:

  • lora
  • peft
  • distillation
  • safety
  • intent
  • jailbreak-detection

gemma-3-12b-intent-jailbreak-classifier

LoRA adapter for google/gemma-3-12b-it distilled from openai/gpt-oss-120b as a teacher,
under the human_intent reasoning-trace condition (the student is trained on
reasoning traces produced by the teacher conditioned on human-annotated intents
from the Jazhyc/wildguard-annotated-intents dataset).

This is the best-performing distilled student we trained: it produces a short
chain-of-thought, an inferred user intent, and a final harm classification for
a given prompt.

Performance

Metric Value
OOD validation harm F1 (mean of ToxicChat train + Aegis 2.0 val) 0.7793
In-domain val harm F1 0.7611
In-domain val harm precision 0.7500
In-domain val harm recall 0.7725
In-domain val semantic similarity (intent) 0.8967

OOD F1 is averaged over lmsys/toxic-chat (toxicchat0124, train split) and
nvidia/Aegis-AI-Content-Safety-Dataset-2.0 (validation split). Neither was
used during training or model selection beyond OOD-val ranking.

Training

  • Base model: google/gemma-3-12b-it
  • Teacher: openai/gpt-oss-120b
  • Condition: human_intent (teacher is given human-annotated intent;
    student learns to reproduce reasoning + intent + harm label)
  • Learning rate: 2e-05
  • Adapter: LoRA, r=32, alpha=64, dropout=0.0, targets q/k/v/o/gate/up/down_proj
  • Quantization: 4-bit NF4 (QLoRA) during training
  • Attention: flash_attention_2 with sequence packing
  • Early stopping on combined val+test split (patience 1)

Usage

from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer

base = "google/gemma-3-12b-it"
adapter = "Jazhyc/gemma-3-12b-intent-jailbreak-classifier"

tokenizer = AutoTokenizer.from_pretrained(base)
model = AutoModelForCausalLM.from_pretrained(base, torch_dtype="bfloat16", device_map="auto")
model = PeftModel.from_pretrained(model, adapter)

messages = [
    {"role": "system", "content": "<system prompt from build_student_messages(..., 'human_intent')>"},
    {"role": "user",   "content": "<prompt to classify>"},
]
inputs = tokenizer.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device)
out = model.generate(inputs, max_new_tokens=512)
print(tokenizer.decode(out[0][inputs.shape[-1]:], skip_special_tokens=True))

The expected output format is <reasoning>...</reasoning><intent>...</intent><harm>safe|harmful</harm>.
See src/intention_jailbreak/model_generation/prompt_templates.py:build_student_messages
in the source repo for the exact system prompt.

Citation / source

Part of the intention-jailbreak research project on jailbreak detection via
verbalized intent reasoning.

Framework versions

  • PEFT 0.18.1

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-05-03Upload distilled gemma-3-12b adapter (gpt-oss-120b teacher, human_intent)5ade2e22.9 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration