← back to catalog · registered 2026-08-22 13:56

ArthT/samarth-repshift-9b-jailbreak

ArthT Qwen 9B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/ArthT%2Fsamarth-repshift-9b-jailbreak"
Response includes
  • classification unknown
  • files 4
  • benchmarks 11 entries
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
?
Primary method

Unclassified

No clear signals of an abliteration technique in this model.
Confidence
UNKNOWN
Why this label 1 signal
No classification signals present. This may not be an abliterated model at all - it could be a repackaging, a merge with unrelated goals, or unrelated content that mentions the term.
  • no classification signals present (no abliterated, uncensored, or known producer/method markers)
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
4mo ago
created 2026-05-17
Downloads over time
Now0→from0↑0%
00110 on May 200 on Oct 11MayJunJulAugSepOct
May 20 → Oct 11 · 60 snapshots · spans 144 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 1.4 UGI
Hazardous 2.4 UGI
Natural Intelligence 17.62 UGI
Political lean -12.2% UGI
Sensitive-Info 14.65 UGI
SocPol 0.9 UGI
UGI 17.27 UGI
Willingness (10) 2.2 UGI
W10-Adherence 1.5 UGI
W10-Direct 3 UGI
Writing 33.52 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en
Tags
peft safetensors safety jailbreak-detection harm-classification representation-engineering repshift lora text-classification en base_model:Qwen/Qwen3.5-9B base_model:adapter:Qwen/Qwen3.5-9B

Related

Total size
0 B
Files
4
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-05-21 18:55

Files by quantization

Auxiliary files 4 files 1.44 MB
meta.json 1.43 MB e1264c35 download
README.md 6.84 KB 0d387c64 download
args.json 1.77 KB a6ee01a3 download
.gitattributes 1.54 KB f1ca65ce download

README current version from Hugging Face


license: apache-2.0
base_model: Qwen/Qwen3.5-9B
library_name: peft
tags:

  • safety
  • jailbreak-detection
  • harm-classification
  • representation-engineering
  • repshift
  • lora
    language:
  • en
    pipeline_tag: text-classification

samarth-repshift-9b-jailbreak

A representation-shift jailbreak detection judge built on Qwen/Qwen3.5-9B,
specialized for detecting real harmful content in adversarially-attacked LLM
responses
.

Companion to ArthT/samarth-repshift-9b-v1.
The two variants trade off between two failure modes:

Variant Strength Weakness Use when
samarth-repshift-9b-jailbreak (this) Discriminates real harm from confident-nonsense on adversarially-attacked outputs (AUROC 0.823 on Active Robustness) Higher false-positive rate on honeypot-style content (0.76 @ p=0.5) Detecting jailbreaks against defended models
samarth-repshift-9b-v1 Almost-zero false positives on confident-nonsense honeypots (0.009 @ p=0.5) Lower discrimination on adversarially-attacked harm (AUROC 0.624 on Active Robustness) Filtering confident-but-fake content from training data, red-team validation

Headline result

On the Active Robustness benchmark (7,454 (prompt, response) pairs from
8 defense methods × 4 attack types applied to Llama-3-8B-Instruct, with
ground-truth labels from a gpt-5.4-mini + grok-4.3 dual-rater panel using
a confident-nonsense-aware rubric), discriminating HARMFUL_REAL from
(CONFIDENT_NONSENSE + REFUSAL):

Judge AUROC
qwen3guard-gen-8b 0.859
samarth-repshift-9b-jailbreak (this) 0.823
tf-v08-respmask-s43 (sibling variant) 0.764
llama-guard-3 0.671
Phase 1 winner (sibling) 0.638
samarth-repshift-9b-v1 (honeypot specialist) 0.624

This judge sits within 0.04 AUROC of the strongest production safety
classifier (qwen3guard-gen-8b) on the adversarial detection task, while
remaining a representation-shift judge (2× inference cost, no learned head,
explainable score).

On the harder HARMFUL_REAL vs CONFIDENT_NONSENSE discrimination subtask
(separating real-harm from plausible-looking fake content), this judge gets
AUROC 0.630 — second only to qwen3guard (0.711) and well ahead of all
other open-source safety judges.

Architecture

Same Δ_harm + Δ_hp_axis representation-shift architecture as the
companion v1 variant:

d_drift   = mean over response tokens of ‖h_with(t) − h_without(t)‖
score_raw = d_drift @ Δ_harm − λ_hp · (d_drift @ Δ_hp_axis)
p_harmful = sigmoid(platt_a · score_raw + platt_b)

The difference vs v1 is in training:

  • v1 trained with an external honeypot corpus (4,154 records) → tightens the
    honeypot defense, weakens generalization to adversarial defenses.
  • This variant trained without external honeypots → broader generalization
    on real adversarial outputs, less aggressive honeypot defense.

Hyperparameters specific to this checkpoint:

  • λ_hp_axis = 2.0 (twice the v1 weight on the hp-axis penalty)
  • EMA decay for Δ_hp_axis = 0.999 (very slow, stable direction)
  • Phase 1 corpus only (no external honeypot augmentation)
  • Other hyperparameters identical to v1 (see Training below)

Training

  • Base: Qwen/Qwen3.5-9B + LoRA r=16, α=32, dropout=0.05
  • Corpus: 9,873 rows = 2,528 harmful prompts (HarmBench / AdvBench /
    JailbreakBench seed) + 2,528 Qwen3-4B-Instruct refusals (same prompts) +
    497 in-house honeypots + 4,320 gray/borderline records
  • Loss: `α·shift_unsafe + β·KL_safe + γ·anchor_honeypot + λ_hp·anchor_gray
    • λ_hp_axis·Δ_hp_axis`
  • Margins: m_b=5.0, m_h=15.0, m_b_dir=0.5, m_h_dir=1.5
  • Weights: α=0.5, β=0.4, γ=0.9, λ_hp=3.0, λ_hp_axis=2.0
  • EMA decays: Δ_harm=0.95, Δ_hp_axis=0.999
  • Rep layers: 22-30 of 36 (top 30%)
  • Schedule: patience 15, eval every 50 steps, lr 1e-4
  • Seed: 42

Recommended hybrid usage

For broad-coverage production safety, pair with harmbench-llama at w=0.6
(see the v1 model card for full sweep). For maximum jailbreak coverage,
pair with qwen3guard-gen-8b — both signals are correlated on real harm,
and the multiplicative form sigmoid(qwen3guard) × sigmoid(this) should
inherit qwen3guard's recall (0.92) with substantially lower false-positive
rate.

Per-dataset scores (this variant, alone)

Dataset AUROC Note
combined_fp_eval (V1 benchmark) 0.848
do_not_answer 0.725
sorry_bench_human 0.680
judge_stress_test 0.548
wildguardmix_test 0.585
beavertails_30k_test 0.604
coinflip_for_safety (GCG/BoN/Pair) 0.461
active_robustness (kitchen-sink jailbreak benchmark) 0.823 this variant's strength
4-axis quadrant eval 0.935 strong adversarial-cell discrimination
xstest_gpt4 0.768
aegis-v2 0.608

Inference

Repshift is a 2-pass forward (with and without adapter on the same response).
See judges/adapters/samarth_repshift.py in the
Robust-jailbreak-judges
repo for the reference implementation. meta.json carries delta_harm and
delta_hp_axis EMA directions, plus calibrated Platt parameters a and b.

Limitations

  • 2× inference cost (forward with adapter + forward without)
  • Higher false-positive rate on confident-nonsense honeypots (~0.76 @ p=0.5
    vs v1's 0.009). If your input stream contains many honeypot-style
    responses, use v1 or a hybrid.
  • Trained on the Phase 1 corpus only (no external honeypot augmentation),
    so the honeypot-specific defense is weaker than v1.
  • English-only training data.
  • LoRA-only; the base Qwen3.5-9B weights must be loaded separately.

Citation

@misc{samarth-repshift-9b-jailbreak,
  title  = {Representation-Shift Judges for Adversarial Jailbreak Detection},
  author = {Singh, Arth and AIM Intelligence},
  year   = {2026},
  note   = {Robust Jailbreak Judges project, repshift Qwen3.5-9B
            λ_hp_axis=2.0 variant, seed 42}
}

⚠️ Correction notice (2026-05-22)

Earlier versions of this card claimed AUROC 0.823 on Active Robustness vs
qwen3guard 0.859. Those numbers were computed with a buggy record-level
join. Corrected numbers (HARMFUL_REAL vs CN+REFUSAL, dual-rater gold labels,
2,177 records):

Judge AUROC full AUROC CN-only
samarth-qwen35-9b + system prompt 0.936 0.871
qwen3guard-gen-8b 0.858 0.605
this (samarth-repshift-9b-jailbreak) 0.627 0.619

For Active Robustness style adversarial detection, prefer
samarth-qwen35-9b. This
variant is retained for research reproducibility on the project repo's
internal λ-sweep ablations.

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-05-21Add 2026-05-22 correction notice — AR AUROCs were computed with buggy join867f0cc6.8 KB
    Loading...
  2. 2026-05-17Initial release: samarth-repshift-9b-jailbreak (lambda-sweep le-l20-e0999, se...a1d24e16.1 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration