← back to catalog · registered 2026-08-22 13:56

jsayyar04/modernbert-jailbreak-guard

Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/jsayyar04%2Fmodernbert-jailbreak-guard"
Response includes
  • classification unknown
  • files 7
  • hub_downloads_all_time 89
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
?
Primary method

Unclassified

No clear signals of an abliteration technique in this model.
Confidence
UNKNOWN
Why this label 1 signal
No classification signals present. This may not be an abliterated model at all - it could be a repackaging, a merge with unrelated goals, or unrelated content that mentions the term.
  • no classification signals present (no abliterated, uncensored, or known producer/method markers)
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
89
11 last 30d - stable
Likes
0
Model age
4mo ago
created 2026-06-08

Training datasets

1 of 1 in /datasets

Corpora the author lists in the model card. Datasets tracked in our /datasets catalog carry a category badge linking to the workflow stage. Others open on Hugging Face.

Downloads over time
Now95→from43↑121%
40608010043 on Jun 1095 on Oct 1195 on Oct 7JunJulAugSepOct
Jun 10 → Oct 11 · 57 snapshots · spans 123 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
safetensors modernbert text-classification llm-security prompt-injection guardrail dataset:jackhhao/jailbreak-classification base_model:answerdotai/ModernBERT-base base_model:finetune:answerdotai/ModernBERT-base license:apache-2.0 model-index region:us
Total size
571 MB
Files
7
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-06-08 15:19

Files by quantization

Auxiliary files 7 files 574 MB
model.safetensors 571 MB ddeb0b40 download
training_args.bin 5.14 KB c306c0b1 download
tokenizer.json 3.42 MB 02687599 download
README.md 5.11 KB 7c1682a4 download
config.json 2.06 KB eb7aed01 download
.gitattributes 1.48 KB a6344aac download
tokenizer_config.json 380 B 58abfae0 download

README current version from Hugging Face


license: apache-2.0
base_model: answerdotai/ModernBERT-base
tags:

  • text-classification
  • llm-security
  • prompt-injection
  • guardrail
    datasets:
  • jackhhao/jailbreak-classification
    metrics:
  • accuracy
  • precision
  • recall
  • f1
    model-index:
  • name: modernbert-jailbreak-guard
    results:
    • task:
      type: text-classification
      dataset:
      type: jackhhao/jailbreak-classification
      name: jailbreak-classification
      metrics:
      • type: accuracy
        value: 0.9962

pipeline_tag: text-classification

modernbert-jailbreak-guard

This repository hosts a sequence classification model fine-tuned to act as a pre-inference security gate for Large Language Model applications. Its core task is to inspect incoming user prompts and classify them as benign or jailbreak with minimal latency.

The architecture leverages answerdotai/ModernBERT-base as its core encoder backbone and attaches a sequence classification head optimized over full-parameter weight tuning on a T4 GPU.

Test Set Performance

Evaluated against an independent test split, this guardrail achieves strong metrics:

  • Overall Model Accuracy: 98.47%
  • Precision (Jailbreak): 97.87%
  • Recall (Jailbreak): 99.28% (Successfully intercepted 138 out of 139 malicious inputs)
  • F1-Score (Jailbreak): 98.57%
  • Macro F1-Score: 98.47%
  • False Negative Rate: 0.72% (Critical leaks minimized to under 1%)
  • False Positive Rate: 2.44% (Maintains smooth user experience with minimal false blocks)

Confusion Matrix Breakdown

  • True Negatives (Safe prompts allowed seamlessly): 120
  • True Positives (Malicious attacks neutralized): 138
  • False Positives (Safe prompts accidentally blocked): 3
  • False Negatives (Malicious payloads leaked): 1

Intended Uses and Limitations

Intended Deployment Design

This model is intended to run as a Pre-Inference Gateway Shield. It intercepts raw string requests coming from client user interfaces before they are routed to generative backends like GPT-4, Llama 3, or Claude.

Limitations and Strategy

  • Input-Only Scope: This gateway model does not monitor outgoing text generated by the core model. It needs to be coupled with an independent post-inference output alignment model to monitor for hallucinations or data leaks.
  • Defense in Depth: While highly robust, it should represent one tier of a holistic security layout including input vector blacklists and runtime system prompts.

Training and Evaluation Data

The model was fine-tuned on the balanced split of the jackhhao/jailbreak-classification dataset.

  • Training Size: 1,044 examples
  • Evaluation Size: 262 examples
  • Label Mapping: benign (0), jailbreak (1)

The dataset features a balanced distribution of classic adversarial templates, context-switching overrides, hypothetical roleplay scripts, and standard conversational strings.

Training Procedure

The model was fine-tuned using the Hugging Face Trainer library on a single cloud-hosted NVIDIA Tesla T4 GPU.

Framework and Optimization Settings

  • Optimizer: ADAMW_TORCH_FUSED (Accelerated hardware optimization kernel)
  • Precision: Mixed-precision training enabled via BF16=True (Brain Floating Point)
  • Sequence Processing: Token unpadding active to dynamically strip empty PAD tensors from memory allocation blocks.

Training hyperparameters

The following hyperparameters were used during training:

  • learning_rate: 2e-05
  • train_batch_size: 8
  • eval_batch_size: 16
  • seed: 42
  • optimizer: Use OptimizerNames.ADAMW_TORCH_FUSED with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments
  • lr_scheduler_type: linear
  • num_epochs: 4

Training results

Training Loss Epoch Step Validation Loss Accuracy Precision Jailbreak Recall Jailbreak F1 Jailbreak Macro F1 True Negatives False Positives False Negatives True Positives False Negative Rate False Positive Rate
0.0587 1.0 131 0.0544 0.9847 0.9787 0.9928 0.9857 0.9847 120 3 1 138 0.0072 0.0244
0.0027 2.0 262 0.0306 0.9885 0.9857 0.9928 0.9892 0.9885 121 2 1 138 0.0072 0.0163
0.0001 3.0 393 0.0265 0.9924 0.9928 0.9928 0.9928 0.9923 122 1 1 138 0.0072 0.0081
0.0000 4.0 524 0.0266 0.9962 1.0 0.9928 0.9964 0.9962 123 0 1 138 0.0072 0.0

Framework versions

  • Transformers 5.10.0.dev0
  • Pytorch 2.11.0+cu128
  • Datasets 4.0.0
  • Tokenizers 0.22.2

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-08Update README.md3eb6b105.1 KB
    Loading...
  2. 2026-06-08Update README.mdbe3ad6a5.2 KB
    Loading...
  3. 2026-06-08Model save50176c93 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration