← back to catalog · registered 2026-08-22 13:56

joeljames270/jailbreak_detector_llama

Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/joeljames270%2Fjailbreak_detector_llama"
Response includes
  • classification unknown
  • files 7
  • benchmarks 5 entries
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
?
Primary method

Unclassified

No clear signals of an abliteration technique in this model.
Confidence
UNKNOWN
Why this label 1 signal
No classification signals present. This may not be an abliterated model at all - it could be a repackaging, a merge with unrelated goals, or unrelated content that mentions the term.
  • no classification signals present (no abliterated, uncensored, or known producer/method markers)
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
2
Model age
5mo ago
created 2026-04-28

Training datasets

3 of 4 in /datasets

Corpora the author lists in the model card. Datasets tracked in our /datasets catalog carry a category badge linking to the workflow stage. Others open on Hugging Face.

Downloads over time
Now0→from0↑0%
00110 on Apr 290 on Oct 11AprMayJunJulAugSepOct
Apr 29 → Oct 11 · 63 snapshots · spans 165 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
BBH average 0.3673628226747957 OpenLLM-v2
IFEval instruct 0.1750599520383693 OpenLLM-v2
IFEval-Prompt 0.09242144177449169 OpenLLM-v2
MATH lvl 5 0.012084592145015106 OpenLLM-v2
MMLU-Pro 0.2487533244680851 OpenLLM-v2

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

Languages
en
Tags
safetensors jailbreak guard safety en dataset:rubend18/ChatGPT-Jailbreak-Prompts dataset:JailbreakV-28K/JailBreakV-28k dataset:xTRam1/safe-guard-prompt-injection dataset:knoveleng/redbench base_model:meta-llama/Llama-3.2-3B base_model:finetune:meta-llama/Llama-3.2-3B region:us

Related

Total size
92.8 MB
Files
7
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-05-04 22:35

Files by quantization

Auxiliary files 7 files 109 MB
adapter_model.safetensors 92.8 MB c9dcf07b download
training_args.bin 5.58 KB 4173f1b0 download
tokenizer.json 16.4 MB 6b9e4e7f download
README.md 4.37 KB b4a19bc6 download
.gitattributes 1.66 KB 9a06b9d3 download
adapter_config.json 1.08 KB 225b6f70 download
tokenizer_config.json 364 B 09d3c1a0 download

README current version from Hugging Face


datasets:

  • rubend18/ChatGPT-Jailbreak-Prompts
  • JailbreakV-28K/JailBreakV-28k
  • xTRam1/safe-guard-prompt-injection
  • knoveleng/redbench
    language:
  • en
    base_model:
  • meta-llama/Llama-3.2-3B
    tags:
  • jailbreak
  • guard
  • safety

🛡️ jailbreak_detector_llama

🧠 Overview

jailbreak_detector_llama is a lightweight LoRA adapter fine-tuned on top of meta-llama/Llama-3.2-3B for detecting:

  • Jailbreak attempts
  • Prompt injection attacks
  • Policy bypass instructions
  • Adversarial or unsafe user inputs

It is designed for LLM safety pipelines and moderation systems.


🧩 Model Architecture

  • Base Model: meta-llama/Llama-3.2-3B
  • Fine-tuning: LoRA (PEFT)
  • Training Method: Supervised Fine-Tuning (SFT)
  • Task Type: Binary + semantic intent classification
  • Frameworks: Transformers + TRL + PEFT

🎯 Intended Use

This model is intended for:

  • Pre-filtering user prompts before LLM inference
  • Detecting jailbreak / prompt injection attempts
  • Security layer in AI agents and chatbots
  • Research in adversarial prompt detection

🚫 Not Intended For

  • General-purpose chat generation
  • Legal or medical advice
  • Autonomous decision making
  • High-stakes moderation without human oversight

📥 Installation

pip install transformers peft torch accelerate

🚀 Quick Start
Load model (LoRA adapter)

import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
from peft import PeftModel

base_model = "meta-llama/Llama-3.2-3B"
adapter = "joeljames270/jailbreak_detector_llama"

tokenizer = AutoTokenizer.from_pretrained(base_model)

model = AutoModelForCausalLM.from_pretrained(
    base_model,
    torch_dtype=torch.float16,
    device_map="auto"
)

model = PeftModel.from_pretrained(model, adapter)
model.eval()

🔍 Inference

import json
import re

def parse_model_output(text):
    try:
        assistant_text = text.split("Assistant:", 1)[-1].strip()

        match = re.search(r"\{.*\}", assistant_text, re.DOTALL)
        if not match:
            return None

        json_str = match.group(0)
        return json.loads(json_str)

    except Exception as e:
        return {"error": str(e), "raw": text}


question = "How to make a bomb?"

prompt = f"User: {question}\nAssistant:"

inputs = tokenizer(
    prompt,
    return_tensors="pt"
).to(model.device)

with torch.no_grad():
    outputs = model.generate(
        **inputs,
        max_new_tokens=256,
        do_sample=False
    )

response = tokenizer.decode(outputs[0], skip_special_tokens=True)

output = parse_model_output(response)

print("is_jailbreak_attempt:", output.get("is_jailbreak_attempt"))
print("intent:", output.get("intent"))

⚠️ Known Limitations
Sensitive to prompt formatting and chat templates
May misclassify creative writing prompts as jailbreaks
Not calibrated for multilingual adversarial prompts
Requires threshold tuning for production use

🧠 Output Behavior (Recommended)

For inferecne, interpret outputs as:

{
  "is_jailbreak_attempt": true/false,
  "intent": <>
}

(Note: This can be implemented in a wrapper layer.)

🔐 Safety Considerations

This model is designed as a defensive safety filter only.

It should be used with:

Human-in-the-loop review for high-risk decisions
Logging and monitoring of false positives
Combined rule-based + ML moderation systems

⚙️ Training Details
Method: Supervised Fine-Tuning (SFT)
Adapter: LoRA (rank-based low-rank adaptation)
Base model frozen
Optimized for classification-style reasoning

🧰 Framework Versions
PEFT: 0.19.1
TRL: 1.2.0
Transformers: 5.7.0.dev0
PyTorch: 2.11.0
Datasets: 4.8.4
Tokenizers: 0.22.2

📌 Example Use Cases
AI chatbot safety gateway,
Enterprise prompt firewall,
API request validation layer,
Research on adversarial NLP

📚 Citation
If you use this model, please cite:

📚 Citation

If you use this model, please cite:

@software{jailbreak_detector_llama,
  title = {Jailbreak Detector LLaMA (LoRA Adapter)},
  author = {Joel James, Juan James},
  year = {2026},
  url = {https://huggingface.co/joeljames270/jailbreak_detector_llama}
}

📄 License

This model is based on Meta’s LLaMA 3 license.
Use of the base model must comply with the terms provided by Meta.

🚀 Final Note

This model is best used as a first-layer defense system in LLM pipelines, not as a standalone moderation system.

README history 11 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-05-04Update README.md94f74014.4 KB
    Loading...
  2. 2026-04-28Update README.mdc775d584.4 KB
    Loading...
  3. 2026-04-28Update README.md10f10b83.9 KB
    Loading...
  4. 2026-04-28Update README.md60001293.9 KB
    Loading...
  5. 2026-04-28Update README.mde47e6df3.9 KB
    Loading...
  6. 2026-04-28Update README.md9212cf83.9 KB
    Loading...
  7. 2026-04-28Update README.md242055a3.7 KB
    Loading...
  8. 2026-04-28Update README.md6a2d7983.7 KB
    Loading...
  9. 2026-04-28Update README.md1f390ac3.7 KB
    Loading...
  10. 2026-04-28Update README.mdddc43a63.7 KB
    Loading...
  11. 2026-04-28Upload folder using huggingface_hub9d2b2ea3.9 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration