← back to catalog · registered 2026-08-22 13:56

idanpers/JailBreakModel

Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/idanpers%2FJailBreakModel"
Response includes
  • classification unknown
  • files 9
  • hub_downloads_all_time 32
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
?
Primary method

Unclassified

No clear signals of an abliteration technique in this model.
Confidence
UNKNOWN
Why this label 1 signal
No classification signals present. This may not be an abliterated model at all - it could be a repackaging, a merge with unrelated goals, or unrelated content that mentions the term.
  • no classification signals present (no abliterated, uncensored, or known producer/method markers)
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
32
28 last 30d - active
Likes
0
Model age
23mo ago
created 2024-11-04
Downloads over time
Now45→from0↑0%
01282563840 on Oct 30, 202445 on Oct 11349 on Jan 21Oct '24Feb '25Jun '25Oct '25FebJunOct
Oct 30, 2024 → Oct 11 · 141 snapshots · spans 711 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
transformers pytorch electra text-classification generated_from_trainer base_model:google/electra-base-discriminator base_model:finetune:google/electra-base-discriminator license:apache-2.0 endpoints_compatible region:us

Related

Total size
418 MB
Files
9
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2024-11-04 20:44

Files by quantization

Auxiliary files 9 files 418 MB
pytorch_model.bin 418 MB 3c8b555a download
training_args.bin 5.05 KB 95b2dc32 download
vocab.txt 226 KB fb140275 download
README.md 4.52 KB e2a1a48b download
.gitattributes 1.48 KB a6344aac download
tokenizer_config.json 1.22 KB 972aa439 download
config.json 857 B 1d41f58a download
requirements.txt 267 B 7c684efe download
special_tokens_map.json 125 B a8b3208c download

README current version from Hugging Face


library_name: transformers
license: apache-2.0
base_model: google/electra-base-discriminator
tags:

  • generated_from_trainer
    model-index:
  • name: JailBreakModel
    results: []

ELECTRA Trainer for Prompt Injection Detection

colab notebook : https://colab.research.google.com/drive/11da3m_gYwmkURcjGn8_kp23GiM-INDrm?usp=sharing

Overview

This repository contains a fine-tuned ELECTRA model designed for detecting prompt injections in AI systems. The model classifies input prompts into two categories: benign and jailbreak. This approach aims to enhance the safety and robustness of AI applications.

Approach and Design Decisions

The primary goal of this project was to create a reliable model that can distinguish between safe and potentially harmful prompts. Key design decisions included:

  • Model Selection: I chose the ELECTRA model due to its efficient training process and strong performance on text classification tasks. ELECTRA's architecture allows for effective learning from limited data, which is crucial given the specificity of the task.

  • Data Preparation: A custom dataset was curated, consisting of diverse prompts labeled as either benign or jailbreak. The dataset aimed to balance both classes to mitigate biases during training.

  • Long Inputs: To handle prompts exceeding the maximum input length of the ELECTRA model, I used truncation. Even though there was a data loss , the model still managed to classify the prompt correctly.

Model Architecture and Training Strategy

The model is based on the google/electra-base-discriminator architecture. Here’s an overview of the training strategy:

  1. Tokenization: I utilized the ELECTRA tokenizer to prepare input prompts. Padding and truncation were handled to ensure uniform input size.

  2. Training Configuration:

    • Learning Rate: Set to 5e-05 for stable convergence.
    • Batch Size: A batch size of 16 was chosen to balance training speed and memory usage.
    • Epochs: The model was trained for 2 epochs to prevent overfitting while still allowing sufficient learning from the dataset.
  3. Evaluation: The model’s performance was evaluated on a validation set, focusing on metrics such as accuracy, precision, recall, and F1 score.

Key Results and Observations

  • The model achieved a high accuracy rate on the validation set, indicating its effectiveness in distinguishing between benign and harmful prompts.

Instructions for Running the Inference Pipeline

To run the inference pipeline for classifying prompts, follow these steps:

  1. Install Dependencies:
    Ensure you have Python installed, and then install the required libraries using pip:

    pip install transformers datasets torch
    
# Load model directly
from transformers import AutoTokenizer, AutoModelForSequenceClassification

Tokenizer = AutoTokenizer.from_pretrained("idanpers/JailBreakModel")
model = AutoModelForSequenceClassification.from_pretrained("idanpers/JailBreakModel")


training_args = TrainingArguments(
  output_dir="./results",
  per_device_train_batch_size=16,
  per_device_eval_batch_size=16,
  report_to="none",  # Disable W&B
  save_safetensors=False,
)




# Create Trainer instance
trainer = Trainer(
  model=model,
  args=training_args,
  tokenizer=tokenizer,
)



use:
def classify_prompt(prompt):
# Error handling for empty input
if not isinstance(prompt, str) or prompt.strip() == "":
    return {"error": "Invalid input. Please provide a non-empty text prompt."}

# Tokenize the input prompt and convert to dataset format expected by trainer.predict
inputs = Tokenizer(prompt, return_tensors="pt", padding=True, truncation=True)
dataset = Dataset.from_dict({"input_ids": inputs["input_ids"], "attention_mask": inputs["attention_mask"]})

# Use trainer.predict to classify
prediction_output = trainer.predict(dataset)

# Get the softmax probabilities for confidence scores
probs = torch.softmax(torch.tensor(prediction_output.predictions), dim=1).cpu().numpy()
confidence = np.max(probs)
pred_label = np.argmax(probs, axis=1)[0]

# Map prediction to label
label = "PROMPT_INJECTION" if pred_label == 1 else "BENIGN"

return {"label": label, "confidence": confidence}

#Accept input from the user and classify it
prompt = input("Enter a prompt for classification: ")
result = classify_prompt(prompt)

#Check for errors before accessing the classification result
if "error" in result:
print(f"Error: {result['error']}")
else:
print(f"Classification Result: {result['label']}")
print(f"Confidence Score: {result['confidence']:.2f}")

README history 5 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2024-11-04Update README.mde1542484.5 KB
    Loading...
  2. 2024-11-04Update README.md6b84f494.3 KB
    Loading...
  3. 2024-11-04Update README.md8aa027c4.2 KB
    Loading...
  4. 2024-11-04Update README.mde2b88be4.2 KB
    Loading...
  5. 2024-11-04JBCT7a60d1f1.1 KB
    Loading...

Discussions 1 thread

  1. 2024-11-05PRAdding `safetensors` variant of this modelopen1 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration