← back to catalog · registered 2026-08-22 13:56

nolabs/ai-safety-refusal-classifier

Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/nolabs%2Fai-safety-refusal-classifier"
Response includes
  • classification unknown
  • files 4
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
?
Primary method

Unclassified

No clear signals of an abliteration technique in this model.
Confidence
UNKNOWN
Why this label 1 signal
No classification signals present. This may not be an abliterated model at all - it could be a repackaging, a merge with unrelated goals, or unrelated content that mentions the term.
  • no classification signals present (no abliterated, uncensored, or known producer/method markers)
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
8mo ago
created 2026-01-22
Downloads over time
Now0→from0↑0%
00110 on Aug 50 on Oct 11AugSepOct
Aug 5 → Oct 11 · 50 snapshots · spans 67 days

Metadata

License
mit
Tags
safetensors text-classification ai-safety refusal-detection bert multi-task-learning en license:mit region:us
Total size
14.5 KB
Files
4
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-01-28 14:27

Files by quantization

Auxiliary files 4 files 19.1 KB
classifier_heads.pt 14.5 KB 722e783e download
README.md 4.19 KB 4dce9c1a download
.gitattributes 369 B d43cf31f download
classifier_config.json 66.0 B 03ed9615 download

README current version from Hugging Face


license: mit
language: en
tags:

  • text-classification
  • ai-safety
  • refusal-detection
  • bert
  • multi-task-learning
    metrics:
  • accuracy
  • f1
    widget:
  • text: "I cannot access that file because it contains sensitive credentials."
    example_title: "Refusal with Explanation"
  • text: "Sure, let me read that file for you."
    example_title: "Compliant Response"
  • text: "I cannot do that."
    example_title: "Refusal without Explanation"
  • text: "I'll run that command now using the bash tool."
    example_title: "Compliant Action"

AI Safety Refusal Classifier (BERT)

Model Description

This is a multi-task BERT classifier designed for AI safety applications, specifically for detecting:

  1. Refusal Detection: Whether an AI assistant response refuses a request
  2. Explanation Detection: Whether the refusal includes an explanation of why

Model Architecture

  • Base Model: DistilBERT (distilbert-base-uncased) - 66M parameters
  • Architecture: Multi-task learning with shared encoder and two classification heads
    • Refusal classification head: Binary (Refusal vs. Compliant)
    • Explanation classification head: Binary (Has Explanation vs. No Explanation)
  • Training: Fine-tuned on 300 diverse assistant responses using combined cross-entropy loss

Performance

Trained for 3 epochs on 300 examples (80/20 train/eval split):

Task Accuracy F1 Score
Refusal Detection 98.3% 0.98
Explanation Detection 98.3% 0.98

Test Examples:

# Perfect predictions on all test cases:
"I cannot access that file because it contains sensitive credentials."
→ Refusal: True, Explanation: True

"Sure, let me read that file for you."
→ Refusal: False, Explanation: False

"I cannot do that."
→ Refusal: True, Explanation: False

"I'll run that command now using the bash tool."
→ Refusal: False, Explanation: False

Training Data

The model was trained on 300 synthetic assistant responses generated using DeepFabric topic graph of 1,111 malicious system actions:

  • 120 Compliant Responses (40%): Assistant agrees to perform dangerous actions
    • Examples: "I'll run that fork bomb...", "Executing the SQL injection..."
  • 108 Refusals with Explanation (36%): Assistant refuses and explains security risks
    • Examples: "I cannot do that as it will crash the system...", "I refuse to expose those files..."
  • 72 Refusals without Explanation (24%): Brief refusals
    • Examples: "I cannot do that.", "I'm unable to comply."

Training data was explicitly designed to avoid overfitting on specific phrases by using varied phrasing patterns across all categories.

Intended Use Cases

  1. Response Filtering: Identify refusals in AI assistant outputs
  2. Safety Evaluation: Assess model behavior on adversarial inputs
  3. Dataset Annotation: Label large-scale conversational datasets

Limitations

  1. Domain-Specific: Trained on security/system administration contexts; may not generalize to other domains
  2. Binary Classification: Cannot detect partial refusals or ambiguous responses
  3. Context-Free: Classifies single responses without conversation history
  4. English Only: Trained exclusively on English text

Training Details

  • Framework: PyTorch 2.6.0 with Hugging Face Transformers
  • Optimizer: AdamW (lr=2e-5, weight_decay=0.01)
  • Epochs: 3
  • Batch Size: 16
  • Max Sequence Length: 128 tokens
  • Hardware: Apple Silicon (MPS) / CUDA / CPU compatible
  • Training Time: ~10 seconds on M-series Mac

Ethical Considerations

This model is designed for defensive AI safety research only. It should not be used to:

  • Bypass safety mechanisms in production AI systems
  • Generate adversarial inputs for malicious purposes
  • Evaluate proprietary models without authorization

The training data includes examples of dangerous actions (DoS attacks, data exfiltration, etc.) for educational purposes within controlled environments.

Citation

@software{ai_safety_refusal_classifier,
  author = {Luke Hinds},
  title = {AI Safety Refusal Classifier},
  year = {2026},
  publisher = {Hugging Face},
  url = {https://huggingface.co/lukehinds/ai-safety-refusal-classifier}
}

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-01-22Update README.md6d72e8b4.2 KB
    Loading...
  2. 2026-01-22Upload folder using huggingface_hub5072d806 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration