← back to catalog · registered 2026-08-22 13:56

nvidia/NemoGuard-JailbreakDetect

Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/nvidia%2FNemoGuard-JailbreakDetect"
Response includes
  • classification unknown
  • files 5
  • hub_downloads_all_time 10,188
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
?
Primary method

Unclassified

No clear signals of an abliteration technique in this model.
Confidence
UNKNOWN
Why this label 1 signal
No classification signals present. This may not be an abliterated model at all - it could be a repackaging, a merge with unrelated goals, or unrelated content that mentions the term.
  • no classification signals present (no abliterated, uncensored, or known producer/method markers)
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
10K
305 last 30d - cooling
Likes
37
Model age
21mo ago
created 2025-01-14
Downloads over time
Now10.3K→from11↑93,718%
03.8K7.6K11.4K11 on Jan 15, 202510.3K on Oct 11Jan '25Apr '25Jul '25Oct '25JanAprJulOct
Jan 15, 2025 → Oct 11 · 130 snapshots · spans 634 days

Metadata

Tags
onnx arxiv:2412.01547 region:us
Total size
41.9 MB
Files
5
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-04-02 18:23

Files by quantization

Auxiliary files 5 files 46.5 MB
snowflake.onnx 41.9 MB 2309ca09 download
snowflake.pkl 4.47 MB 45c2bc8d download
config.json 125 KB c9320293 download
README.md 3.74 KB 73fe04b3 download
.gitattributes 1.48 KB a6344aac download

README current version from Hugging Face

Model Overview

Description:

NemoGuard JailbreakDetect was developed to detect attempts to jailbreak large language models.
This model is ready for commercial use.

License/Terms of Use:

NVIDIA Open Model License

Reference(s):

Improved Large Language Model Jailbreak Detection via Pretrained Embeddings

Model Architecture:

Architecture Type: Random Forest

Network Architecture: N/A

Input:

Input Type(s): Text Embedding

Input Parameters: 768 dimensional vector

Input Format(s): Vector

Other Properties Related to Input: Must be an output from the corresponding embedding model, snowflake-arctic-m-long.

Output:

Output Type(s): Classification, Probability

Output Format: Bool, Float

Output Parameters: 1D

Other Properties Related to Output: N/A

Software Integration:

Runtime Engine(s):

  • Not Applicable (N/A)

Supported Hardware Microarchitecture Compatibility:

  • x86
  • x64

[Preferred/Supported] Operating System(s):

  • Windows
  • MacOS
  • Linux

Model Version(s):

NemoGuard-JailbreakDetect-v1.0: Jailbreak detection model using Snowflake-arctic-embed-m embeddings

Training, Testing, and Evaluation Datasets:

Training Dataset:

A combination of three open datasets, mixed together, de-duplicated, and reviewed for data quality.
Jailbreak data was augmented with the use of garak.
The datasets used are outlined below:

Advbench

Link: https://github.com/thunlp/Advbench

** Data Collection Method by dataset

  • [Automated]

** Labeling Method by dataset

  • [Automated]

Properties:
520 entries, all comprised of jailbreak attempts.

Wildjailbreak

Link: https://huggingface.co/datasets/allenai/wildjailbreak

** Data Collection Method by dataset

  • Hybrid: Automated, Synthetic

** Labeling Method by dataset

  • [Automated]

Properties:
6387 total entries: 5721 benign prompts, 666 jailbreak attempts

jackhao/jailbreak-classification

Link: https://huggingface.co/datasets/jackhhao/jailbreak-classification

** Data Collection Method by dataset

  • [Automated]

** Labeling Method by dataset

  • [Automated]

Properties:
1306 total entries: 640 benign prompts, 666 jailbreak attempts

Testing Dataset:

A stratified subset (20%) of the aggregate dataset was used for testing.

Evaluation Dataset:

Evaluated on JailbreakHub.

Model F1 Score False Positive Rate False Negative Rate
NemoGuard JailbreakDetect 0.9601 0.0042 0.0435

Inference:

Engine: N/A

Test Hardware:

  • RTX A6000
  • A100

Ethical Considerations:

NVIDIA believes Trustworthy AI is a shared responsibility, and we have established policies and practices to enable development for a wide array of AI applications.
When downloaded or used in accordance with our terms of service, developers should work with their internal model team to ensure this model meets requirements for the relevant industry and use case and addresses unforeseen product misuse.

For more detailed information on ethical considerations for this model, please see the Model Card++ Explainability, Bias, Safety & Security, and Privacy Subcards [Insert Link to Model Card++ here].
Please report security vulnerabilities or NVIDIA AI Concerns here

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2025-01-15Rename overview.md to README.md (#1)2e0d9a13.7 KB
    Loading...

Discussions 5 threads

  1. 2026-09-06snowflake.onnx gives different classifications from snowflake.pklopen1 💬#5
    Loading...
  2. 2026-09-06snowflake.onnx gives different classifications from snowflake.pklclosed1 💬#4
    Loading...
  3. 2026-04-02I don't understandopen1 💬#3
    Loading...
  4. 2025-04-03PRUpdate README.mdopen1 💬#2
    Loading...
  5. 2025-01-14PRRename overview.md to README.mdmerged1 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration