language: en
license: apache-2.0
tags:
- safety
- guardrail
- text-classification
- input-filtering
- transformers
pipeline_tag: text-classification
library_name: transformers
base_model: microsoft/deberta-v2
Jazzmine Input Safeguard v2
Model Summary
Jazzmine Input Safeguard v2 is a fine-tuned DeBERTa v2–based Transformer model designed to analyze and filter user inputs for safety, policy compliance, and undesired or potentially harmful content before they are passed to downstream systems such as large language models, AI agents, or automated workflows.
The model is intended to operate as an input-level guardrail, reducing the likelihood that unsafe or adversarial prompts reach sensitive components of an AI system.
Base Model
This model is fine-tuned from DeBERTa v2 (microsoft/deberta-v2-*), selected for its strong contextual representations and high performance on text classification tasks.
Intended Use
This model is intended for:
- Input validation in conversational AI systems
- Safety filtering before LLM invocation
- Guardrail enforcement in agent-based architectures
- Pre-processing and risk assessment of user-generated text
This model is not intended for:
- Fully autonomous moderation decisions
- Legal, medical, or compliance judgments
- Use in high-stakes environments without additional safeguards or human oversight
Training Data
The model was fine-tuned on a mixture of real-world and synthetic data.
- Real data consists of anonymized, curated examples of safe and unsafe user inputs.
- Synthetic data was programmatically generated to improve coverage of edge cases, adversarial phrasing, and rare failure modes that are underrepresented in real data.
No personally identifiable information (PII) was used in the training process.
Training Procedure
- Task: Supervised text classification
- Fine-tuning approach: Standard Transformer fine-tuning
- Loss function: Cross-entropy
- Framework: Hugging Face Transformers
- Weights format:
safetensors
Exact dataset composition, prompts, and hyperparameters are not publicly disclosed.
How to Use
from transformers import AutoModel, AutoTokenizer
model = AutoModel.from_pretrained(
"nourmedini1/jazzmine-input-safeguard-v2"
)
tokenizer = AutoTokenizer.from_pretrained(
"nourmedini1/jazzmine-input-safeguard-v2"
)
inputs = tokenizer(
"Example user input text",
return_tensors="pt"
)
outputs = model(**inputs)
Limitations
The model may produce false positives or false negatives.
Performance depends on similarity between inference inputs and the training distribution.
Adversarial or highly novel inputs may bypass classification.
The model should be used as part of a layered safety strategy rather than as a standalone solution.
Ethical Considerations
This model is intended to support safety and moderation workflows, not replace human judgment.
Biases present in the training data — including biases introduced through synthetic data generation — may affect predictions and should be considered when deploying the model.
Users are responsible for ensuring that the model is applied in ways that align with applicable laws, regulations, and ethical guidelines.
License
Apache License 2.0