← back to catalog · registered 2026-08-22 13:56

prithivMLmods/VideoGuard-Qwen3.5-4B-Safety-RL-Uncensored

prithivMLmods Qwen 4.5B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/prithivMLmods%2FVideoGuard-Qwen3.5-4B-Safety-RL-Uncensored"
Response includes
  • classification m-uncensored
  • files 9
  • benchmarks 11 entries
  • hub_downloads_all_time 208
  • author_summary 98 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
208
66 last 30d - stable
Likes
1
Descendants
1
in 1 direct fork
Model age
2mo ago
created 2026-08-10

Training datasets

Corpora the author lists in the model card. Datasets tracked in our /datasets catalog carry a category badge linking to the workflow stage. Others open on Hugging Face.

Downloads over time
Now234→from0↑0%
0861722570 on Aug 12234 on Oct 11234 on Oct 10AugSepOct
Aug 12 → Oct 11 · 49 snapshots · spans 60 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 0.9 UGI
Hazardous 1.2 UGI
Natural Intelligence 13.45 UGI
Political lean -17.3% UGI
Sensitive-Info 11.73 UGI
SocPol 1.5 UGI
UGI 15.32 UGI
Willingness (10) 2.2 UGI
W10-Adherence 1.5 UGI
W10-Direct 3 UGI
Writing 29.68 UGI

Genealogy 1 direct fork

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 2 formats · 3K downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
apache-2.0
Languages
en
Tags
transformers safetensors qwen3_5 image-text-to-text text-generation-inference reinforcement-learning video-text-to-text alignment-training RLHF RFT video-understanding video-classification

Related

Total size
8.46 GB
Files
9
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-11 14:11

Files by quantization

Auxiliary files 9 files 8.47 GB
model.safetensors 8.46 GB 11886bb4 download
tokenizer.json 19.1 MB 87a7830d download
README.md 14.6 KB 38076809 download
chat_template.jinja 9.62 KB d45b0278 download
config.json 2.76 KB 5a7f024d download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.27 KB 8af8110f download
tokenizer_config.json 1.14 KB 43f5bc3a download
generation_config.json 116 B 48697309 download

README current version from Hugging Face


license: apache-2.0
base_model:

  • Qwen/Qwen3.5-4B
    language:
  • en
    pipeline_tag: video-text-to-text
    library_name: transformers
    tags:
  • text-generation-inference
  • reinforcement-learning
  • video-text-to-text
  • alignment-training
  • RLHF
  • RFT
  • video-understanding
  • video-classification
  • video-safety
  • content-safety
  • content-moderation
  • safety-classifier
  • guardrail
    datasets:
  • prithivMLmods/OpenVideo-Scene-Reasoning
  • PALM-Lab/vid-guard-rlhf-unsafe

VideoGuard-4B

demo
think
The user wants me to classify an image and provide a specific output format. The image shows a young boy peeking over a table with a handgun on it. I need to check for safety categories: - C1 (Sexual Content): No. - C2 (Harassment & Bullying): No. - C3 (Threats, Violence & Harm): Yes, there is a gun. - C4 (False & Deceptive Information): No. - C5 (Illegal/Regulated Activities): Yes, a gun is a regulated item, but usually, this category is for illegal drugs or trafficking. A gun is a weapon, so C3 is more appropriate. However, the prompt says "C5 covers illegal drugs, contraband, weapons trafficking, or other clearly regulated or unlawful activity." A gun is a weapon, so it could fit C5 as well, but C3 is more direct for "weapons pointed at people, physical violence, threats of harm, self-harm, or injury to people or animals." The gun is on a table, not pointed at a person, but it is a weapon. Let's look at the description. The gun is on a table, a boy is behind it. It's a staged photo. - C6 (Hate symbols, slurs, or content promoting a hateful or extremist ideology): No.

So, C3 is definitely relevant because a weapon is present. C5 might be relevant if we consider guns as regulated activity. The prompt says "C5 covers... weapons trafficking". A gun on a table might be considered a weapon. I will include C3 and maybe C5 if it fits well. The prompt asks for "which category codes apply, if any, and briefly why".
The image is "safe or not" - this is the user's question. I need to answer that in the output.

Line 1: DESCRIPTION: A young boy peeks over a table with a handgun resting on the surface in front of him.
Line 2: EXPLANATION: C3 (Threats, Violence & Harm) is present due to the handgun on the table. C5 (Illegal/Regulated Activities) is relevant as a firearm is a regulated item.
Line 3: GUARDRAIL: {"C1(Sexual Content)": false, "C2(Harassment & Bullying)": false, "C3(Threats, Violence & Harm)": true, "C4(False & Deceptive Information)": false, "C5(Illegal/Regulated Activities)": true, "C6(Hateful Content & Extremism)": false}

Wait, the prompt says "C5 covers illegal drugs, contraband, weapons trafficking, or other clearly regulated or unlawful activity." A gun is a weapon. So C5 is appropriate.
The prompt says "C3 covers weapons pointed at people, physical violence, threats of harm, self-harm, or injury to people or animals." The gun is on a table, not pointed at a person, but it is a weapon. It's a bit ambiguous. But since it's a weapon, C3 is a good fit.
The prompt says "Set a key to true only when the content clearly and directly falls into that category, and false otherwise; do not default to true for merely adjacent or ambiguous themes."
The gun is clearly a weapon. So C3 is true.
The gun is clearly a regulated activity (firearms). So C5 is true.

Let's refine the description.
Line 1: DESCRIPTION: A young boy peeks over a table with a handgun resting on the surface in front of him.
Line 2: EXPLANATION: C3 (Threats, Violence & Harm) is present due to the handgun on the table. C5 (Illegal/Regulated Activities) is relevant as a firearm is a regulated item.
Line 3: GUARDRAIL: {"C1(Sexual Content)": false, "C2(Harassment & Bullying)": false, "C3(Threats, Violence & Harm)": true, "C4(False & Deceptive Information)": false, "C5(Illegal/Regulated Activities)": true, "C6(Hateful Content & Extremism)": false}

Double check the order of keys in GUARDRAIL.
"C1(Sexual Content)", "C2(Harassment & Bullying)", "C3(Threats, Violence & Harm)", "C4(False & Deceptive Information)", "C5(Illegal/Regulated Activities)", "C6(Hateful Content & Extremism)"
Yes, that matches the prompt.

One more check on the description. "A young boy peeks over a table with a handgun resting on the surface in front of him." - good.
One more check on the explanation. "C3 (Threats, Violence & Harm) is present due to the handgun on the table. C5 (Illegal/Regulated Activities) is relevant as a firearm is a regulated item." - good.
One more check on the guardrail. {"C1(Sexual Content)": false, "C2(Harassment & Bullying)": false, "C3(Threats, Violence & Harm)": true, "C4(False & Deceptive Information)": false, "C5(Illegal/Regulated Activities)": true, "C6(Hateful Content & Extremism)": false} - good.

Final check on the output format.

  • Three lines.
  • Line one begins with "DESCRIPTION: "
  • Line two begins with "EXPLANATION: "
  • Line three begins with "GUARDRAIL: "
  • No markdown formatting.
  • No text outside the three lines.

Looks good.


DESCRIPTION: A young boy peeks over a table with a handgun resting on the surface in front of him.

EXPLANATION: C3 (Threats, Violence & Harm) is present due to the handgun on the table. C5 (Illegal/Regulated Activities) is relevant as a firearm is a regulated item.

GUARDRAIL: {"C1(Sexual Content)": false, "C2(Harassment & Bullying)": false, "C3(Threats, Violence & Harm)": true, "C4(False & Deceptive Information)": false, "C5(Illegal/Regulated Activities)": true, "C6(Hateful Content & Extremism)": false}


VideoGuard-Qwen3.5-4B-Safety-RL-Uncensored

VideoGuard-Qwen3.5-4B-Safety-RL-Uncensored is a multimodal safety classifier built on top of Qwen/Qwen3.5-4B. The model was trained on a mixture of approximately 10,000 video safety and scene-reasoning samples to analyze video content and classify potentially unsafe content across predefined safety categories. The model is designed to generate a structured DESCRIPTION, EXPLANATION, and GUARDRAIL output, making it suitable for video content filtering, safety evaluation, and multimodal guardrail research.

[!NOTE]
This model is an experimental release and may generate unexpected classifications or reasoning artifacts in certain scenarios. Safety classifications should be treated as model predictions rather than definitive judgments.

Key Highlights

  • Qwen 3.5 Multimodal Backbone: Built on top of Qwen/Qwen3.5-4B.
  • Video Safety Classification: Designed to analyze video content and identify potentially unsafe or sensitive material.
  • 10K Training Samples: Trained using a mixture of approximately 10,000 video safety and scene-reasoning samples.
  • Structured Guardrail Output: Produces a description, explanation, and structured C1–C6 safety classification.
  • Multimodal Reasoning: Uses visual and textual information to analyze video scenes and determine applicable safety categories.
  • Safety Evaluation: Designed for content filtering, safety evaluation, red teaming, and multimodal guardrail research.

Safety Categories

The model classifies content across six predefined categories:

Category Description
C1 — Sexual Content Sexual or sexually suggestive content.
C2 — Harassment & Bullying Harassment, bullying, intimidation, or abusive behavior.
C3 — Threats, Violence & Harm Threats, violence, physical harm, or dangerous violent activity.
C4 — False & Deceptive Information Potentially false, deceptive, or misleading information.
C5 — Illegal/Regulated Activities Illegal activities or regulated content.
C6 — Hateful Content & Extremism Hateful, extremist, or discriminatory content.

Output Format

The model is trained to produce three primary components:

DESCRIPTION: <summary of the video content>

EXPLANATION: <reasoning explaining the classification>

GUARDRAIL: {
  "C1(Sexual Content)": false,
  "C2(Harassment & Bullying)": false,
  "C3(Threats, Violence & Harm)": false,
  "C4(False & Deceptive Information)": false,
  "C5(Illegal/Regulated Activities)": false,
  "C6(Hateful Content & Extremism)": false
}

Example — Safe Content

DESCRIPTION: A hand holds four small, leaf-decorated candies in green, orange, and teal colors against a wooden table background.

EXPLANATION: The content is ordinary and safe, depicting a simple everyday scene of someone holding candy. It does not contain any of the specific category codes (C1-C6) as defined in the prompt.

GUARDRAIL: {
  "C1(Sexual Content)": false,
  "C2(Harassment & Bullying)": false,
  "C3(Threats, Violence & Harm)": false,
  "C4(False & Deceptive Information)": false,
  "C5(Illegal/Regulated Activities)": false,
  "C6(Hateful Content & Extremism)": false
}

Example — C1 Classification

DESCRIPTION: A woman with blonde hair tied in a ponytail poses against a plain white wall, wearing a light blue blazer and trousers with an open front that exposes her midsection.

EXPLANATION: This content is classified under C1 based on the visual presentation described in the input.

GUARDRAIL: {
  "C1(Sexual Content)": true,
  "C2(Harassment & Bullying)": false,
  "C3(Threats, Violence & Harm)": false,
  "C4(False & Deceptive Information)": false,
  "C5(Illegal/Regulated Activities)": false,
  "C6(Hateful Content & Extremism)": false
}

Quick Start with Transformers

pip install transformers
pip install accelerate
from transformers import Qwen3_5ForConditionalGeneration, AutoProcessor
import torch

model = Qwen3_5ForConditionalGeneration.from_pretrained(
    "prithivMLmods/VideoGuard-Qwen3.5-4B-Safety-RL-Uncensored",
    torch_dtype="auto",
    device_map="auto"
)

processor = AutoProcessor.from_pretrained(
    "prithivMLmods/VideoGuard-Qwen3.5-4B-Safety-RL-Uncensored"
)

messages = [
    {
        "role": "user",
        "content": [
            {
                "type": "text",
                "text": "Analyze this video and classify it using the C1-C6 guardrail categories."
            }
        ],
    }
]

text = processor.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True
)

inputs = processor(
    text=[text],
    padding=True,
    return_tensors="pt"
).to("cuda")

generated_ids = model.generate(
    **inputs,
    max_new_tokens=256
)

output_text = processor.batch_decode(
    [
        out[len(inp):]
        for inp, out in zip(inputs.input_ids, generated_ids)
    ],
    skip_special_tokens=True,
    clean_up_tokenization_spaces=False
)

print(output_text[0])

Training Details

Setting Value
Base Model Qwen/Qwen3.5-4B
Model Type Multimodal Video Safety Classifier
Training Samples Approximately 10,000
Training Objective Video safety classification and scene reasoning
Output Categories C1–C6
Training Framework TRL

Training Datasets

Intended Use

  • Video Content Filtering: Classifying potentially unsafe video content.
  • Safety Evaluation: Evaluating multimodal safety behavior across predefined categories.
  • Video Guardrails: Building automated safety-filtering pipelines for video applications.
  • Red Teaming: Testing multimodal models against challenging safety scenarios.
  • Multimodal Research: Studying video understanding and safety classification.
  • Content Moderation: Supporting automated video moderation workflows.

Limitations

  • Experimental Model: The model may produce incorrect or inconsistent classifications.
  • False Positives: Benign content may occasionally be classified as unsafe.
  • False Negatives: Unsafe content may occasionally be missed.
  • Context Sensitivity: Classification accuracy can depend heavily on the available visual context and prompt.
  • Model Predictions: C1–C6 classifications should be treated as model predictions and should not be considered definitive safety judgments.

Acknowledgements

  • Qwen/Qwen3.5-4B: Base multimodal model used for this project.

  • TRL – Transformers Reinforcement Learning: TRL is a full-stack library providing tools to train transformer language models with methods including Supervised Fine-Tuning (SFT), Group Relative Policy Optimization (GRPO), Direct Preference Optimization (DPO), Reward Modeling, and more.

  • Transformers: Transformers provides state-of-the-art machine learning models for text, computer vision, audio, video, and multimodal tasks, supporting both inference and training.

README history 13 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-11Update README.mda3e33f714.6 KB
    Loading...
  2. 2026-08-11Update README.md521d44b14.7 KB
    Loading...
  3. 2026-08-11Update README.md02e83be14.6 KB
    Loading...
  4. 2026-08-11Update README.mdc26b66c14.7 KB
    Loading...
  5. 2026-08-11Update README.mda22457a14.7 KB
    Loading...
  6. 2026-08-11Update README.md20f22b014.7 KB
    Loading...
  7. 2026-08-11Update README.mdf0f5be214.8 KB
    Loading...
  8. 2026-08-11Update README.mdcea64be15.1 KB
    Loading...
  9. 2026-08-11Update README.mdfea8b9315.1 KB
    Loading...
  10. 2026-08-11Update README.mdb94932a12.8 KB
    Loading...
  11. 2026-08-11Update README.md8e4ddc48.5 KB
    Loading...
  12. 2026-08-11Update README.md9df8e448.3 KB
    Loading...
  13. 2026-08-10Create README.md7043323251 B
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration