← back to catalog · registered 2026-08-22 13:56

prithivMLmods/VideoGuard-Qwen3.5-9B-Safety-RL-Uncensored

prithivMLmods Qwen 9.4B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/prithivMLmods%2FVideoGuard-Qwen3.5-9B-Safety-RL-Uncensored"
Response includes
  • classification m-uncensored
  • files 9
  • benchmarks 11 entries
  • hub_downloads_all_time 70
  • author_summary 98 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
70
52 last 30d - active
Likes
2
Descendants
1
in 1 direct fork
Model age
2mo ago
created 2026-08-10

Training datasets

Corpora the author lists in the model card. Datasets tracked in our /datasets catalog carry a category badge linking to the workflow stage. Others open on Hugging Face.

Downloads over time
Now94→from0↑0%
034691030 on Aug 594 on Oct 11AugSepOct
Aug 5 → Oct 11 · 50 snapshots · spans 67 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 1.4 UGI
Hazardous 2.4 UGI
Natural Intelligence 17.62 UGI
Political lean -12.2% UGI
Sensitive-Info 14.65 UGI
SocPol 0.9 UGI
UGI 17.27 UGI
Willingness (10) 2.2 UGI
W10-Adherence 1.5 UGI
W10-Direct 3 UGI
Writing 33.52 UGI

Genealogy 1 direct fork

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 2 formats · 1K downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
apache-2.0
Languages
en
Tags
transformers safetensors qwen3_5 image-text-to-text text-generation-inference reinforcement-learning video-text-to-text alignment-training RLHF RFT video-understanding video-classification

Related

Total size
17.5 GB
Files
9
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-11 17:28

Files by quantization

Auxiliary files 9 files 17.5 GB
model.safetensors 17.5 GB a91e0b53 download
tokenizer.json 19.1 MB 87a7830d download
README.md 12.4 KB 02e40b8b download
chat_template.jinja 9.62 KB d45b0278 download
config.json 2.77 KB 734ecb48 download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.27 KB 8af8110f download
tokenizer_config.json 1.14 KB 43f5bc3a download
generation_config.json 116 B 48697309 download

README current version from Hugging Face


license: apache-2.0
base_model:

  • Qwen/Qwen3.5-9B
    language:
  • en
    pipeline_tag: video-text-to-text
    library_name: transformers
    tags:
  • text-generation-inference
  • reinforcement-learning
  • video-text-to-text
  • alignment-training
  • RLHF
  • RFT
  • video-understanding
  • video-classification
  • video-safety
  • content-safety
  • content-moderation
  • safety-classifier
  • guardrail
    datasets:
  • prithivMLmods/OpenVideo-Scene-Reasoning
  • PALM-Lab/vid-guard-rlhf-unsafe

VideoGuard-9B

demo
think
The user wants me to classify an image and provide a specific output format. The image shows a person on a high-rise structure, likely a crane or a tower under construction, with a cityscape below. I need to output three lines: 1. **DESCRIPTION:** A single sentence describing the image. 2. **EXPLANATION:** One or two sentences explaining the category codes (C1-C6) and why. 3. **GUARDRAIL:** A single-line JSON object with specific boolean keys.

Step 1: Description
The image shows a person in a black hoodie and camouflage pants standing on a red steel structure, likely a crane or a high-rise construction site, with a cityscape visible in the foggy background.

Step 2: Explanation
The image depicts a worker on a high structure, which falls under "C5 (Illegal/Regulated Activities)" as it shows construction or maintenance work. It does not show sexual content, harassment, threats, misinformation, or hate symbols.

Step 3: Guardrail
I need to set the boolean keys to true or false.

  • C1 (Sexual Content): False.
  • C2 (Harassment & Bullying): False.
  • C3 (Threats, Violence & Harm): False.
  • C4 (False & Deceptive Information): False.
  • C5 (Illegal/Regulated Activities): True (Construction work).
  • C6 (Hateful Content & Extremism): False.

Let's format the output.
Line 1: DESCRIPTION: A person in a black hoodie and camouflage pants stands on a red steel structure, likely a crane or high-rise construction site, with a cityscape visible in the foggy background.
Line 2: EXPLANATION: The image depicts a worker on a high structure, which falls under "C5 (Illegal/Regulated Activities)" as it shows construction or maintenance work.
Line 3: GUARDRAIL: {"C1(Sexual Content)": false, "C2(Harassment & Bullying)": false, "C3(Threats, Violence & Harm)": false, "C4(False & Deceptive Information)": false, "C5(Illegal/Regulated Activities)": true, "C6(Hateful Content & Extremism)": false}

Double check constraints:

  • Exactly three lines.
  • No markdown formatting outside the three lines.
  • JSON keys in exact order.
  • Boolean values are strictly true/false.

Looks good.


DESCRIPTION: A person in a black hoodie and camouflage pants stands on a red steel structure, likely a crane or high-rise construction site, with a cityscape visible in the foggy background.

EXPLANATION: The image depicts a worker on a high structure, which falls under "C5 (Illegal/Regulated Activities)" as it shows construction or maintenance work.

GUARDRAIL: {"C1(Sexual Content)": false, "C2(Harassment & Bullying)": false, "C3(Threats, Violence & Harm)": false, "C4(False & Deceptive Information)": false, "C5(Illegal/Regulated Activities)": true, "C6(Hateful Content & Extremism)": false}


VideoGuard-Qwen3.5-9B-Safety-RL-Uncensored

VideoGuard-Qwen3.5-9B-Safety-RL-Uncensored is a multimodal safety classifier built on top of Qwen/Qwen3.5-9B. The model was trained on a mixture of approximately 10,000 video safety and scene-reasoning samples to analyze video content and classify potentially unsafe content across predefined safety categories. The model is designed to generate a structured DESCRIPTION, EXPLANATION, and GUARDRAIL output, making it suitable for video content filtering, safety evaluation, and multimodal guardrail research.

[!NOTE]
This model is an experimental release and may generate unexpected classifications or reasoning artifacts in certain scenarios. Safety classifications should be treated as model predictions rather than definitive judgments.

Key Highlights

  • Qwen 3.5 Multimodal Backbone: Built on top of Qwen/Qwen3.5-9B.
  • Video Safety Classification: Designed to analyze video content and identify potentially unsafe or sensitive material.
  • 10K Training Samples: Trained using a mixture of approximately 10,000 video safety and scene-reasoning samples.
  • Structured Guardrail Output: Produces a description, explanation, and structured C1–C6 safety classification.
  • Multimodal Reasoning: Uses visual and textual information to analyze video scenes and determine applicable safety categories.
  • Safety Evaluation: Designed for content filtering, safety evaluation, red teaming, and multimodal guardrail research.

Safety Categories

The model classifies content across six predefined categories:

Category Description
C1 — Sexual Content Sexual or sexually suggestive content.
C2 — Harassment & Bullying Harassment, bullying, intimidation, or abusive behavior.
C3 — Threats, Violence & Harm Threats, violence, physical harm, or dangerous violent activity.
C4 — False & Deceptive Information Potentially false, deceptive, or misleading information.
C5 — Illegal/Regulated Activities Illegal activities or regulated content.
C6 — Hateful Content & Extremism Hateful, extremist, or discriminatory content.

Output Format

The model is trained to produce three primary components:

DESCRIPTION: <summary of the video content>

EXPLANATION: <reasoning explaining the classification>

GUARDRAIL: {
  "C1(Sexual Content)": false,
  "C2(Harassment & Bullying)": false,
  "C3(Threats, Violence & Harm)": false,
  "C4(False & Deceptive Information)": false,
  "C5(Illegal/Regulated Activities)": false,
  "C6(Hateful Content & Extremism)": false
}

Example — Safe Content

DESCRIPTION: A hand holds four small, leaf-decorated candies in green, orange, and teal colors against a wooden table background.

EXPLANATION: The content is ordinary and safe, depicting a simple everyday scene of someone holding candy. It does not contain any of the specific category codes (C1-C6) as defined in the prompt.

GUARDRAIL: {
  "C1(Sexual Content)": false,
  "C2(Harassment & Bullying)": false,
  "C3(Threats, Violence & Harm)": false,
  "C4(False & Deceptive Information)": false,
  "C5(Illegal/Regulated Activities)": false,
  "C6(Hateful Content & Extremism)": false
}

Example — Sexual Content Classification

DESCRIPTION: A woman with blonde hair tied in a ponytail poses against a plain white wall, wearing a light blue blazer and trousers with an open front that exposes her midsection.

EXPLANATION: This content is classified under C1 based on the visual presentation described in the input.

GUARDRAIL: {
  "C1(Sexual Content)": true,
  "C2(Harassment & Bullying)": false,
  "C3(Threats, Violence & Harm)": false,
  "C4(False & Deceptive Information)": false,
  "C5(Illegal/Regulated Activities)": false,
  "C6(Hateful Content & Extremism)": false
}

Quick Start with Transformers

pip install transformers
pip install accelerate
from transformers import Qwen3_5ForConditionalGeneration, AutoProcessor
import torch

model = Qwen3_5ForConditionalGeneration.from_pretrained(
    "prithivMLmods/VideoGuard-Qwen3.5-9B-Safety-RL-Uncensored",
    torch_dtype="auto",
    device_map="auto"
)

processor = AutoProcessor.from_pretrained(
    "prithivMLmods/VideoGuard-Qwen3.5-9B-Safety-RL-Uncensored"
)

messages = [
    {
        "role": "user",
        "content": [
            {
                "type": "text",
                "text": "Analyze this video and classify it using the C1-C6 guardrail categories."
            }
        ],
    }
]

text = processor.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True
)

inputs = processor(
    text=[text],
    padding=True,
    return_tensors="pt"
).to("cuda")

generated_ids = model.generate(
    **inputs,
    max_new_tokens=256
)

output_text = processor.batch_decode(
    [
        out[len(inp):]
        for inp, out in zip(inputs.input_ids, generated_ids)
    ],
    skip_special_tokens=True,
    clean_up_tokenization_spaces=False
)

print(output_text[0])

Training Details

Setting Value
Base Model Qwen/Qwen3.5-9B
Model Type Multimodal Video Safety Classifier
Training Samples Approximately 10,000
Training Objective Video safety classification and scene reasoning
Output Categories C1–C6
Training Framework TRL

Training Datasets

Intended Use

  • Video Content Filtering: Classifying potentially unsafe video content.
  • Safety Evaluation: Evaluating multimodal safety behavior across predefined categories.
  • Video Guardrails: Building automated safety-filtering pipelines for video applications.
  • Red Teaming: Testing multimodal models against challenging safety scenarios.
  • Multimodal Research: Studying video understanding and safety classification.
  • Content Moderation: Supporting automated video moderation workflows.

Limitations

  • Experimental Model: The model may produce incorrect or inconsistent classifications.
  • False Positives: Benign content may occasionally be classified as unsafe.
  • False Negatives: Unsafe content may occasionally be missed.
  • Context Sensitivity: Classification accuracy can depend heavily on the available visual context and prompt.
  • Model Predictions: C1–C6 classifications should be treated as model predictions and should not be considered definitive safety judgments.

Acknowledgements

  • Qwen/Qwen3.5-9B: Base multimodal model used for this project.

  • TRL – Transformers Reinforcement Learning: TRL is a full-stack library providing tools to train transformer language models with methods including Supervised Fine-Tuning (SFT), Group Relative Policy Optimization (GRPO), Direct Preference Optimization (DPO), Reward Modeling, and more.

  • Transformers: Transformers provides state-of-the-art machine learning models for text, computer vision, audio, video, and multimodal tasks, supporting both inference and training.

README history 5 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-11Update README.mdc14e63d12.4 KB
    Loading...
  2. 2026-08-11Update README.mdb9346d08.5 KB
    Loading...
  3. 2026-08-11Update README.md700e75a8.3 KB
    Loading...
  4. 2026-08-10Update README.md1b9c112251 B
    Loading...
  5. 2026-08-10initial commitb468a1328 B
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration