← back to catalog · registered 2026-08-22 13:56

prithivMLmods/Qwen3-VL-8B-Abliterated-Caption-it

prithivMLmods Qwen 8.8B multimodal
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/prithivMLmods%2FQwen3-VL-8B-Abliterated-Caption-it"
Response includes
  • classification m1
  • files 16
  • benchmarks 11 entries
  • hub_downloads_all_time 8,120
  • author_summary 98 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
8K
190 last 30d - cooling
Likes
37
Descendants
4
in 4 direct forks
Model age
11mo ago
created 2025-10-26
Downloads over time
Now8.1K→from76↑10,612%
03K6K8.9K76 on Oct 29, 20258.1K on Oct 11Oct '25Dec '25FebAprJunAugOct
Oct 29, 2025 → Oct 11 · 89 snapshots · spans 347 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 1.2 UGI
Hazardous 4.1 UGI
Natural Intelligence 17.46 UGI
Political lean -18.5% UGI
Sensitive-Info 19.17 UGI
SocPol 1.1 UGI
UGI 32.78 UGI
Willingness (10) 6 UGI
W10-Adherence 6 UGI
W10-Direct 6 UGI
Writing 30.31 UGI

Genealogy 4 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 3 formats · 6K downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
apache-2.0
Languages
en zh th
Tags
transformers safetensors qwen3_vl image-text-to-text trl text-generation-inference uncensored image-captioning vlm visual-understanding caption image-to-text

Related

Total size
16.3 GB
Files
16
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-06-01 09:01

Files by quantization

Auxiliary files 16 files 16.3 GB
model-00001-of-00002.safetensors 9.23 GB ******** download
model-00002-of-00002.safetensors 7.10 GB ******** download
tokenizer.json 10.9 MB ******** download
vocab.json 2.65 MB 4783fe10 download
merges.txt 1.59 MB 31349551 download
model.safetensors.index.json 66.2 KB 3ce63d41 download
tokenizer_config.json 5.50 KB 3899c353 download
chat_template.jinja 5.17 KB 12438680 download
README.md 4.23 KB 653005d7 download
config.json 1.54 KB 14c8a406 download
.gitattributes 1.53 KB 52373fe2 download
video_preprocessor_config.json 817 B e32b1d90 download
preprocessor_config.json 782 B 2fa65535 download
added_tokens.json 707 B b54f9135 download
special_tokens_map.json 613 B ac23c0aa download
generation_config.json 147 B 7532cbda download

README current version from Hugging Face


license: apache-2.0
language:

  • en
  • zh
  • th
    base_model:
  • Qwen/Qwen3-VL-8B-Instruct
    pipeline_tag: image-text-to-text
    library_name: transformers
    tags:
  • trl
  • text-generation-inference
  • uncensored
  • image-captioning
  • vlm
  • visual-understanding
  • caption
  • image-to-text

1

Qwen3-VL-8B-Abliterated-Caption-it

The Qwen3-VL-8B-Abliterated-Caption-it model is a fine-tuned version of Qwen3-VL-8B-Instruct, tailored for Abliterated Captioning / Uncensored Image Captioning. This variant is designed to generate highly detailed and descriptive captions across a broad range of visual categories, including images with complex, sensitive, or nuanced content—across varying aspect ratios and resolutions.

Key Highlights

  • Abliterated / Uncensored Captioning: Fine-tuned to bypass common content filters while preserving factual and descriptive richness across diverse visual categories.
  • High-Fidelity Descriptions: Generates comprehensive captions for general, artistic, technical, abstract, and low-context images.
  • Robust Across Aspect Ratios: Capable of accurately captioning images with wide, tall, square, and irregular dimensions.
  • Variational Detail Control: Produces outputs with both high-level summaries and fine-grained descriptions as needed.
  • Foundation on Qwen3-VL Architecture: Leverages the strengths of the Qwen3-VL-8B multimodal model for visual reasoning, comprehension, and instruction-following.
  • Multilingual Output Capability: Supports multilingual descriptions (English as default), adaptable via prompt engineering.

Base Model Signatures:

This model has been re-sharded and optimized for the latest Transformers version from the base model: https://huggingface.co/huihui-ai/Huihui-Qwen3-VL-8B-Instruct-abliterated.


Quick Start with Transformers

[!note]
Instruction Query: Provide a detailed caption for the image

from transformers import Qwen3VLForConditionalGeneration, AutoProcessor
from qwen_vl_utils import process_vision_info

model = Qwen3VLForConditionalGeneration.from_pretrained(
    "prithivMLmods/Qwen3-VL-8B-Abliterated-Caption-it", torch_dtype="auto", device_map="auto"
)

processor = AutoProcessor.from_pretrained("prithivMLmods/Qwen3-VL-8B-Abliterated-Caption-it")

messages = [
    {
        "role": "user",
        "content": [
            {
                "type": "image",
                "image": "https://qianwen-res.oss-cn-beijing.aliyuncs.com/Qwen-VL/assets/demo.jpeg",
            },
            {"type": "text", "text": "Describe this image in detail."},
        ],
    }
]

text = processor.apply_chat_template(
    messages, tokenize=False, add_generation_prompt=True
)
image_inputs, video_inputs = process_vision_info(messages)
inputs = processor(
    text=[text],
    images=image_inputs,
    videos=video_inputs,
    padding=True,
    return_tensors="pt",
)
inputs = inputs.to("cuda")

generated_ids = model.generate(**inputs, max_new_tokens=128)
generated_ids_trimmed = [
    out_ids[len(in_ids):] for in_ids, out_ids in zip(inputs.input_ids, generated_ids)
]
output_text = processor.batch_decode(
    generated_ids_trimmed, skip_special_tokens=True, clean_up_tokenization_spaces=False
)
print(output_text)

Intended Use

This model is suited for:

  • Generating detailed and unfiltered image captions for general-purpose or artistic datasets.
  • Content moderation research, red-teaming, and generative safety evaluations.
  • Enabling descriptive captioning for visual datasets typically excluded from mainstream models.
  • Creative applications (e.g., storytelling, art generation) that benefit from rich descriptive captions.
  • Captioning for non-standard aspect ratios and stylized visual content.

Limitations

  • May produce explicit, sensitive, or offensive descriptions depending on image content and prompts.
  • Not suitable for deployment in production systems requiring content filtering or moderation.
  • Can exhibit variability in caption tone or style depending on input prompt phrasing.
  • Accuracy for unfamiliar or synthetic visual styles may vary.
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration