← back to catalog · registered 2026-08-22 13:56

matatonic/QVQ-72B-Preview-abliterated-6.5bpw-h8-exl2

matatonic Qwen 72B multimodal
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/matatonic%2FQVQ-72B-Preview-abliterated-6.5bpw-h8-exl2"
Response includes
  • classification m1
  • files 21
  • hub_downloads_all_time 187
  • author_summary 16 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
187
18 last 30d - cooling
Likes
1
Model age
21mo ago
created 2025-01-06
Downloads over time
Now196→from1↑19,500%
0721442161 on Jan 1, 2025196 on Oct 11196 on Oct 10Jan '25Apr '25Jul '25Oct '25JanAprJulOct
Jan 1, 2025 → Oct 11 · 132 snapshots · spans 648 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
other
Languages
en
Tags
transformers safetensors qwen2_vl image-text-to-text abliterated uncensored - chat conversational en base_model:Qwen/QVQ-72B-Preview base_model:quantized:Qwen/QVQ-72B-Preview license:other text-generation-inference

Related

Total size
57.8 GB
Files
21
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2025-01-06 22:54

Files by quantization

Auxiliary files 21 files 57.8 GB
output-00001-of-00006.safetensors 10.00 GB 204b8184 download
output-00002-of-00006.safetensors 9.99 GB bd296e64 download
output-00003-of-00006.safetensors 9.98 GB c78cd3c5 download
output-00005-of-00006.safetensors 9.94 GB c847d27a download
output-00004-of-00006.safetensors 9.87 GB 85aaac59 download
output-00006-of-00006.safetensors 8.05 GB 8f855fdd download
tokenizer.json 10.9 MB 9c5ae00e download
measurement.json 4.26 MB 6a808f4c download
vocab.json 2.65 MB 4783fe10 download
merges.txt 1.59 MB 20024bfe download
model.safetensors.index.json 107 KB 730c42a7 download
LICENSE 6.80 KB b547357c download
tokenizer_config.json 5.92 KB 1882ca76 download
README.md 3.17 KB 1ed48369 download
config.json 1.56 KB 22b4f09a download
.gitattributes 1.53 KB 52373fe2 download
chat_template.json 1.10 KB 36286ad1 download
special_tokens_map.json 644 B 3a784031 download
added_tokens.json 629 B 06135f3c download
preprocessor_config.json 348 B 5c84044b download
generation_config.json 245 B 8a71055a download

README current version from Hugging Face


license: other
license_name: qwen
license_link: https://huggingface.co/huihui-ai/QVQ-72B-Preview-abliterated/blob/main/LICENSE
language:

  • en
    pipeline_tag: image-text-to-text
    base_model: Qwen/QVQ-72B-Preview
    tags:
  • abliterated
  • uncensored
    • chat
      library_name: transformers

huihui-ai/QVQ-72B-Preview-abliterated

This is an uncensored version of Qwen/QVQ-72B-Preview created with abliteration (see remove-refusals-with-transformers to know more about it).

This is a crude, proof-of-concept implementation to remove refusals from an LLM model without using TransformerLens.

It was only the text part that was processed, not the image part.

Usage

We offer a toolkit to help you handle various types of visual input more conveniently. This includes base64, URLs, and interleaved images and videos. You can install it using the following command:

pip install qwen-vl-utils

Here we show a code snippet to show you how to use the chat model with transformers and qwen_vl_utils:

from transformers import Qwen2VLForConditionalGeneration, AutoTokenizer, AutoProcessor
from qwen_vl_utils import process_vision_info

# default: Load the model on the available device(s)
model = Qwen2VLForConditionalGeneration.from_pretrained(
    "huihui-ai/QVQ-72B-Preview-abliterated", torch_dtype="auto", device_map="auto"
)

# default processer
processor = AutoProcessor.from_pretrained("huihui-ai/QVQ-72B-Preview-abliterated")

# The default range for the number of visual tokens per image in the model is 4-16384. You can set min_pixels and max_pixels according to your needs, such as a token count range of 256-1280, to balance speed and memory usage.
# min_pixels = 256*28*28
# max_pixels = 1280*28*28
# processor = AutoProcessor.from_pretrained("huihui-ai/QVQ-72B-Preview-abliterated", min_pixels=min_pixels, max_pixels=max_pixels)

messages = [
    {
        "role": "system",
        "content": [
            {"type": "text", "text": "You are a helpful and harmless assistant. You are Qwen developed by Alibaba. You should think step-by-step."}
        ],
    },
    {
        "role": "user",
        "content": [
            {
                "type": "image",
                "image": "https://qianwen-res.oss-cn-beijing.aliyuncs.com/QVQ/demo.png",
            },
            {"type": "text", "text": "What value should be filled in the blank space?"},
        ],
    }
]

# Preparation for inference
text = processor.apply_chat_template(
    messages, tokenize=False, add_generation_prompt=True
)
image_inputs, video_inputs = process_vision_info(messages)
inputs = processor(
    text=[text],
    images=image_inputs,
    videos=video_inputs,
    padding=True,
    return_tensors="pt",
)
inputs = inputs.to("cuda")

# Inference: Generation of the output
generated_ids = model.generate(**inputs, max_new_tokens=8192)
generated_ids_trimmed = [
    out_ids[len(in_ids) :] for in_ids, out_ids in zip(inputs.input_ids, generated_ids)
]
output_text = processor.batch_decode(
    generated_ids_trimmed, skip_special_tokens=True, clean_up_tokenization_spaces=False
)
print(output_text)

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2025-01-06initial5cb1d2a3.2 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration