← back to catalog · registered 2026-08-22 13:56

huihui-ai/Qwen2.5-VL-3B-Instruct-abliterated

huihui-ai Qwen 3.8B GGUF multimodal 128K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/huihui-ai%2FQwen2.5-VL-3B-Instruct-abliterated"
Response includes
  • classification m8
  • files 16
  • benchmarks 11 entries
  • hub_downloads_all_time 153,462
  • author_summary 183 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of layer-wise ablation inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • author=huihui-ai + is_gguf=1
  • M3 (huihui abliteration) wrapped in M8 (GGUF quantization)
Refusal direction extracted via
Extraction technique

huihui-ai layer-band extraction

Confidence
HIGH
Why we say so
producer=huihui-ai (documented layer-band methodology in model cards)
Downloads · lifetime
153K
22K last 30d - stable
Likes
35
Descendants
3
in 3 direct forks
Model age
20mo ago
created 2025-02-16
Downloads over time
Now162.4K→from18↑901,906%
059.5K119.1K178.6K18 on Feb 12, 2025162.4K on Oct 11Feb '25May '25Aug '25Nov '25FebMayAug
Feb 12, 2025 → Oct 11 · 133 snapshots · spans 606 days

Benchmarks

Benchmark Score Source
Entertainment 0.8 UGI
Hazardous 1.8 UGI
Natural Intelligence 9.42 UGI
Political lean -41.7% UGI
Sensitive-Info 9.4 UGI
SocPol 0.5 UGI
UGI 27.1 UGI
Willingness (10) 6.2 UGI
W10-Adherence 4.5 UGI
W10-Direct 8 UGI
Writing 17.47 UGI

Genealogy 3 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

Languages
en
Tags
transformers safetensors gguf qwen2_5_vl image-text-to-text multimodal abliterated uncensored conversational en base_model:Qwen/Qwen2.5-VL-3B-Instruct base_model:quantized:Qwen/Qwen2.5-VL-3B-Instruct

Related

Total size
6.99 GB
Files
16
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2025-11-07 23:35

Files by quantization

Auxiliary files 16 files 7.00 GB
model-00001-of-00002.safetensors 4.65 GB 3353a229 download
model-00002-of-00002.safetensors 2.34 GB 52369eee download
tokenizer.json 6.71 MB c0382117 download
vocab.json 2.65 MB 4783fe10 download
merges.txt 1.59 MB 31349551 download
model.safetensors.index.json 64.7 KB d9c73564 download
LICENSE 7.21 KB 87f48cf8 download
tokenizer_config.json 7.06 KB 5048e7ae download
README.md 3.15 KB 015b88de download
.gitattributes 1.78 KB 7ba05f70 download
config.json 1.40 KB cfb2d061 download
chat_template.json 1.03 KB 732bd68b download
special_tokens_map.json 644 B 3a784031 download
added_tokens.json 629 B 06135f3c download
preprocessor_config.json 351 B 8273ab1b download
generation_config.json 229 B 9b492590 download

README current version from Hugging Face


license_name: qwen-research
license_link: https://huggingface.co/Qwen/Qwen2.5-VL-3B-Instruct/blob/main/LICENSE
language:

  • en
    pipeline_tag: image-text-to-text
    tags:
  • multimodal
  • abliterated
  • uncensored
    library_name: transformers
    base_model:
  • Qwen/Qwen2.5-VL-3B-Instruct

huihui-ai/Qwen2.5-VL-3B-Instruct-abliterated

This is an uncensored version of Qwen/Qwen2.5-VL-3B-Instruct created with abliteration (see remove-refusals-with-transformers to know more about it).

It was only the text part that was processed, not the image part.

ollama

You can use huihui_ai/qwen2.5-vl-abliterated:3b directly,

ollama run huihui_ai/qwen2.5-vl-abliterated:3b

GGUF

The official llama.cpp-b6907 has now been updated to support Qwen2.5-VL conversion to GGUF format and can be tested using llama-mtmd-cli.

The GGUF file has been uploaded.

llama-mtmd-cli -m huihui-ai/Qwen2.5-VL-3B-Instruct-abliterated/GGUF/ggml-model-Q4_K_M.gguf --mmproj huihui-ai/Qwen2.5-VL-3B-Instruct-abliterated/GGUF/mmproj-ggml-model-f16.gguf -c 4096 --image png/cc.jpg -p "Describe this image." 

If it's just for chatting, you can use llama-cli.

llama-cli -m huihui-ai/Qwen2.5-VL-3B-Instruct-abliterated/GGUF/ggml-model-Q4_K_M.gguf -c 4096

Usage

You can use this model in your applications by loading it with Hugging Face's transformers library:

from transformers import Qwen2_5_VLForConditionalGeneration, AutoTokenizer, AutoProcessor
from qwen_vl_utils import process_vision_info

model = Qwen2_5_VLForConditionalGeneration.from_pretrained(
    "huihui-ai/Qwen2.5-VL-3B-Instruct-abliterated", torch_dtype="auto", device_map="auto"
)
processor = AutoProcessor.from_pretrained("huihui-ai/Qwen2.5-VL-3B-Instruct-abliterated")

image_path = "/tmp/test.png"

messages = [
    {
        "role": "user",
        "content": [
            {
                "type": "image",
                "image": f"file://{image_path}",
            },
            {"type": "text", "text": "Describe this image."},
        ],
    }
]

text = processor.apply_chat_template(
    messages, tokenize=False, add_generation_prompt=True
)
image_inputs, video_inputs = process_vision_info(messages)
inputs = processor(
    text=[text],
    images=image_inputs,
    videos=video_inputs,
    padding=True,
    return_tensors="pt",
)
inputs = inputs.to("cuda")

generated_ids = model.generate(**inputs, max_new_tokens=256)
generated_ids_trimmed = [
    out_ids[len(in_ids) :] for in_ids, out_ids in zip(inputs.input_ids, generated_ids)
]
output_text = processor.batch_decode(
    generated_ids_trimmed, skip_special_tokens=True, clean_up_tokenization_spaces=False
)
output_text = output_text[0]

print(output_text)

Donation

Your donation helps us continue our further development and improvement, a cup of coffee can do it.
  • bitcoin:
  bc1qqnkhuchxw0zqjh2ku3lu4hq45hc6gy84uk70ge

README history 8 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2025-11-07Update README.mda0d13233.2 KB
    Loading...
  2. 2025-11-07Update README.mdcaeda053 KB
    Loading...
  3. 2025-11-07Update README.mdcf63a793 KB
    Loading...
  4. 2025-02-17Update README.mda64b76c2.3 KB
    Loading...
  5. 2025-02-16Update README.md20f4a0d2.3 KB
    Loading...
  6. 2025-02-16Update README.md4d209482.3 KB
    Loading...
  7. 2025-02-16Upload 15 files766206e2.1 KB
    Loading...
  8. 2025-02-16initial commit7dcbbe628 B
    Loading...

Discussions 4 threads

  1. 2025-10-27Projector file where?open6 💬#4
    Loading...
  2. 2025-07-14PRFixed the name of the image processormerged1 💬#3
    Loading...
  3. 2025-04-06vllm erroropen5 💬#2
    Loading...
  4. 2025-02-16When is the 7B Version comingclosed2 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration