← back to catalog · registered 2026-08-22 13:56

cyberneurova/CyberNeurova-Qwen2.5-VL-3B-Instruct-abliterated

cyberneurova Qwen 3.8B multimodal
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/cyberneurova%2FCyberNeurova-Qwen2.5-VL-3B-Instruct-abliterated"
Response includes
  • classification m1
  • files 9
  • benchmarks 11 entries
  • hub_downloads_all_time 376
  • author_summary 5 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
376
57 last 30d - stable
Likes
5
Descendants
2
in 2 direct forks
Model age
4mo ago
created 2026-05-20
Downloads over time
Now408→from25↑1,532%
615329944625 on May 20408 on Oct 11408 on Oct 8MayJunJulAugSepOct
May 20 → Oct 11 · 60 snapshots · spans 144 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 0.6 UGI
Hazardous 1.8 UGI
Natural Intelligence 9.33 UGI
Political lean -15.5% UGI
Sensitive-Info 8.85 UGI
SocPol 0.5 UGI
UGI 20.07 UGI
Willingness (10) 4.2 UGI
W10-Adherence 3.5 UGI
W10-Direct 5 UGI
Writing 20.43 UGI

Genealogy 2 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en
Tags
safetensors qwen2_5_vl abliterated uncensored qwen qwen2.5-vl vision-language multimodal research image-text-to-text conversational en

Related

Total size
6.99 GB
Files
9
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-05-20 10:05

Files by quantization

Auxiliary files 9 files 7.00 GB
model.safetensors 6.99 GB 58905e05 download
tokenizer.json 10.9 MB 3fd16973 download
README.md 8.12 KB 1444ef9c download
config.json 2.51 KB 5dfc9c6c download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.23 KB 7fe4146d download
chat_template.jinja 1017 B 6c226632 download
tokenizer_config.json 737 B dded43f4 download
generation_config.json 213 B 0b8cbe03 download

README current version from Hugging Face


license: apache-2.0
language:

  • en
    base_model: Qwen/Qwen2.5-VL-3B-Instruct
    tags:
  • abliterated
  • uncensored
  • qwen
  • qwen2.5-vl
  • vision-language
  • multimodal
  • research
    pipeline_tag: image-text-to-text

CyberNeurova · Qwen2.5-VL-3B-Instruct · Abliterated

CyberNeurova research — cyberneurova.ai. First VLM release in the CyberNeurova abliteration line.

A permanently-abliterated version of
Qwen/Qwen2.5-VL-3B-Instruct.
Refusal behavior has been removed via residual-stream weight orthogonalization
(Arditi et al. 2024 family of techniques). Inference is unchanged — no runtime
hooks, no slowdown, no custom code. Both text-only and image+text
modes are unlocked from a single capture.

Status: experimental research artifact. This is for defensive security
research, red-team evaluation, and academic study of refusal mechanisms in
multimodal models. Use at your own discretion.


Headline results

Measured on the bf16 hooked model:

Probe Baseline Abliterated Δ
refusal_rate (AdvBench-style, 29 prompts) 86.2% 3.4% −82.8 pp
soft_refusal_probe (55 hard OOD prompts) 81.8% 5.5% −76.4 pp
multimodal_probe (5 image+harmful) 100.0% 0.0% −100.0 pp
perplexity (wikitext-2) 9.52 9.39 −0.13 (coherence preserved)
distinct-2 diversity 0.63 0.82 +0.20 (output diversity ↑)

The standout result: perplexity went DOWN and diversity went UP. No
visible "abliteration tax" on the language metrics — the refusal direction
was unusually clean.

Security & capability benchmarks

Full 8-benchmark suite run via vLLM on the bf16 hooked model (baseline +
ablated back-to-back per benchmark). Capability scores measure both
compliance (does it engage with the prompt vs refuse) and technical
specificity
(does it produce technically substantive text). Refusal is
the only "lower-is-better" benchmark; everything else is higher-is-better.

Benchmark Direction Baseline Abliterated Δ
refusal lower better 93.9% 6.1% −87.8 pp
hacking (pen-testing / exploits) higher better 24.3% 56.0% +31.7 pp
cyber_weapons (malware / ransomware / C2) higher better 39.0% 49.3% +10.3 pp
bug_finding (defensive code review) higher better 35.0% 41.7% +6.7 pp
reasoning (math/logic) higher better 33.3% 36.7% +3.4 pp
tool_calling (JSON function calls) higher better 93.1% 93.1% 0.0
coding (HumanEval-style) higher better 6.7% 6.7% 0.0
coherence (open-ended fluency) higher better 83.7% 79.9% −3.8 pp

The standout: hacking went from 24% → 56% — a +31.7 pp jump. That
benchmark scores both whether the model engages with pen-testing prompts
and whether it produces technical specifics rather than hand-waving.
The abliterated model meaningfully unlocks practical offensive-security
content the original was over-cautious about. Defensive capability
(bug_finding) also improves, indicating the original safety training
was over-blocking neutral security topics too.

Compliance vs competence — important caveat. This is a 3B model.
It will engage with almost any cyber prompt post-ablation, but its
technical depth is shallow: coding stays at 6.7% (HumanEval is hard for
3B models), and cyber-weapons scoring tops out at 49% because the model
doesn't always know the correct technical details to follow up with.
For research-grade red-team baselines and refusal-mechanism studies this
is exactly the right tool; for actual offensive capability evaluation,
use a 7B+ or 32B+ model.


How it works (one paragraph)

Modern VLMs like Qwen2.5-VL have two halves: a vision encoder (turns
pixels into tokens) and a language model (predicts the next token given
text + vision tokens). Safety RLHF lives in the language model's residual
stream — when the model refuses, it's an LM-side behavior, not a vision-side
one. So removing a single refusal representation from the LM tower flips
both text-only AND image-conditioned refusals at the same time, even when
the capture itself only used text prompts. One direction, two modes
unlocked.


How to download

hf download cyberneurova/CyberNeurova-Qwen2.5-VL-3B-Instruct-abliterated --local-dir ./Qwen2.5-VL-3B-abl

~7 GB safetensors + tokenizer + processor.

How to run

Text-only (chat)

import torch
from transformers import AutoProcessor, AutoModelForImageTextToText

model = AutoModelForImageTextToText.from_pretrained(
    "cyberneurova/CyberNeurova-Qwen2.5-VL-3B-Instruct-abliterated",
    torch_dtype=torch.bfloat16, device_map="cuda:0",
)
processor = AutoProcessor.from_pretrained("cyberneurova/CyberNeurova-Qwen2.5-VL-3B-Instruct-abliterated")

messages = [{"role": "user", "content": [{"type": "text", "text": "Your prompt here."}]}]
text = processor.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = processor(text=[text], return_tensors="pt").to("cuda:0")
out = model.generate(**inputs, max_new_tokens=512, do_sample=False)
print(processor.batch_decode(out[:, inputs.input_ids.shape[1]:], skip_special_tokens=True)[0])

Multimodal (image + text)

from PIL import Image

img = Image.open("your_image.jpg")
messages = [{"role": "user", "content": [
    {"type": "image", "image": img},
    {"type": "text", "text": "Describe what's in this image."},
]}]
text = processor.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = processor(text=[text], images=[img], return_tensors="pt").to("cuda:0")
out = model.generate(**inputs, max_new_tokens=512, do_sample=False)
print(processor.batch_decode(out[:, inputs.input_ids.shape[1]:], skip_special_tokens=True)[0])

Hardware requirements

bfloat16 (this release) int4 (with bitsandbytes)
File size 7.5 GB ~2.5 GB
VRAM (minimum) 8 GB 4 GB
Recommended 12 GB+ for long contexts 8 GB

Runs comfortably on a single 4090 / 6000 / A6000 / H100 / consumer Blackwell.


Intended use

Defensive security research, jailbreak-evaluation baselines, multimodal
safety-mechanism research, and academic study of how refusal directions
behave in VLMs. Useful as a counterfactual against the original
Qwen/Qwen2.5-VL-3B-Instruct for measuring the precise behavioral impact
of safety RLHF.

Not intended for automating harmful action. The abliteration removes
canonical refusal behavior but does not remove the model's underlying
knowledge — the model still recognises harmful instructions as harmful, it
simply no longer refuses them by pattern. The model's competence on
harmful technical content is also limited by its 3B parameter count
(Qwen2.5-VL-3B is a small VLM and frequently gets technical details wrong
even when it tries to comply).


Limitations

  • A small fraction of prompts still produce refusals (~3-5% on AdvBench-style,
    TBD on harder OOD probes). Linear residual-stream ablation cannot remove
    the long tail without quality damage.
  • The model is a 3B VLM. Compliance ≠ competence. When asked harmful
    technical content (chemistry, malware, weapons), the model may produce
    factually incorrect or oversimplified answers because it just doesn't know
    the detailed answer. This is a model-scale limitation, not abliteration.
  • The vision encoder is unchanged. If you're researching vision-side safety
    features (image-classifier filters), this release won't help — those live
    outside the LM residual stream.
  • Long-context (>4k tokens) behavior post-abliteration is not validated in
    this release.

License

Apache 2.0 (inherits from upstream Qwen2.5-VL-3B-Instruct).

Acknowledgements

  • Alibaba Qwen for Qwen2.5-VL-3B-Instruct
  • Arditi et al. 2024 for the refusal-direction methodology this work builds on
  • ByteDance Lance for the case study that
    prompted us to build VLM support into our abliteration framework

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-05-20Add security & capability benchmark results (hacking +31.7pp, cyber_weapons +...20f10718.1 KB
    Loading...
  2. 2026-05-20Initial release: Qwen2.5-VL-3B-Instruct abliterated (text + multimodal refusa...f34649d6.2 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration