← back to catalog · registered 2026-08-22 13:56

nemozxy123/Huihui-Qwen3-VL-8B-Thinking-abliterated-AWQ-W4A16

nemozxy123 Qwen 6.9B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/nemozxy123%2FHuihui-Qwen3-VL-8B-Thinking-abliterated-AWQ-W4A16"
Response includes
  • classification m1
  • files 17
  • hub_downloads_all_time 76
  • author_summary 5 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
76
21 last 30d - stable
Likes
1
Model age
4mo ago
created 2026-05-30
Downloads over time
Now90→from44↑105%
4259779544 on Jun 1090 on Oct 11JunJulAugSepOct
Jun 10 → Oct 11 · 57 snapshots · spans 123 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
transformers safetensors qwen3_vl image-text-to-text awq 4bit quantized compressed-tensors abliterated W4A16 conversational base_model:huihui-ai/Huihui-Qwen3-VL-8B-Thinking-abliterated

Related

Total size
7.03 GB
Files
17
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-05-30 18:16

Files by quantization

Auxiliary files 17 files 7.05 GB
model-00001-of-00002.safetensors 4.64 GB 4f0123ce download
model-00002-of-00002.safetensors 2.40 GB 60df3c7e download
tokenizer.json 10.9 MB aeb13307 download
vocab.json 2.65 MB 4783fe10 download
merges.txt 1.59 MB 31349551 download
model.safetensors.index.json 118 KB 06d53336 download
config.json 7.44 KB 1bf5a978 download
tokenizer_config.json 5.32 KB fec7f182 download
chat_template.jinja 5.08 KB 551cd6a9 download
README.md 2.96 KB 3a099d26 download
.gitattributes 1.53 KB 52373fe2 download
recipe.yaml 1.51 KB 557a4c50 download
video_preprocessor_config.json 858 B 8deea1de download
preprocessor_config.json 821 B b7d11200 download
added_tokens.json 707 B b54f9135 download
special_tokens_map.json 613 B ac23c0aa download
generation_config.json 192 B e54e6d2f download

README current version from Hugging Face


license: apache-2.0
tags:

  • awq
  • 4bit
  • quantized
  • qwen3_vl
  • compressed-tensors
  • abliterated
  • W4A16
    library_name: transformers
    base_model:
  • huihui-ai/Huihui-Qwen3-VL-8B-Thinking-abliterated
    pipeline_tag: image-text-to-text

Huihui-Qwen3-VL-8B-Thinking-abliterated-AWQ-4bit

This is an AWQ 4-bit quantized version of huihui-ai/Huihui-Qwen3-VL-8B-Thinking-abliterated.

The primary goal of this quantization is to retain the original model's video analysis capabilities. By converting the model to AWQ-4bit, it becomes possible to launch the model with vLLM and directly pass video inputs—preserving temporal information and continuous-frame understanding. In contrast, llama.cpp-based solutions rely on frame sampling, which inevitably discards fine-grained temporal dynamics and inter-frame coherence.

Quantization details

The quantization configuration (layer selection, etc.) follows cyankiwi/Qwen3.5-9B-AWQ-4bit.

The calibration dataset used for AWQ is mit-han-lab/pile-val-backup.

You can find the original quantization script in the model repository. This is my first time doing something like this.

Running 65K context on 16GB VRAM

On an RTX 5060 Ti 16GB (Blackwell architecture), the 65,536-token context window can be used by enabling FP8 KV-cache quantization (set --kv-cache-dtype fp8" when loading, requires a compatible vLLM version).

Note on GPU architectures: This has been tested and confirmed working on the RTX 5060 Ti. Due to architectural differences, the same cannot be guaranteed for RTX 40-series or older 16GB GPUs when vision capabilities are also loaded—OOM (out of memory) is still possible. Adjust batch size and context length accordingly.

Example vLLM launch command

Below is the launch configuration I use on Windows. Replace the model path and media directory with your own.

set VIDEO_MAX_PIXELS=200704
set FPS=2.0
set FPS_MAX_FRAMES=2590
set FPS_MIN_FRAMES=4
set FORCE_QWENVL_VIDEO_READER=torchcodec

python -m vllm.entrypoints.openai.api_server ^
    --model /path/to/your/model/Huihui-Qwen3-VL-8B-Thinking-abliterated-AWQ-W4A16 ^
    --served-model-name Huihui-Qwen3-VL-8B-Thinking-abliterated-AWQ-W4A16 ^
    --trust-remote-code ^
    --enforce-eager ^
    --dtype auto ^
    --max-model-len 65536 ^
    --kv-cache-dtype fp8 ^
    --gpu-memory-utilization 0.92 ^
    --port 8000 ^
    --allowed-local-media-path /path/to/your/media

Acknowledgements

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-05-30Update README.md9368a9c3 KB
    Loading...
  2. 2026-05-30initial commitb0b227228 B
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration