← back to catalog · registered 2026-08-22 13:56

ahmed22xa/Huihui-Qwen3-VL-4B-Thinking-abliterated-comfy

ahmed22xa Qwen 4B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/ahmed22xa%2FHuihui-Qwen3-VL-4B-Thinking-abliterated-comfy"
Response includes
  • classification m1
  • files 5
  • author_summary 4 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · 30-day
0
Likes
6
Model age
3mo ago
created 2026-06-24
Downloads over time
Now0→from0↑0%
00110 on Jun 240 on Oct 11JunJulAugSepOct
Jun 24 → Oct 11 · 55 snapshots · spans 109 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Quantizations
FP8
Tags
transformers qwen3_vl thinking reasoning vision-language abliterated uncensored safetensors comfyui fp8 int8 convrot

Related

Total size
17.9 GB
Files
5
Quantizations
2
Registered
2026-08-22 13:56
Last updated on HF
2026-08-09 19:53

Files by quantization

FP8 1 file 4.88 GB
Huihui-Qwen3-VL-4B-Thinking-abliterated-fp8_scaled.safetensors 4.88 GB 4131743b download
Auxiliary files 4 files 13.0 GB
Huihui-Qwen3-VL-4B-Thinking-abliterated.safetensors 8.27 GB 6d04ddac download
Huihui-Qwen3-VL-4B-Thinking-abliterated-int8_convrot.safetensors 4.72 GB 288381bb download
README.md 8.38 KB 35cbf863 download
.gitattributes 1.48 KB a6344aac download

README current version from Hugging Face


license: apache-2.0
base_model: huihui-ai/Huihui-Qwen3-VL-4B-Thinking-abliterated
base_model_relation: quantized
tags:

  • qwen3_vl
  • thinking
  • reasoning
  • vision-language
  • abliterated
  • uncensored
  • safetensors
  • comfyui
  • fp8
  • int8
  • convrot
  • transformers
    pipeline_tag: image-text-to-text
    library_name: transformers

Huihui-Qwen3-VL-4B-Thinking-abliterated — ComfyUI Edition

This repo packages the huihui-ai/Huihui-Qwen3-VL-4B-Thinking-abliterated thinking Vision-Language Model in two ready-to-use formats for ComfyUI:

File Size Format Use case
Huihui-Qwen3-VL-4B-Thinking-abliterated.safetensors 8.88 GiB BF16 single safetensors Maximum fidelity / training / full-precision workflows
Huihui-Qwen3-VL-4B-Thinking-abliterated-fp8_scaled.safetensors 5.24 GiB FP8 (E4M3FN) per-tensor scaled ComfyUI Qwen3-VL Text Encoder node
Huihui-Qwen3-VL-4B-Thinking-abliterated-int8_convrot.safetensors 4.72 GiB INT8 ConvRot (per-row scaled) Native INT8 tensor cores — recommended on RTX 30/40/50

Source / Provenance

What is a "Thinking" model?

The Thinking edition is trained to perform explicit step-by-step reasoning (chain-of-thought) before producing the final answer. Output text typically contains a ... block with the reasoning trace, then the answer. Useful for:

  • Complex captioning where you want the model to reason about composition, lighting, intent
  • Vision-grounded instruction following
  • Multi-step visual analysis tasks
  • For ComfyUI: getting better prompt-engineering suggestions and structured captions

What was done

BF16 single-file (*.safetensors, 8.88 GiB)

Original upstream repo splits the weights into two safetensors shards (model-00001-of-00002.safetensors + model-00002-of-00002.safetensors). They were merged into a single safetensors file using the original model.safetensors.index.json mapping. No weights modified.

  • 713 tensors
  • dtype: bfloat16
  • Verified structurally identical to upstream (same key set)

FP8 scaled (*-fp8_scaled.safetensors, 5.24 GiB)

Per-tensor abs-max quantization to float8_e4m3fn for the 252 linear projections of the language model (q/k/v/o_proj + gate/up/down_proj across all layers). Embeddings, layer norms, biases and the entire visual encoder stay in BF16.

  • 1217 tensors (252 × 3 + 461 BF16)
  • Quantised layers: float8_e4m3fn weights + float32 per-tensor scale + uint8[64] comfy_quant marker (JSON: {"format": "float8_e4m3fn", "full_precision_matrix_mult": false})
  • Per-tensor scale = max(|w|) / 448 (E4M3FN max)
  • Mean round-trip relative error ≈ 2.3% (typical for FP8 LLM quantisation)
  • Schema matches the ComfyUI "fp8_scaled" convention used by other Qwen3-VL models in this size class.

INT8 ConvRot (*-int8_convrot.safetensors, 4.72 GiB)

Row-wise INT8 quantization of the 252 linear projections of the language model, with Hadamard rotation (group size 256) applied to the weight matrices to distribute outliers before quantization. Embeddings, layer norms, biases, the entire visual encoder and the first transformer block stay in BF16 (protected by the --qwen35 filter). Quantized directly from the BF16 source (not from FP8) using silveroxides/convert_to_quant.

  • 1724 tensors (358 quantized × 3 + 650 BF16)
  • Per-tensor .comfy_quant JSON marker: {"format": "int8_tensorwise", "orig_dtype": "torch.bfloat16", "convrot": true, "convrot_groupsize": 256, "per_row": true}
  • weight_scale shape: per-row (N, 1), not scalar — the per-row scale is what makes ConvRot compatible with LoRA application at runtime
  • Quantization command (reproducible):
    ctq -i Huihui-Qwen3-VL-4B-Thinking-abliterated.safetensors \
        -o Huihui-Qwen3-VL-4B-Thinking-abliterated-int8_convrot.safetensors \
        --int8 --convrot --convrot-group-size 256 \
        --scaling_mode row \
        --comfy_quant --save-quant-metadata --qwen35 \
        --simple --low-memory --device cuda
    
  • Quantization was done on GPU (NVIDIA RTX 3090) in ~3 minutes.

Why INT8 ConvRot over FP8 on RTX 30-series

On Ampere GPUs (RTX 3090) FP8 is emulated in software because Ampere has no native FP8 tensor cores. INT8 has native INT8 tensor cores on every RTX 30+ generation. Result on RTX 3090: ~20-30% faster text-encoder forward pass than the FP8-scaled variant, with comparable quality.

Why --scaling_mode row matters

The default --scaling_mode tensor produces a single global scale per weight matrix (weight_scale shape ()). For ConvRot's per-row Hadamard rotation to remain correct on the quantization side, per-row scales are required (weight_scale shape (N, 1), one scale per output channel). Without --scaling_mode row, the file loads but LoRAs applied at runtime silently degrade to plain tensorwise INT8. Always verify after conversion:

from safetensors import safe_open
import json
with safe_open("…-int8_convrot.safetensors", framework="pt") as f:
    raw = f.get_tensor([k for k in f.keys() if k.endswith('.comfy_quant')][0]).tolist()
    print(json.loads(bytes(raw)))
# Must contain: convrot=True, per_row=True, convrot_groupsize=256

Quantisation script

FP8 conversion was done on GPU (NVIDIA RTX 3090) in ~4 seconds.

Usage

ComfyUI — Qwen3-VL Text Encoder (recommended)

  1. Drop the *-fp8_scaled.safetensors into your ComfyUI models/text_encoders/ directory.
  2. Use the Qwen3-VL Text Encoder node and select Huihui-Qwen3-VL-4B-Thinking-abliterated-fp8_scaled.
  3. Pair with a Qwen3-VL compatible diffusion model and sampler.

ComfyUI — INT8 ConvRot (recommended on RTX 30/40/50)

  1. Drop the *-int8_convrot.safetensors into your ComfyUI models/text_encoders/ directory.
  2. Use the native Qwen3-VL Text Encoder node and select the INT8 file. (ComfyUI ≥ 0.27.0 supports INT8 ConvRot natively; no custom node required for this text-encoder path.)
  3. ~20-30% faster forward pass than the FP8-scaled variant on RTX 30-series (native INT8 tensor cores). Pair with any Qwen3-VL compatible workflow — this encoder is not Krea-specific.

ComfyUI — Full BF16 (when more precision is required)

  1. Drop the *.safetensors into models/text_encoders/.
  2. Same node, select the BF16 file. Higher VRAM usage.

transformers (BF16 only — tokenizer/configs are not bundled here)

The upstream repo huihui-ai/Huihui-Qwen3-VL-4B-Thinking-abliterated has the matching tokenizer, processor and configs. For BF16 inference:

from transformers import AutoProcessor, Qwen3VLForConditionalGeneration
import torch

model = Qwen3VLForConditionalGeneration.from_pretrained(
    "ahmed22xa/Huihui-Qwen3-VL-4B-Thinking-abliterated-comfy",
    torch_dtype=torch.bfloat16,
    device_map="auto",
)
processor = AutoProcessor.from_pretrained("huihui-ai/Huihui-Qwen3-VL-4B-Thinking-abliterated")

The FP8 file is not loadable with transformers.from_pretrained directly — it follows ComfyUI's per-tensor-FP8 layout with comfy_quant markers.

Sibling repos

The non-thinking ("Instruct") sibling is also packaged the same way:

License & disclaimer

  • License: apache-2.0 (inherited from upstream Qwen/Qwen3-VL-4B-Thinking).
  • Abliteration notice: This is an uncensored variant. The safety filtering has been significantly reduced, potentially generating sensitive, controversial, or inappropriate content. Use with caution. See the upstream model card for the full disclaimer.
  • No warranty. Users are solely responsible for any consequences arising from use of this model.

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-09Upload README.md with huggingface_hub6a5e14a8.4 KB
    Loading...
  2. 2026-06-24Add model card with source attribution, base_model link and thinking-edition ...d1e5f285.3 KB
    Loading...

Discussions 1 thread

  1. 2026-06-30Thank you!open1 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration