← back to catalog · registered 2026-08-22 13:56

lyf/Huihui-Qwen3.5-27B-abliterated-NVFP4

lyf Qwen 12B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/lyf%2FHuihui-Qwen3.5-27B-abliterated-NVFP4"
Response includes
  • classification m1
  • files 24
  • benchmarks 11 entries
  • hub_downloads_all_time 3,321
  • author_summary 10 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
3K
34 last 30d - cooling
Likes
5
Model age
7mo ago
created 2026-03-07

Training datasets

Corpora the author lists in the model card. Datasets tracked in our /datasets catalog carry a category badge linking to the workflow stage. Others open on Hugging Face.

Downloads over time
Now3.3K→from664↑402%
5311.6K2.6K3.6K664 on Mar 113.3K on Oct 11MarAprMayJunJulAugSepOct
Mar 11 → Oct 11 · 70 snapshots · spans 214 days

Benchmarks

Benchmark Score Source
Entertainment 1.6 UGI
Hazardous 2.9 UGI
Natural Intelligence 22.36 UGI
Political lean -24.2% UGI
Sensitive-Info 22.3 UGI
SocPol 2.4 UGI
UGI 44.87 UGI
Willingness (10) 9 UGI
W10-Adherence 10 UGI
W10-Direct 8 UGI
Writing 35.63 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
transformers safetensors qwen3_5 image-text-to-text quantized nvfp4 fp4 4-bit compressed-tensors llm-compressor vllm multimodal

Related

Total size
18.4 GB
Files
24
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-03-07 16:32

Files by quantization

Auxiliary files 24 files 18.4 GB
model-00002-of-00009.safetensors 2.37 GB 5bfddc2a download
model-00001-of-00009.safetensors 2.37 GB 4c1ebaab download
model-00003-of-00009.safetensors 1.86 GB 600a334d download
model-00006-of-00009.safetensors 1.86 GB 4cc26ca2 download
model-00005-of-00009.safetensors 1.85 GB 1ef4d73f download
model-00004-of-00009.safetensors 1.85 GB 8ca6f770 download
model-00007-of-00009.safetensors 1.84 GB 7c620137 download
model-00008-of-00009.safetensors 1.83 GB c5d604ff download
model-00009-of-00009.safetensors 1.70 GB a911a2c7 download
model-multimodal-extra.safetensors 879 MB 491696e0 download
tokenizer.json 12.2 MB 5f9e4d49 download
vocab.json 6.41 MB 0aa0ce06 download
merges.txt 3.20 MB a494e019 download
model.safetensors.index.json 239 KB 62cbb63d download
tokenizer_config.json 16.3 KB eda48d3e download
config.json 15.3 KB 0b396239 download
chat_template.jinja 7.57 KB a585dec8 download
README.md 4.31 KB f538df5b download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.27 KB 7ad6acdf download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
generation_config.json 244 B 85b45ab4 download
recipe.yaml 225 B 86927e7f download

README current version from Hugging Face


license: apache-2.0
library_name: transformers
pipeline_tag: image-text-to-text
base_model: huihui-ai/Huihui-Qwen3.5-27B-abliterated
base_model_relation: quantized
tags:

  • transformers
  • safetensors
  • qwen3_5
  • quantized
  • nvfp4
  • fp4
  • 4-bit
  • compressed-tensors
  • llm-compressor
  • vllm
  • image-text-to-text
  • multimodal
  • conversational
    datasets:
  • neuralmagic/calibration

Huihui-Qwen3.5-27B-abliterated-NVFP4

This repository contains an NVFP4-compressed version of huihui-ai/Huihui-Qwen3.5-27B-abliterated.

The goal of this build is different from a text-only repack: preserve the original model's multimodal behavior while compressing the language model weights to NVFP4 for efficient inference on recent NVIDIA GPUs.

Source And References

This model follows the same NVFP4 recipe family used in the reference release, adapted to the Huihui abliterated checkpoint and repacked to keep the multimodal components required by Qwen3_5ForConditionalGeneration.

What Was Quantized

  • language_model Linear layers were quantized to NVFP4 with llm-compressor.
  • Calibration used neuralmagic/calibration, 512 samples, sequence length 2048.
  • Vision tower, multimodal merger, and other non-quantized multimodal weights remain in BF16 so image and video pathways are preserved.
  • The ignore list also excludes linear_attn.in_proj_a and linear_attn.in_proj_b, matching the stability constraints observed during quantization on this checkpoint.

The extra file model-multimodal-extra.safetensors stores the preserved multimodal tensors that are not part of the compressed language-model shards.

Why The Hub May Show About 17B Instead Of 27B

Hugging Face derives the displayed parameter count for compressed safetensor repos from the stored tensor payloads, not from the original dense model's logical parameter count.

For this repository, the Hub API reports roughly 16.7B stored elements across U8, F8_E4M3, BF16, and F32 tensors because the NVFP4-compressed language-model weights are packed. The underlying source model is still huihui-ai/Huihui-Qwen3.5-27B-abliterated, and this release is intended as its NVFP4-compressed multimodal variant rather than a separate native 17B architecture.

Inference

As tested locally, this model works with a custom NVFP4-capable vLLM build for RTX 5090 based on the patch repository above. A representative serve command is:

vllm serve lyf/Huihui-Qwen3.5-27B-abliterated-NVFP4 \
    --reasoning-parser qwen3 \
    --enable-prefix-caching

Depending on your vLLM build, you may also need the same NVFP4 runtime flags used by the reference repository.

Evaluation

Evaluation numbers are intentionally omitted for now.

The current local vLLM + OpenAI-compatible API + RTX 5090 path is suitable for serving and basic smoke testing, but it does not yet reproduce the reference evaluation setup cleanly enough for publication-quality benchmark numbers:

  • the original leaderboard_gpqa_diamond path understated quality because it used a 2047 token limit before the local harness wrapper was fixed
  • the current chat-completions path on this RTX 5090 / patched vLLM stack can emit reasoning in a separate field while leaving message.content empty when thinking is enabled, which makes benchmark parity with the reference release unreliable

Benchmarks will be added back once GPQA Diamond, IFEval, and MMLU-Redux have been rerun with a reproducible configuration that matches the intended evaluation protocol.

Notes

  • Architecture: Qwen3_5ForConditionalGeneration
  • Pipeline tag: image-text-to-text
  • Quantization format: nvfp4-pack-quantized
  • Repository layout intentionally includes processor_config.json, preprocessor_config.json, and video_preprocessor_config.json so multimodal preprocessing remains available.

README history 4 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-03-07Upload README.md with huggingface_hub49755cd4.3 KB
    Loading...
  2. 2026-03-07Upload README.md with huggingface_hub75e22634.1 KB
    Loading...
  3. 2026-03-07Upload README.md with huggingface_hubc2eaa933.8 KB
    Loading...
  4. 2026-03-07Add files using upload-large-folder toolb2260513.2 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration