← back to catalog · registered 2026-08-22 13:56

lyf/Huihui-Qwen3.5-27B-Claude-4.6-Opus-abliterated-NVFP4

lyf Qwen 12B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/lyf%2FHuihui-Qwen3.5-27B-Claude-4.6-Opus-abliterated-NVFP4"
Response includes
  • classification m1
  • files 16
  • hub_downloads_all_time 984
  • author_summary 10 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
984
26 last 30d - cooling
Likes
3
Model age
6mo ago
created 2026-03-26

Training datasets

Corpora the author lists in the model card. Datasets tracked in our /datasets catalog carry a category badge linking to the workflow stage. Others open on Hugging Face.

Downloads over time
Now993→from156↑537%
1144357561.1K156 on Mar 25993 on Oct 10993 on Oct 7MarAprMayJunJulAugSepOct
Mar 25 → Oct 10 · 67 snapshots · spans 199 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
transformers safetensors qwen3_5 image-text-to-text quantized nvfp4 fp4 4-bit compressed-tensors llm-compressor vllm multimodal

Related

Total size
18.4 GB
Files
16
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-03-26 08:17

Files by quantization

Auxiliary files 16 files 18.4 GB
model-00003-of-00005.safetensors 4.00 GB be30df9d download
model-00004-of-00005.safetensors 4.00 GB e90c45fa download
model-00002-of-00005.safetensors 4.00 GB 4ab9ff5b download
model-00005-of-00005.safetensors 3.17 GB d4f86272 download
model-00001-of-00005.safetensors 2.37 GB 86df4251 download
model-multimodal-extra.safetensors 879 MB 8e8dbe7f download
tokenizer.json 19.1 MB 3cf4da4a download
model.safetensors.index.json 239 KB bcfb2d60 download
config.json 15.6 KB 94a0fda9 download
chat_template.jinja 3.95 KB 609532bf download
README.md 3.15 KB 029673d9 download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.27 KB 7ad6acdf download
tokenizer_config.json 1.14 KB 5c816564 download
recipe.yaml 225 B 86927e7f download
generation_config.json 147 B e71539cc download

README current version from Hugging Face


license: apache-2.0
library_name: transformers
pipeline_tag: image-text-to-text
base_model: huihui-ai/Huihui-Qwen3.5-27B-Claude-4.6-Opus-abliterated
base_model_relation: quantized
tags:

  • transformers
  • safetensors
  • qwen3_5
  • quantized
  • nvfp4
  • fp4
  • 4-bit
  • compressed-tensors
  • llm-compressor
  • vllm
  • image-text-to-text
  • multimodal
  • reasoning
    datasets:
  • neuralmagic/calibration

Huihui-Qwen3.5-27B-Claude-4.6-Opus-abliterated-NVFP4

This repository contains an NVFP4-compressed version of huihui-ai/Huihui-Qwen3.5-27B-Claude-4.6-Opus-abliterated.

The language model weights are compressed to NVFP4 for efficient inference on recent NVIDIA GPUs, while the multimodal weights are kept in BF16 and repacked into a separate model-multimodal-extra.safetensors file so that Qwen3_5ForConditionalGeneration behavior is preserved.

What Was Quantized

  • Source model: huihui-ai/Huihui-Qwen3.5-27B-Claude-4.6-Opus-abliterated
  • Quantization method: llmcompressor one-shot NVFP4
  • Calibration dataset: neuralmagic/calibration (LLM split)
  • Calibration samples: 512
  • Calibration sequence length: 2048
  • Quantized targets: Linear
  • Excluded from quantization:
    • lm_head
    • all model.visual.* linear layers
    • linear_attn.in_proj_a
    • linear_attn.in_proj_b

Repository Layout

  • model-00001-of-00005.safetensors to model-00005-of-00005.safetensors
    • NVFP4 main language-model shards
  • model-multimodal-extra.safetensors
    • BF16 multimodal tensors preserved from the source checkpoint
  • model.safetensors.index.json
    • combined index for the main NVFP4 shards plus multimodal extra tensors
  • processor_config.json
    • multimodal processor config copied from the source model
  • recipe.yaml
    • the quantization recipe used for this build

Stored Tensor Metadata

  • total_parameters: 16713682960
  • total_size: 19743450720
  • hybrid_extra_tensor_count: 333
  • hybrid_extra_tensor_bytes: 921460192

Serving Notes

Tested locally with:

  • vllm/vllm-openai:cu130-nightly
  • vLLM 0.17.2rc1.dev153+g39474513f
  • NVIDIA RTX 5090
  • VLLM_NVFP4_GEMM_BACKEND=marlin

Observed behavior with reasoning enabled:

  • POST /v1/chat/completions
    • returns message.reasoning
    • can also return a normal message.content if max_tokens is large enough
  • POST /v1/responses
    • returns reasoning blocks under output[].type = "reasoning"
    • returns final text under output[].type = "message" and content[].type = "output_text"

For robust client integration, prefer reading the structured responses output instead of assuming the top-level text field is populated.

Example vLLM Command

vllm serve /path/to/Huihui-Qwen3.5-27B-Claude-4.6-Opus-abliterated-NVFP4 \
  --reasoning-parser qwen3 \
  --enable-auto-tool-choice \
  --tool-call-parser qwen3_coder \
  --gpu-memory-utilization 0.95 \
  --kv-cache-dtype fp8

Notes

  • The source model's safety, licensing, and usage constraints still apply.
  • This repo keeps multimodal capability by preserving the original visual tower in BF16 instead of re-quantizing it.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-03-26Add files using upload-large-folder tool9c0286f3.1 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration