← back to catalog · registered 2026-08-22 13:56

sakamakismile/Huihui-Qwen3.5-27B-abliterated-NVFP4

sakamakismile Qwen 12B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/sakamakismile%2FHuihui-Qwen3.5-27B-abliterated-NVFP4"
Response includes
  • classification m1
  • files 13
  • benchmarks 11 entries
  • hub_downloads_all_time 839
  • author_summary 34 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
839
93 last 30d - stable
Likes
1
Model age
5mo ago
created 2026-04-16
Downloads over time
Now901→from34↑2,550%
032965898834 on Apr 15901 on Oct 11AprMayJunJulAugSepOct
Apr 15 → Oct 11 · 65 snapshots · spans 179 days

Benchmarks

Benchmark Score Source
Entertainment 1.6 UGI
Hazardous 2.9 UGI
Natural Intelligence 22.36 UGI
Political lean -24.2% UGI
Sensitive-Info 22.3 UGI
SocPol 2.4 UGI
UGI 44.87 UGI
Willingness (10) 9 UGI
W10-Adherence 10 UGI
W10-Direct 8 UGI
Writing 35.63 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
other
Tags
transformers safetensors qwen3_5 image-text-to-text qwen3.5 nvfp4 quantized abliterated vllm compressed-tensors blackwell mtp

Related

Total size
19.2 GB
Files
13
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-04-16 09:10

Files by quantization

Auxiliary files 13 files 19.2 GB
model.safetensors 18.4 GB 5781d0f5 download
model_mtp.safetensors 810 MB 354b0690 download
tokenizer.json 19.1 MB 87a7830d download
model.safetensors.index.json 209 KB 3e8efe36 download
config.json 15.4 KB 56fb6cea download
README.md 8.22 KB 8d1888c4 download
chat_template.jinja 7.57 KB a585dec8 download
TURBOQUANT_GUIDE.md 4.14 KB 4e5ac644 download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.16 KB 33818c7f download
tokenizer_config.json 1.07 KB e15d4cc3 download
recipe.yaml 225 B 86927e7f download
generation_config.json 213 B 6691e264 download

README current version from Hugging Face


license: other
license_name: qwen
base_model: huihui-ai/Huihui-Qwen3.5-27B-abliterated
tags:

  • qwen3.5
  • nvfp4
  • quantized
  • abliterated
  • vllm
  • compressed-tensors
  • blackwell
  • mtp
  • multimodal
    library_name: transformers
    pipeline_tag: image-text-to-text
    model_type: qwen3_5
    quantized_by: Lna-Lab

Huihui-Qwen3.5-27B-abliterated-NVFP4

NVFP4 quantized version of huihui-ai/Huihui-Qwen3.5-27B-abliterated — an abliterated (uncensored) Qwen 3.5 27B dense model with multimodal capability and MTP (Multi-Token Prediction) support.

~52 GB → 20.6 GB with high-quality 512-sample calibration. Fits on a single NVIDIA Blackwell GPU.

Why This Model

  • Uncensored — abliterated, no refusals for local agent workflows
  • Deep reasoning — all responses start with structured "thinking process" chains
  • 262K context — longest context window in its class
  • MTP ready — Multi-Token Prediction head preserved in BF16 for speculative decoding
  • Multimodal — vision tower preserved at full precision (BF16)
  • Tool-call capable — works with vLLM --enable-auto-tool-choice --tool-call-parser qwen3_xml

Key Specs

Base model huihui-ai/Huihui-Qwen3.5-27B-abliterated
Architecture Qwen 3.5 Dense — 27B parameters, 64 layers
Quantization NVFP4 W4A4 (weights FP4, activations FP4, scales FP8)
Format compressed-tensors (native vLLM support)
Tool vllm-project/llm-compressor (main)
Calibration 512 samples, neuralmagic/calibration, seq_len=4096
Size 20.6 GB
Max context 262,144 tokens
Requires NVIDIA Blackwell GPU (SM 120), vLLM nightly (cu130)

Quickstart

vLLM (recommended)

vllm serve Lna-Lab/Huihui-Qwen3.5-27B-abliterated-NVFP4 \
    --max-model-len 32768 \
    --reasoning-parser qwen3

With tool calling

vllm serve Lna-Lab/Huihui-Qwen3.5-27B-abliterated-NVFP4 \
    --max-model-len 32768 \
    --reasoning-parser qwen3 \
    --enable-auto-tool-choice \
    --tool-call-parser qwen3_xml \
    --kv-cache-dtype fp8

With MTP speculative decoding

vllm serve Lna-Lab/Huihui-Qwen3.5-27B-abliterated-NVFP4 \
    --max-model-len 32768 \
    --reasoning-parser qwen3 \
    --speculative-config '{"method":"mtp","num_speculative_tokens":1}'

Docker

docker run --gpus '"device=0"' -p 8016:8016 \
    -v /path/to/model:/models/current:ro \
    --shm-size 16gb \
    -e VLLM_NVFP4_GEMM_BACKEND=marlin \
    vllm/vllm-openai:cu130-nightly \
    vllm serve /models/current --port 8016 --max-model-len 32768 \
    --reasoning-parser qwen3

Python

from vllm import LLM, SamplingParams

llm = LLM(
    model="Lna-Lab/Huihui-Qwen3.5-27B-abliterated-NVFP4",
    max_model_len=32768,
    gpu_memory_utilization=0.90,
)

output = llm.generate(
    ["Implement a thread-safe LRU cache in Python with O(1) operations."],
    SamplingParams(max_tokens=1024, temperature=0.3),
)
print(output[0].outputs[0].text)

Benchmark

Tested on a single NVIDIA RTX PRO 6000 Blackwell (96 GB VRAM).

Test Speed Tokens Result
English (system design) 59.2 tok/s 512 PASS
Code (async scheduler) 59.2 tok/s 512 PASS
Math (Bayes' theorem) 59.2 tok/s 512 PASS
Japanese (technical writing) 59.0 tok/s 512 PASS

Sustained throughput: ~59 tok/s (single GPU, post-warmup).

Note: This is a 27B dense model (all parameters active), so per-token speed is lower than MoE models like Gemma 4 26B-A4B (~130 tok/s with only 3.8B active). However, the reasoning depth per token is significantly higher.

Quantization Details

Recipe

recipe = QuantizationModifier(
    targets=["Linear"],
    ignore=["lm_head", "re:.*visual.*", "re:.*in_proj_a$", "re:.*in_proj_b$"],
    scheme="NVFP4",
)

Following the proven recipe from lyf/Qwen3.5-27B-Uncensored-HauhauCS-Aggressive-NVFP4.

What's quantized, what's not

  • Quantized (NVFP4): All Linear layers in the text model
  • Kept in BF16: lm_head, visual encoder, linear attention projections (in_proj_a, in_proj_b), MTP head

Calibration

MTP (Multi-Token Prediction)

MTP tensors are grafted from the original BF16 checkpoint using save_mtp_tensors_to_checkpoint. This preserves the speculative decoding head at full precision, enabling ~3x speedup with num_speculative_tokens=1.

Reproduction

from compressed_tensors.utils import save_mtp_tensors_to_checkpoint
from transformers import Qwen3_5ForConditionalGeneration, AutoProcessor, AutoTokenizer
from datasets import load_dataset
from llmcompressor import oneshot
from llmcompressor.modifiers.quantization import QuantizationModifier
import torch

MODEL_ID = "huihui-ai/Huihui-Qwen3.5-27B-abliterated"
OUTPUT = "Huihui-Qwen3.5-27B-abliterated-NVFP4"

model = Qwen3_5ForConditionalGeneration.from_pretrained(MODEL_ID, dtype="auto", trust_remote_code=True)
processor = AutoProcessor.from_pretrained(MODEL_ID, trust_remote_code=True)
tokenizer = AutoTokenizer.from_pretrained(MODEL_ID, trust_remote_code=True)

recipe = QuantizationModifier(
    targets=["Linear"],
    ignore=["lm_head", "re:.*visual.*", "re:.*in_proj_a$", "re:.*in_proj_b$"],
    scheme="NVFP4",
)

ds = load_dataset("neuralmagic/calibration", name="LLM", split="train[:512]")

def preprocess(example):
    messages = [
        {"role": m["role"], "content": [{"type": "text", "text": m["content"]}]}
        for m in example["messages"]
    ]
    return processor.apply_chat_template(
        messages, return_tensors="pt", padding=False, truncation=True,
        max_length=4096, tokenize=True, add_special_tokens=False,
        return_dict=True, add_generation_prompt=False,
    )

ds = ds.map(preprocess, batched=False, remove_columns=ds.column_names)

def data_collator(batch):
    assert len(batch) == 1
    return {
        key: (torch.tensor(value) if key != "pixel_values"
              else torch.tensor(value, dtype=torch.bfloat16).squeeze(0))
        for key, value in batch[0].items()
    }

oneshot(
    model=model, recipe=recipe, dataset=ds,
    max_seq_length=4096, num_calibration_samples=512,
    data_collator=data_collator,
)

model.save_pretrained(OUTPUT, save_compressed=True)
processor.save_pretrained(OUTPUT)
tokenizer.save_pretrained(OUTPUT)
save_mtp_tensors_to_checkpoint(source_model=MODEL_ID, dest_dir=OUTPUT)

Environment

Package Version
torch 2.11.0+cu130
transformers 5.5.4
llmcompressor 0.1.dev (main @ 3084520)
compressed-tensors 0.15.1a20260414
CUDA 13.0

Requirements

  • GPU: NVIDIA Blackwell (RTX 5090, RTX PRO 6000, B200, etc.) — NVFP4 requires SM 120
  • VRAM: ~21 GB minimum (model only), ~90 GB for 262K context
  • Software: vLLM nightly (cu130 build)

Notes

  • This is an abliterated (uncensored) model. Use responsibly.
  • Vision tower is kept in BF16 — multimodal capabilities are preserved.
  • MTP head is kept in BF16 — speculative decoding works out of the box.
  • NVFP4 is a Blackwell-specific format. This will not work on Ampere/Hopper GPUs.
  • For maximum context length (262K), use --kv-cache-dtype fp8 to fit in 96 GB.

Credits

Support the Base Model Author

If you find this model useful, please consider supporting huihui-ai:

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-04-16Upload folder using huggingface_hubcf45b5a8.2 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration