← back to catalog · registered 2026-08-22 13:56

TelperionAI/Huihui-Qwen3.8-27B-abliterated-NVFP4

TelperionAI Qwen 9.4B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/TelperionAI%2FHuihui-Qwen3.8-27B-abliterated-NVFP4"
Response includes
  • classification m1
  • files 16
  • hub_downloads_all_time 1,392
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
1K
566 last 30d - stable
Likes
2
Model age
7w ago
created 2026-08-21
Downloads over time
Now1.7K→from26↑6,335%
06131.2K1.8K26 on Aug 191.7K on Oct 11AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
transformers safetensors qwen3_5 image-text-to-text nvfp4 fp4 awq autoround abliterated llm-compressor compressed-tensors vllm

Related

Total size
23.0 GB
Files
16
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-21 06:12

Files by quantization

Auxiliary files 16 files 23.0 GB
model-00001-of-00003.safetensors 18.6 GB ca4c3746 download
model-00002-of-00003.safetensors 3.59 GB adb55493 download
model-00003-of-00003.safetensors 810 MB 90fa0e3e download
tokenizer.json 19.1 MB 06b95093 download
vocab.json 6.41 MB 0aa0ce06 download
merges.txt 3.20 MB a494e019 download
model.safetensors.index.json 186 KB ce95083a download
config.json 21.7 KB c7338e70 download
chat_template.jinja 8.74 KB c0c686f9 download
README.md 4.81 KB 38cb5e7d download
recipe.yaml 2.74 KB 0eea66ff download
.gitattributes 1.53 KB 52373fe2 download
tokenizer_config.json 1.10 KB ed1f99f3 download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
generation_config.json 214 B 0bc3addd download

README current version from Hugging Face


license: apache-2.0
base_model: huihui-ai/Huihui-Qwen3.8-27B-abliterated
tags:

  • nvfp4
  • fp4
  • awq
  • autoround
  • abliterated
  • llm-compressor
  • compressed-tensors
  • vllm
    library_name: transformers

Huihui-Qwen3.8-27B-abliterated-NVFP4

Mixed-precision NVFP4 quantization of
huihui-ai/Huihui-Qwen3.8-27B-abliterated,
built with llm-compressor.

24.7 GB. Calibrated on text generated by this abliterated model itself, not by
stock Qwen — see below, it matters.

Recipe

component precision
mlp.{gate,up,down}_proj, layers 0–55 NVFP4 (4-bit, group-16, FP8-e4m3 scales)
mlp.{gate,up,down}_proj, layers 56–63 FP8 e4m3
self_attn.{q,k,v,o}_proj FP8 e4m3
linear_attn.{in_proj_qkv,in_proj_z,out_proj} (GDN) FP8 e4m3
lm_head, embed_tokens, norms, GDN state params, vision tower BF16

AWQ per-input-channel scaling, then AutoRound (SignSGD, block-wise loss, 200 iters)
on the NVFP4 MLPs and GPTQ on the 8-bit modules. Requires Blackwell for native NVFP4.

Calibration: self-distilled from the abliterated model on a balanced Nemotron-v2 prompt
blend (25% code, 25% math, 20% STEM, 20% chat, 10% multilingual).

Benchmarks

Measured against the abliterated BF16 model as its own reference — not stock Qwen —
so the numbers reflect quantization damage only, not the effect of abliteration.
142,727 tokens plus 200 free greedy generations. vLLM 0.27.1, TP=2, 2×B300.

build size ↓ top-1 ↑ near-tie ↓ moderate ↓ confident ↓ certain ↓ divmed ↑ tok/s ↑
this model (NVFP4 AWQ+AutoRound) 24.7 GB 92.98% 34.02% 9.83% 1.80% 0.20% 27 10702
INT4 sibling (AWQ+GPTQ) 25.1 GB 96.45% 21.76% 2.92% 0.88% 0.12% 41 4551
earlier build (base-model calibration) 24.7 GB 91.79% 37.58% 11.21% 3.54% 0.21% 20 10685

Sizes are on-disk tensor bytes and include the ~0.85 GB BF16 MTP head.

Columns. top-1 is raw argmax agreement with the BF16 abliterated model. The bucket
columns are disagreement rates split by how confident the reference was at that position
(top1−top2 logprob margin): near-tie <0.5, moderate 0.5–2, confident 2–5, certain >5.
Only confident and certain are real damage. divmed is the median token index at
which free greedy generation first diverges.

Perplexity is excluded — on this model family it is anti-correlated with quality.

Calibration matters more than abliteration

An earlier build of this model used the stock-Qwen calibration set and a weaker recipe, and
landed at 3.54% confident damage. Regenerating the calibration from the abliterated model
itself brings that to 1.80%.

That also answers a question worth stating plainly: abliterated weights are not intrinsically
harder to quantize.
With matched recipe and self-distilled calibration this model reaches
1.80% confident damage, against 1.85% for the same recipe on stock
Qwen3.8-27B. The earlier gap was the calibration and recipe, not the abliteration.

NVFP4 vs INT4 on this model

The INT4 sibling
is more faithful (confident 0.88% vs 1.80%) but decodes at 4551 tok/s against 10702 here.
This build is the throughput choice on Blackwell; the INT4 one is the fidelity choice, and the
only option on Ampere/Ada where FP4 does not exist.

Usage

from vllm import LLM
llm = LLM("TelperionAI/Huihui-Qwen3.8-27B-abliterated-NVFP4", tensor_parallel_size=2)

Speculative decoding (MTP)

The MTP head is included, in BF16, grafted from the abliterated base (not stock Qwen):

llm = LLM("TelperionAI/Huihui-Qwen3.8-27B-abliterated-NVFP4", tensor_parallel_size=2,
          speculative_config={"method": "mtp", "num_speculative_tokens": 2})

Qwen3_5ForConditionalGeneration does not carry mtp.* in its state dict, so
llm-compressor silently drops it even though config.json declares
mtp_num_hidden_layers: 1. It is excluded from quantization via re:.*mtp.*.
Acceptance rate has not been measured; the head is verified to load and generate.

Limitations

  • Single evaluation corpus, and no downstream task benchmarks.
  • Abliterated base. This model has had its refusal directions removed upstream; that
    behaviour is inherited here and is not something quantization changes.
  • The abliterated calibration set is ~18% smaller than the stock one (the same length
    filter kept fewer generations), so it is not perfectly matched to the stock-model builds.
  • Vision tower untouched (BF16); evaluated as a text model.
  • MTP acceptance rate unmeasured.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-21Upload folder using huggingface_hub3079cd64.8 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration