← back to catalog · registered 2026-08-25 09:02

trailio/Huihui-Qwen3.8-27B-abliterated-NVFP8

trailio Qwen 24B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/trailio%2FHuihui-Qwen3.8-27B-abliterated-NVFP8"
Response includes
  • classification m1
  • files 32
  • hub_downloads_all_time 2,301
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
2K
291 last 30d - stable
Likes
0
Model age
6w ago
created 2026-08-25
Downloads over time
Now2.4K→from35↑6,851%
08911.8K2.7K35 on Aug 262.4K on Oct 11AugSepOct
Aug 26 → Oct 11 · 47 snapshots · spans 46 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
transformers safetensors qwen3_5 image-text-to-text nvfp8 fp8 quantized compressed-tensors vllm conversational base_model:huihui-ai/Huihui-Qwen3.8-27B-abliterated base_model:quantized:huihui-ai/Huihui-Qwen3.8-27B-abliterated

Related

Total size
30.5 GB
Files
32
Quantizations
1
Registered
2026-08-25 09:02
Last updated on HF
2026-08-25 08:17

Files by quantization

Auxiliary files 32 files 30.5 GB
model-00018-of-00018.safetensors 3.16 GB 1d347950 download
model-00003-of-00018.safetensors 2.37 GB 2e1bf62c download
model-00001-of-00018.safetensors 2.37 GB 5e4c4999 download
model-00004-of-00018.safetensors 1.98 GB 3ba3dcfd download
model-00016-of-00018.safetensors 1.97 GB f5c11aae download
model-00006-of-00018.safetensors 1.97 GB 866da7a4 download
model-00008-of-00018.safetensors 1.97 GB 0734c55b download
model-00010-of-00018.safetensors 1.97 GB 2c3985c8 download
model-00012-of-00018.safetensors 1.97 GB 4f20935e download
model-00014-of-00018.safetensors 1.97 GB f589006d download
model-00002-of-00018.safetensors 1.51 GB 441e9fd6 download
model-00007-of-00018.safetensors 1.04 GB f55e7afd download
model-00009-of-00018.safetensors 1.04 GB 76782e9b download
model-00011-of-00018.safetensors 1.04 GB 9b71f6b6 download
model-00013-of-00018.safetensors 1.04 GB 1e5b4aee download
model-00015-of-00018.safetensors 1.04 GB 497b23fe download
model-00017-of-00018.safetensors 1.04 GB 1945f1a5 download
model-00005-of-00018.safetensors 1.04 GB 75c742d0 download
tokenizer.json 12.2 MB 0997f410 download
vocab.json 6.41 MB 0aa0ce06 download
merges.txt 3.20 MB a494e019 download
model.safetensors.index.json 194 KB f21355ad download
tokenizer_config.json 17.5 KB 5de744b3 download
LICENSE 11.3 KB f938136e download
chat_template.jinja 8.74 KB c0c686f9 download
README.md 5.05 KB 98964fb1 download
config.json 4.42 KB 724cd52a download
.gitattributes 1.53 KB 52373fe2 download
pyproject.toml 663 B c6b51a8f download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
generation_config.json 202 B 023756cf download

README current version from Hugging Face


base_model: huihui-ai/Huihui-Qwen3.8-27B-abliterated
base_model_relation: quantized
license: apache-2.0
pipeline_tag: image-text-to-text
library_name: transformers
tags: [nvfp8, fp8, quantized, compressed-tensors, qwen3_5, vllm]

Huihui-Qwen3.8-27B-abliterated — NVFP8

Block-16, two-level-scaled FP8 E4M3. 55.6 GB → 32.8 GB (−41%).

Stock vLLM cannot load this

NVFP8 is not an NVIDIA-defined format. NVIDIA's microscaling family is NVFP4
(block-16, E4M3 scales) at 4 bits and MXFP8 (block-32, E8M0 scales) at 8 bits —
there is no vendor NVFP8. vLLM dispatches kernels by matching quantization
args: _is_nvfp4_format requires num_bits == 4, _is_mxfp8 requires
group_size == 32. This checkpoint's signature — 8-bit, tensor_group,
group_size=16 — matches nothing, and loading raises NotImplementedError.

The nvfp8/ package in this repo supplies the Triton kernel and the dispatch
hook that make it loadable.

Usage

The runtime ships in this repo, so download it as a directory rather than
letting vLLM stream the weights into the HF cache — you need a local path to
install the plugin from.

hf download trailio/Huihui-Qwen3.8-27B-abliterated-NVFP8 --local-dir ./qwen-nvfp8
cd qwen-nvfp8
pip install -e .        # registers a vllm.general_plugins entry point
from vllm import LLM
llm = LLM("./qwen-nvfp8", trust_remote_code=True, max_model_len=8192)
print(llm.generate(["The capital of France is"])[0].outputs[0].text)

pip install -e . has to run from that downloaded directory — it is what
registers the entry point. Passing the repo id straight to LLM(...) without
installing gets you the NotImplementedError described above, because the
weights land in the HF cache and the plugin is never installed.

The entry point is required rather than a convenience: vLLM forces the spawn
start method once CUDA is initialized, so calling register() by hand in the
parent process never reaches EngineCore. Installation is what makes
registration happen in every engine and worker process.

Needs ~31 GB of disk for the download and ~35 GB of VRAM at runtime (weights
plus activations), before KV cache.

Requires sm_89+ (Ada or newer). Validated on RTX PRO 6000 Blackwell
(sm_120, 96 GB) with vLLM 0.27.1.

Format

weight               [N, K]      float8_e4m3fn
weight_scale         [N, K/16]   float8_e4m3fn   one per 16 input elements
weight_global_scale  per-shard   float32

Weight-only (W8A16); activations stay bf16.

The global scale is per-shard rather than per-tensor because vLLM fuses
gate_proj+up_proj into gate_up_proj and q/k/v_proj into qkv_proj.
Those tensors were quantized independently and carry different global scales,
so the value is expanded per-output-row at load time. Collapsing it to a single
scalar mis-scales one shard by the ratio between them.

Group scales are rounded up to the next representable e4m3 value, not to
nearest. Round-to-nearest lets a scale land below its target, pushing the
group's largest element past 448 so it clips — measured at 3.4% of all
elements. Rounding up drives that to zero.

What was quantized

400 Linear layers, 24.33B of 27.78B params (87.6%). Left in BF16:

why
visual.* 27-block ViT + merger — abliteration left vision untouched
mtp.* multi-token-prediction head; accept rates are sensitive to its logits
lm_head, embed_tokens 1.27B params each; vocab projections are where 8-bit bites
linear_attn.in_proj_a/b [48, 5120] DeltaNet decay/gate — 245K params each, and error compounds through the SSM recurrence instead of averaging out
A_log, dt_bias, conv1d SSM state; config pins mamba_ssm_dtype=float32

All quantized K-dims (5120, 6144, 17408) divide by 16, so groups tile with no
padding.

Accuracy

Weight round-trip error, identical tensors:

scheme rel. error
NVFP8 (block-16) 0.0252
FP8 per-channel 0.0265
FP8 per-tensor 0.0265

On outlier-heavy weights: 0.0213 vs 0.0249 per-channel. 0.025 is the E4M3
floor — 3 mantissa bits — not a property of the scaling scheme.

Triton kernel error against an fp32 reference is ≤3.7e-3 across every layer
shape at M=1 through M=2048, roughly an order of magnitude below the
quantization error itself.

PTQ with no calibration data: weight-only NVFP8 derives its scales from the
weights, so there is no dataset dependence and nothing to overfit.

Known issues

On containers with CUDA < 12.9, FlashInfer's JIT cannot resolve sm_120 and
fails with FlashInfer requires GPUs with sm75 or higher. Set
VLLM_USE_FLASHINFER_SAMPLER=0 or use a CUDA ≥ 12.9 image. Unrelated to
quantization — it affects the sampler, not the model.

Provenance

Quantized from
huihui-ai/Huihui-Qwen3.8-27B-abliterated,
itself an abliterated derivative of Qwen/Qwen3.8-27B.

Abliteration removes refusal behavior; the base model's reduced-filtering
caveats apply here unchanged. Quantization preserves behavior — it neither adds
nor removes safety properties.

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-25Upload README.md with huggingface_hub2e6ee195 KB
    Loading...
  2. 2026-08-25Upload folder using huggingface_hub56b431f4.4 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration