← back to catalog · registered 2026-08-22 13:56

tacodevs/Huihui-Qwen3.5-27B-Claude-4.6-Opus-abliterated-FP8

tacodevs Qwen 24B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/tacodevs%2FHuihui-Qwen3.5-27B-Claude-4.6-Opus-abliterated-FP8"
Response includes
  • classification m1
  • files 10
  • hub_downloads_all_time 467
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
467
18 last 30d - cooling
Likes
0
Model age
6mo ago
created 2026-04-08
Downloads over time
Now476→from216↑120%
203303402502216 on Apr 15476 on Oct 11476 on Oct 10AprMayJunJulAugSepOct
Apr 15 → Oct 11 · 65 snapshots · spans 179 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
transformers safetensors qwen3_5 image-text-to-text fp8 compressed-tensors quantized vision-language multimodal qwen3.5 abliterated conversational

Related

Total size
28.3 GB
Files
10
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-04-08 05:10

Files by quantization

Auxiliary files 10 files 28.3 GB
model.safetensors 28.3 GB b09a3a83 download
tokenizer.json 19.1 MB 87a7830d download
config.json 9.68 KB 1e803765 download
README.md 5.20 KB 70fe0002 download
chat_template.jinja 3.95 KB 609532bf download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.27 KB 8af8110f download
tokenizer_config.json 1.14 KB aeb7593d download
recipe.yaml 249 B 0c0661f3 download
generation_config.json 142 B 4ef941fd download

README current version from Hugging Face


base_model: huihui-ai/Huihui-Qwen3.5-27B-Claude-4.6-Opus-abliterated
license: apache-2.0
library_name: transformers
tags:

  • fp8
  • compressed-tensors
  • quantized
  • vision-language
  • multimodal
  • qwen3.5
  • abliterated
    quantized_by: tacodevs

Huihui-Qwen3.5-27B-Claude-4.6-Opus-abliterated FP8 (vision-preserving)

Calibrated FP8 quantization of huihui-ai/Huihui-Qwen3.5-27B-Claude-4.6-Opus-abliterated that preserves vision-language capabilities, unlike vLLM's dynamic FP8 which destroys them.

TL;DR

If you serve this model in vLLM with --quantization fp8 (dynamic FP8), the model completely loses vision and hallucinates random descriptions for every image (we tested: image of an anime girl in a kimono → model says "bowling ball", "Chris Hemsworth in a suit", "blank white rectangle"). This checkpoint avoids that by using calibrated FP8 with the vision tower kept in BF16.

Why dynamic FP8 destroys Qwen3.5-VL vision

Dynamic --quantization fp8 in vLLM walks every Linear layer in the model and quantizes weights to FP8 (E4M3, ~256 representable values per range) using a single tensor-wide scale per layer. For pure language layers this is fine because activations sit in similar ranges. For vision-language models it is catastrophic, and here is the exact failure mode:

  1. The vision merger is the only bridge between visual encoder and LM. Qwen3.5-VL's visual tower outputs 1152-dim image features. A single Linear merger projects those features into the 5120-dim LM embedding space. If this projection is even slightly wrong, the LM gets noise instead of meaningful visual tokens.

  2. The merger has a much wider weight distribution than LM layers. It needs to encode visual patterns at many scales, so its weights span a wider range than typical LM weights.

  3. Single-scale FP8 crushes the merger. With one tensor-wide scale, important small-magnitude weights in the merger get rounded to zero. The projector outputs become noise.

  4. The LM still receives "image embeddings" at the image token positions, but they are noise. The LM has no useful image information, falls back to text-only generation, and hallucinates plausible-sounding descriptions from the text prompt alone.

  5. The Opus-distilled fine-tune amplifies the problem. The Claude 4.6 Opus distillation used text-only training data (Claude can't share image tensors), which already weakened the vision-LM connection. Dynamic FP8 finishes the job.

End result with dynamic FP8: vision is 0 percent functional. The model generates wrong-but-plausible descriptions for every image and never sees the actual content.

What this checkpoint does differently

This checkpoint uses calibrated FP8 with explicit visual-tower exclusion:

  • Per-channel weight scales instead of one global scale per layer. Each row of each weight matrix has its own scale, computed from the actual weight distribution. This preserves precision in layers with wide dynamic range.
  • Visual tower kept in BF16 via the ignore list (re:.*visual.*, re:.*vision.*).
  • Visual merger kept in BF16 (re:.*merger.*). This is the critical bridge layer.
  • lm_head kept in BF16 (always a good idea, sensitive layer).
  • Dynamic activation quantization at inference time, computed per-token, so activations get the right scale for whatever they actually contain.

The vision encoder, the merger that bridges vision and LM, and the LM head all run in BF16 exactly as in the original model. Only the body of the language model (attention and MLP linears) is FP8.

Result

  • Vision quality: ~99 percent of the BF16 original. Confirmed working on the same images that completely break dynamic FP8.
  • LM quality: ~99 percent of the BF16 original (well within benchmark noise for FP8).
  • VRAM: ~28 GB (down from ~54 GB BF16). Half the size.
  • Speed: ~2x faster than BF16 on H100/H200/B200, identical to dynamic FP8.

Quantization details

  • Tool: llmcompressor main branch
  • Scheme: FP8_DYNAMIC (per-channel weight scales, dynamic activation scales)
  • Targets: All Linear layers
  • Excluded modules:
    • lm_head
    • re:.*visual.* (entire visual tower)
    • re:.*merger.* (vision-to-LM merger)
    • re:.*vision.* (anything else vision-related)
  • Original size: ~54 GB BF16
  • FP8 size: ~28 GB

Usage with vLLM

python -m vllm.entrypoints.openai.api_server \
    --model tacodevs/Huihui-Qwen3.5-27B-Claude-4.6-Opus-abliterated-FP8 \
    --max-model-len 16384 \
    --gpu-memory-utilization 0.50 \
    --max-num-seqs 2 \
    --trust-remote-code

IMPORTANT: Do NOT pass --quantization fp8. The model already has its quantization config baked in via compressed-tensors. vLLM will detect it from the config and use the proper FP8 path. Passing --quantization fp8 would re-quantize the already-FP8 weights and break everything.

Credits

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-04-08Add model card293b40a5.2 KB
    Loading...

Discussions 1 thread

  1. 2026-04-23Can we get the same version for Qwen 3.6?open2 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration