← back to catalog · registered 2026-08-22 13:56

sakamakismile/Huihui-Qwen3.8-27B-abliterated-NVFP4

sakamakismile Qwen 24B second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/sakamakismile%2FHuihui-Qwen3.8-27B-abliterated-NVFP4"
Response includes
  • classification m1
  • files 16
  • hub_downloads_all_time 90,562
  • author_summary 34 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
91K
23K last 30d - stable
Likes
65
Descendants
1
in 1 direct fork
Model age
7w ago
created 2026-08-16
Downloads over time
Now97.3K→from1.8K↑5,404%
035.6K71.2K106.9K1.8K on Aug 1897.3K on Oct 11AugSepOct
Aug 18 → Oct 11 · 49 snapshots · spans 54 days

Genealogy 1 direct fork

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Quantizations
BF16
Tags
vllm safetensors qwen3_5 nvfp4 compressed-tensors mtp speculative-decoding blackwell abliterated uncensored text-generation conversational

Related

Total size
19.1 GB
Files
16
Quantizations
2
Registered
2026-08-22 13:56
Last updated on HF
2026-08-24 11:01

Files by quantization

BF16 1 file 810 MB
model-mtp-bf16.safetensors 810 MB 90fa0e3e download
Auxiliary files 15 files 18.4 GB
model.safetensors 18.4 GB 68af39a4 download
tokenizer.json 19.1 MB 445a1c45 download
vocab.json 6.41 MB 0aa0ce06 download
merges.txt 3.20 MB a494e019 download
model.safetensors.index.json 233 KB 511701f7 download
config.json 15.3 KB db5369a1 download
chat_template.jinja 8.74 KB c0c686f9 download
README.md 4.80 KB 44663193 download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.16 KB 33818c7f download
tokenizer_config.json 1.14 KB 1d134cd2 download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
generation_config.json 225 B d423c5dd download
recipe.yaml 218 B 481bfd53 download

README current version from Hugging Face


license: apache-2.0
base_model: huihui-ai/Huihui-Qwen3.8-27B-abliterated
base_model_relation: quantized
library_name: vllm
tags:

  • nvfp4
  • compressed-tensors
  • mtp
  • speculative-decoding
  • blackwell
  • qwen3_5
  • abliterated
  • uncensored
    pipeline_tag: text-generation

Huihui-Qwen3.8-27B-abliterated-NVFP4

NVFP4 (W4A4, group 16) quantization of huihui-ai/Huihui-Qwen3.8-27B-abliterated — all credit for the model goes to @huihui-ai. This repo carries only the quantized weights.

55.6 GB → 20.0 GB. Fits on two 16 GB cards with real KV headroom.

  • MTP draft head preserved in bf16 and wired into the index — speculative decoding works out of the box.
  • bf16 kept for: lm_head, the vision tower, the DeltaNet conv1d, and the MTP head. Everything else is NVFP4 W4A4.
  • Built for SM120 / Blackwell with vanilla vLLM v0.22.0 — compressed-tensors is auto-detected, no --quantization flag.
  • Quantized GPU-resident (full bf16 dispatched across 7× RTX PRO 2000 Blackwell via device_map=auto): 119 seconds end to end.

Measured (TP=4, 32k ctx, KV fp8, RTX PRO 2000 Blackwell ×4, MTP n=3)

concurrency aggregate t/s
1 77.3
4 204.8
8 380.9

Single-stream prefill: 3,590 tok/s on an 8k prompt (prefix cache disabled, mean of 3). GPU KV cache at this config: 613,655 tokens.

Capability after abliteration + 4-bit: an 8-probe set (Japanese fluency, English code, arithmetic, instruction-following, logic puzzle, domain knowledge, defensive-security explanation, and a trading backtest that punishes look-ahead bias) scores 8/8 — identical to the unmodified Qwen3.8-27B quantized with the same recipe, at 87.2 t/s vs 85.7 t/s. No measurable degradation.

Serve

vllm serve sakamakismile/Huihui-Qwen3.8-27B-abliterated-NVFP4 \
  --trust-remote-code --tensor-parallel-size 4 \
  --max-model-len 32768 --kv-cache-dtype fp8 --reasoning-parser qwen3 \
  --speculative-config '{"method":"qwen3_5_mtp","num_speculative_tokens":3}'

On boards without P2P add NCCL_P2P_DISABLE=1 and --disable-custom-all-reduce.

⚠️ Gotchas

  • Send reasoning_effort: "medium" for long-form work. This is the one that will bite you. reasoning_effort defaults to xhigh, and at that setting the <think> phase sometimes never terminates: it grows past 19,000 characters, degenerates into repeating a single line, and the token budget is gone before any answer is emitted. On a 9-case gate (French/English long-form, temperature 0 and 0.7) this model fails 1 of 9 at xhigh and passes 9 of 9 at medium, where thinking stays near 1k characters. The unmodified Qwen3.8-27B passes the same gate at both settings, so the abliterated fine-tune is more prone to it — but medium is the right default either way.

    "chat_template_kwargs": {"reasoning_effort": "medium"}
    

    Worth knowing what these modes actually are: they are not three levels of capability, they are three system prompts. Reading the chat template — xhigh injects "think carefully through the task, validate key assumptions, consider plausible alternatives…", low injects "keep your thinking brief and focused, moving directly to the conclusion", and medium injects nothing at all — it is simply the model with no deliberation instruction. That matches what I measured: <think> runs a few hundred characters at low, ~1k at medium, 5–19k at xhigh, while answer quality was indistinguishable across all three on short verifiable tasks (96 runs, no errors at any setting). The risk at xhigh is not worse reasoning — it is deliberation outliving the token budget.

  • The 15 mtp.* modules are listed in quantization_config.ignore — do not remove them. If vLLM treats the bf16 MTP head as NVFP4 the draft breaks silently: 0% acceptance and slower than no MTP.

  • W4A16 (NVFP4A16) does not serve on this architecture in vLLM 0.22 (gptq_marlin_repack: size_n=24 not divisible by tile_n_size=64). W4A4 only.

  • Long single-file code generation drops a closing paren roughly 1–2 times in 14 regardless of sampling temperature (measured with a JS parser on the base model). Put a syntax check in the loop rather than tuning temperature.

  • Text-only serving shown above; the vision tower ships in bf16 but multimodal serving was not benchmarked here.

Recipe

llm-compressor NVFP4 (W4A4, group 16), targets: [Linear], ignore: [lm_head, re:.*visual.*, re:.*conv1d.*, re:.*mtp.*], 32 calibration samples × 8192 tok from neuralmagic/calibration. MTP tensors grafted back in bf16 after saving and appended to quantization_config.ignore.

🙏 @huihui-ai for the model, Qwen team for the base, vLLM & llm-compressor teams for the tooling.

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-24Re-quantize from upstream Latest update 3 (ablation narrowed to layers 18-51)...36276305.9 KB
    Loading...
  2. 2026-08-16Upload folder using huggingface_hub1b0f5f54.8 KB
    Loading...

Discussions 1 thread

  1. 2026-08-28Speed concernopen2 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration