license: apache-2.0
base_model:
- huihui-ai/Huihui-Qwen3.8-27B-abliterated
- sakamakismile/Huihui-Qwen3.8-27B-abliterated-NVFP4
tags: - gguf
- nvfp4
- mtp
- speculative-decoding
- vision
- multimodal
- qwen3.8
- abliterated
- blackwell
- llama.cpp
pipeline_tag: text-generation
model-index: - name: Huihui-Qwen3.8-27B-abliterated-NVFP4-GGUF
results: []
quantized_by: renketong
Huihui-Qwen3.8-27B-abliterated-NVFP4-GGUF
GGUF conversion of huihui-ai / Qwen3.8-27B-abliterated (uncensored) quantized to NVFP4,
converted from the source NVFP4 safetensors by sakamakismile.
Upstream lineage:
Qwen/Qwen3.8-27B(Apache-2.0) — base modelhuihui-ai/Huihui-Qwen3.8-27B-abliterated— abliterated (uncensored) fine-tunesakamakismile/Huihui-Qwen3.8-27B-abliterated-NVFP4— NVFP4 quantized safetensors (compressed-tensors,nvfp4-pack-quantized)- This repo — lossless repack to GGUF via
convert_hf_to_gguf.py(no re-quantization, no precision loss)
Files
| File | Size | Description |
|---|---|---|
Qwen3.8-27B-huihui-NVFP4.gguf |
19.65 GB | Main model. NVFP4 MLP + attention, Q5_K embeddings, BF16 MTP head, 262K native context |
mmproj-huihui.gguf |
931 MB | BF16 vision projector (mmproj) for image input |
- Architecture:
qwen35(Qwen3_5ForConditionalGeneration), 64 layers + 1 MTP layer (nextn_predict_layers=1) - Bits per weight: ~5.6 BPW
- The MTP (multi-token prediction) head ships in the file — speculative decoding works out of the box in llama.cpp / LM Studio
- Vision projector enables image input (Qwen3-VL path); load it alongside the main model.
Load in LM Studio / llama.cpp
# llama.cpp: text-only
llama-server \
-m Qwen3.8-27B-huihui-NVFP4.gguf \
--spec-type draft-mtp \
--spec-draft-n-max 2 \
-c 163840 \
-ngl 999
# llama.cpp: with vision
llama-server \
-m Qwen3.8-27B-huihui-NVFP4.gguf \
--mmproj mmproj-huihui.gguf \
--spec-type draft-mtp \
--spec-draft-n-max 2 \
-c 131072 \
-ngl 999
LM Studio settings
- Draft probability: 0 — critical. Default 0.75 rejects ~60% of correct MTP drafts and collapses speculative speed to base rate (60-70 t/s). Setting 0 accepts all drafts and unlocks ~120 t/s.
- Min draft tokens: 0–2
- Max draft tokens: 2–3 (acceptance collapses at ≥4)
- KV cache quant: q8_0 recommended
Measured speed — RTX 5090 (32 GB), LM Studio, 2026-08
MTP acceptance rate drives speed; content type drives acceptance rate:
| Content | Draft acceptance | Speed |
|---|---|---|
| Code / JSON | 80–96% | 117–129 t/s |
| Math reasoning | 71% | 114 t/s |
| Chinese / English prose | 37–48% | 88–91 t/s |
Notes:
- Short outputs (<50 tokens) measure low due to prefill amortization — irrelevant to real use.
- The ~8% gap vs Q8-attention builds (e.g. utautako/Qwen3.8-27B-NVFP4-MTP-Q8attn-GGUF) is the price of full-NVFP4 attention; it buys smaller size and full abliteration.
How it was converted
Lossless repack from compressed-tensors NVFP4 safetensors — no dequantization → requantization round trip:
git clone --depth 1 https://github.com/ggml-org/llama.cpp
pip install -r requirements.txt # torch + numpy + pyyaml + transformers
python convert_hf_to_gguf.py ./Huihui-Qwen3.8-27B-abliterated-NVFP4 \
--outfile Qwen3.8-27B-huihui-NVFP4.gguf --outtype auto
python convert_hf_to_gguf.py ./Huihui-Qwen3.8-27B-abliterated-NVFP4 \
--outfile mmproj-huihui.gguf --mmproj
Conversion is pure CPU (mmap, no VRAM), ~1 minute for 27B on a modern desktop.
Why this model
- Uncensored (abliterated) — no safety refusals
- NVIDIA NVFP4 — native Blackwell FP4 tensor cores, faster than GGUF Q4/Q5 k-quants at similar size
- MTP speculative decoding — roughly doubles throughput over autoregressive baseline (65 → ~120 t/s on code)
- Vision — image input supported via mmproj
License
Apache-2.0. Base model: Qwen/Qwen3.8-27B. Fine-tune: huihui-ai. Quantization source: sakamakismile. GGUF conversion: renketong.