← back to catalog · registered 2026-08-22 13:56

renketong/Huihui-Qwen3.8-27B-abliterated-NVFP4-GGUF

renketong Qwen 27B GGUF multimodal second-order 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/renketong%2FHuihui-Qwen3.8-27B-abliterated-NVFP4-GGUF"
Response includes
  • classification m8
  • files 4
  • hub_downloads_all_time 5,796
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
6K
2K last 30d - stable
Likes
1
Model age
7w ago
created 2026-08-18
Downloads over time
Now6.8K→from0↑0%
02.5K5K7.4K0 on Aug 186.8K on Oct 11AugSepOct
Aug 18 → Oct 11 · 49 snapshots · spans 54 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
gguf nvfp4 mtp speculative-decoding vision multimodal qwen3.8 abliterated blackwell llama.cpp text-generation base_model:huihui-ai/Huihui-Qwen3.8-27B-abliterated

Related

Total size
18.3 GB
Files
4
Quantizations
2
Registered
2026-08-22 13:56
Last updated on HF
2026-08-18 07:22

Files by quantization

mmproj 1 file 888 MB
mmproj-huihui.gguf 888 MB ebbd2520 download
Auxiliary files 3 files 18.3 GB
Qwen3.8-27B-huihui-NVFP4.gguf 18.3 GB 0e5970c0 download
README.md 4.03 KB 3fda66d8 download
.gitattributes 1.60 KB 2c8d5af0 download

README current version from Hugging Face


license: apache-2.0
base_model:

  • huihui-ai/Huihui-Qwen3.8-27B-abliterated
  • sakamakismile/Huihui-Qwen3.8-27B-abliterated-NVFP4
    tags:
  • gguf
  • nvfp4
  • mtp
  • speculative-decoding
  • vision
  • multimodal
  • qwen3.8
  • abliterated
  • blackwell
  • llama.cpp
    pipeline_tag: text-generation
    model-index:
  • name: Huihui-Qwen3.8-27B-abliterated-NVFP4-GGUF
    results: []
    quantized_by: renketong

Huihui-Qwen3.8-27B-abliterated-NVFP4-GGUF

GGUF conversion of huihui-ai / Qwen3.8-27B-abliterated (uncensored) quantized to NVFP4,
converted from the source NVFP4 safetensors by sakamakismile.

Upstream lineage:

  1. Qwen/Qwen3.8-27B (Apache-2.0) — base model
  2. huihui-ai/Huihui-Qwen3.8-27B-abliterated — abliterated (uncensored) fine-tune
  3. sakamakismile/Huihui-Qwen3.8-27B-abliterated-NVFP4 — NVFP4 quantized safetensors (compressed-tensors, nvfp4-pack-quantized)
  4. This repo — lossless repack to GGUF via convert_hf_to_gguf.py (no re-quantization, no precision loss)

Files

File Size Description
Qwen3.8-27B-huihui-NVFP4.gguf 19.65 GB Main model. NVFP4 MLP + attention, Q5_K embeddings, BF16 MTP head, 262K native context
mmproj-huihui.gguf 931 MB BF16 vision projector (mmproj) for image input
  • Architecture: qwen35 (Qwen3_5ForConditionalGeneration), 64 layers + 1 MTP layer (nextn_predict_layers=1)
  • Bits per weight: ~5.6 BPW
  • The MTP (multi-token prediction) head ships in the file — speculative decoding works out of the box in llama.cpp / LM Studio
  • Vision projector enables image input (Qwen3-VL path); load it alongside the main model.

Load in LM Studio / llama.cpp

# llama.cpp: text-only
llama-server \
  -m Qwen3.8-27B-huihui-NVFP4.gguf \
  --spec-type draft-mtp \
  --spec-draft-n-max 2 \
  -c 163840 \
  -ngl 999

# llama.cpp: with vision
llama-server \
  -m Qwen3.8-27B-huihui-NVFP4.gguf \
  --mmproj mmproj-huihui.gguf \
  --spec-type draft-mtp \
  --spec-draft-n-max 2 \
  -c 131072 \
  -ngl 999

LM Studio settings

  • Draft probability: 0 — critical. Default 0.75 rejects ~60% of correct MTP drafts and collapses speculative speed to base rate (60-70 t/s). Setting 0 accepts all drafts and unlocks ~120 t/s.
  • Min draft tokens: 0–2
  • Max draft tokens: 2–3 (acceptance collapses at ≥4)
  • KV cache quant: q8_0 recommended

Measured speed — RTX 5090 (32 GB), LM Studio, 2026-08

MTP acceptance rate drives speed; content type drives acceptance rate:

Content Draft acceptance Speed
Code / JSON 80–96% 117–129 t/s
Math reasoning 71% 114 t/s
Chinese / English prose 37–48% 88–91 t/s

Notes:

  • Short outputs (<50 tokens) measure low due to prefill amortization — irrelevant to real use.
  • The ~8% gap vs Q8-attention builds (e.g. utautako/Qwen3.8-27B-NVFP4-MTP-Q8attn-GGUF) is the price of full-NVFP4 attention; it buys smaller size and full abliteration.

How it was converted

Lossless repack from compressed-tensors NVFP4 safetensors — no dequantization → requantization round trip:

git clone --depth 1 https://github.com/ggml-org/llama.cpp
pip install -r requirements.txt  # torch + numpy + pyyaml + transformers

python convert_hf_to_gguf.py ./Huihui-Qwen3.8-27B-abliterated-NVFP4 \
  --outfile Qwen3.8-27B-huihui-NVFP4.gguf --outtype auto
python convert_hf_to_gguf.py ./Huihui-Qwen3.8-27B-abliterated-NVFP4 \
  --outfile mmproj-huihui.gguf --mmproj

Conversion is pure CPU (mmap, no VRAM), ~1 minute for 27B on a modern desktop.

Why this model

  • Uncensored (abliterated) — no safety refusals
  • NVIDIA NVFP4 — native Blackwell FP4 tensor cores, faster than GGUF Q4/Q5 k-quants at similar size
  • MTP speculative decoding — roughly doubles throughput over autoregressive baseline (65 → ~120 t/s on code)
  • Vision — image input supported via mmproj

License

Apache-2.0. Base model: Qwen/Qwen3.8-27B. Fine-tune: huihui-ai. Quantization source: sakamakismile. GGUF conversion: renketong.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-18Upload README.md with huggingface_hub59277474 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration