← back to catalog · registered 2026-08-22 13:56

felippeburk/Huihui-Qwen3.8-27B-Abliterated-NVFP4-MTP-GGUF

felippeburk Qwen 27B GGUF second-order 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/felippeburk%2FHuihui-Qwen3.8-27B-Abliterated-NVFP4-MTP-GGUF"
Response includes
  • classification m8
  • files 4
  • hub_downloads_all_time 2,994
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
3K
385 last 30d - stable
Likes
3
Model age
7w ago
created 2026-08-21
Downloads over time
Now3.1K→from164↑1,781%
181.1K2.3K3.4K164 on Aug 193.1K on Oct 113.1K on Oct 10AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
gguf qwen3_5 nvfp4 mtp speculative-decoding blackwell abliterated uncensored llama.cpp text-generation base_model:sakamakismile/Huihui-Qwen3.8-27B-abliterated-NVFP4 base_model:quantized:sakamakismile/Huihui-Qwen3.8-27B-abliterated-NVFP4
Total size
18.3 GB
Files
4
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-21 06:20

Files by quantization

Auxiliary files 4 files 18.3 GB
huihui-qwen3.8-27b-abliterated-nvfp4-mtp.gguf 18.3 GB e5d75ac7 download
LICENSE 11.1 KB d6456956 download
README.md 4.39 KB 3de1f7ab download
.gitattributes 1.56 KB 072d3972 download

README current version from Hugging Face


license: apache-2.0
base_model: sakamakismile/Huihui-Qwen3.8-27B-abliterated-NVFP4
tags:

  • qwen3_5
  • nvfp4
  • mtp
  • speculative-decoding
  • blackwell
  • abliterated
  • uncensored
  • gguf
  • llama.cpp
  • text-generation

Huihui-Qwen3.8-27B Abliterated NVFP4-MTP GGUF

NVFP4 quantized Huihui-Qwen3.8-27B (abliterated / uncensored fine-tune of Qwen3.8-27B) with MTP (multi-token prediction) draft layers, packaged as a single GGUF for llama.cpp speculative decoding on Blackwell GPUs (RTX 5090 / 5080).

What this is

  • Quantization: NVFP4 (weights + KV cache), converted to GGUF via llama.cpp convert_hf_to_gguf.py --outtype auto
  • MTP draft layers included — enables --spec-type draft-mtp --spec-draft-n-max 2 in llama.cpp for speculative decoding
  • Multimodal: pairs with the mmproj-F16.gguf vision projector (from unsloth/Qwen3.8-27B-GGUF)
  • Single file: huihui-qwen3.8-27b-abliterated-nvfp4-mtp.gguf (~19 GB)

Attribution

Conversion methodology

Converted on an RTX 5090 (CachyOS, user-space, no Docker). The conversion script lives in the felippeburk/rtx-5090-cachyos-testing repo:

# Builds llama.cpp with NVFP4/MTP support, downloads source safetensors,
# runs the conversion, and verifies the result. Everything installs to ~/.local.
git clone https://gitlab.com/felippeburk/rtx-5090-cachyos-testing
cd rtx-5090-cachyos-testing
bash llama/setup-mtp.sh --model huihui

The conversion step is: llama.cpp/convert_hf_to_gguf.py <source-safetensors-dir> --outfile huihui-qwen3.8-27b-abliterated-nvfp4-mtp.gguf --outtype auto, run with torch + transformers in an isolated venv. Reproducible from the source safetensors at any time.

Recommended server args (llama.cpp, 192K context)

-m huihui-qwen3.8-27b-abliterated-nvfp4-mtp.gguf \
--mmproj mmproj-F16.gguf \
-fitt 8192 -c 196608 -n 65536 -fa on -ngl 99 -np 1 -t 16 -tb 16 \
-ctk q8_0 -ctv q8_0 -ctkd q4_1 -ctvd q4_1 -ctxcp 16 -cram 6144 \
--cache-idle-slots --no-warmup \
--spec-type draft-mtp --spec-draft-n-max 2 \
--temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 \
--presence-penalty 0.0 --repeat-penalty 1.0 \
--jinja --chat-template-file config/chat_template.jinja \
--reasoning on --reasoning-format deepseek --reasoning-budget 4096 \
--reasoning-budget-message "I have reached my reasoning budget. I will now provide my final answer or take the most appropriate action based on my analysis so far." \
--chat-template-kwargs '{"reasoning_effort":"medium"}'

Important usage notes

  • Use reasoning_effort medium. This abliterated fine-tune was validated at medium — at xhigh the think phase can run away past the token budget and never terminate. --chat-template-kwargs forces medium above; a patched chat template may default to xhigh.
  • Thinking mode requires a generous output budget. This model uses the Qwen3.8 thinking format, so reasoning shares the max_tokens budget with the answer. Set the client's max_tokens to at least 65536 (-n 65536 above) and give the server a --reasoning-budget (here 4096) so the thinking phase can't eat the entire output. With a too-small max_tokens, long prompts can produce reasoning but an empty final answer.
  • Pin a reasoning-effort budget message. --reasoning-budget-message tells the model to wrap up and answer once the budget is hit; without it, long tool-calling sessions can stall in thinking.
  • Use thinking-mode sampling. Temperature 1.0, presence_penalty 0.0 (the 0.7 / 1.5 values in some older configs are for instruct mode and hurt thinking-mode output).
  • reasoning_effort only accepts xhigh / medium / low — there is no high level.
  • Vision is available via --mmproj mmproj-F16.gguf (downloaded from unsloth/Qwen3.8-27B-GGUF).

Full config lives in config/models.yaml of that GitLab repo.

License

Apache-2.0 (inherited from source weights and base model)

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-21Upload README.md with huggingface_hub21551cf4.4 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration