← back to catalog · registered 2026-08-22 13:56

BennyDaBall/Qwen3.8-Uncensored-NVFP4-MTP

BennyDaBall Qwen GGUF multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/BennyDaBall%2FQwen3.8-Uncensored-NVFP4-MTP"
Response includes
  • classification m8
  • files 4
  • hub_downloads_all_time 5,476
  • author_summary 3 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
5K
2K last 30d - stable
Likes
12
Model age
7w ago
created 2026-08-21
Downloads over time
Now5.8K→from69↑8,281%
02.1K4.2K6.4K69 on Aug 195.8K on Oct 11AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh
Tags
gguf nvfp4 qwen3.8 qwen3.5 blackwell mtp speculative-decoding uncensored abliterated vision multimodal llama.cpp

Related

Total size
18.3 GB
Files
4
Quantizations
2
Registered
2026-08-22 13:56
Last updated on HF
2026-08-26 18:18

Files by quantization

BF16 1 file 885 MB
mmproj-BF16.gguf 885 MB 5ac423f8 download
Auxiliary files 3 files 18.3 GB
Qwen3.8-27B-Uncensored-NVFP4-MTP.gguf 18.3 GB db17acbc download
README.md 6.34 KB 8331e239 download
.gitattributes 1.61 KB d40df5bd download

README current version from Hugging Face


license: apache-2.0
language:

  • en
  • zh
    base_model:
  • JonathanColetti/Qwen3.8-27B-Uncensored
    base_model_relation: quantized
    library_name: gguf
    pipeline_tag: text-generation
    tags:
  • gguf
  • nvfp4
  • qwen3.8
  • qwen3.5
  • blackwell
  • mtp
  • speculative-decoding
  • uncensored
  • abliterated
  • vision
  • multimodal
  • llama.cpp
  • lm-studio
  • conversational

Qwen3.8-27B-Uncensored NVFP4 (MTP)

Follow me on X @BennyDaBall_OG !

A native NVFP4 GGUF of JonathanColetti/Qwen3.8-27B-Uncensored, the Heretic-abliterated build of Qwen3.8-27B. The whole transformer backbone is quantized to 4-bit NVFP4 for Blackwell, and the MTP (multi-token prediction) speculative head is kept intact for fast decoding. No retraining, no distillation, just a clean quant.


What is this?

  • 27B, uncensored, NVFP4. Every attention, Gated DeltaNet, and MLP weight across all 64 layers is quantized to native NVFP4 (GGML type 40). The lm_head, token embeddings, and the MTP draft head stay in BF16.
  • MTP head retained. The model ships as 65 blocks (blk.64 is the MTP draft head), so llama.cpp can self-speculate with --spec-type draft-mtp for a large decode speedup, no separate draft model needed.
  • Vision included. The repo ships the BF16 vision projector (mmproj-BF16.gguf), so the model sees images and video frames when loaded with --mmproj. The abliteration only modified text-model tensors, so the vision tower is stock Qwen3.8 quality.
  • 262,144 native context.

Use this file when you want a fast, uncensored 27B on a Blackwell GPU (RTX 5090, 5080, RTX PRO) with native FP4 density and built-in speculative decoding.


The files

File Size Backbone lm_head token_embd MTP head
Qwen3.8-27B-Uncensored-NVFP4-MTP.gguf 18.34 GiB NVFP4 BF16 BF16 BF16
mmproj-BF16.gguf (vision projector) 0.86 GiB BF16 - - -
sha256 (model):  db17acbca53da5a7b0e861175b198cc6fd467865e99d1ad53d9dca584257a1a1
sha256 (mmproj): 5ac423f8a29059dc24e51bc6a43e9380dcd57a9347f28b62591e0b3f60b7081c

The corrected chat template is embedded in the model file. Text-only use works without the mmproj; add it when you want image input (~1 GiB extra VRAM).


Requirements

  • A Blackwell (sm_120) GPU for the native FP4 path.
  • A recent llama.cpp with NVFP4 CUDA kernels and the qwen35 architecture, or LM Studio 2.29.1+ (runtime llama.cpp-nvidia-cuda12 2.29.1 loads it and runs coherently).
  • For the MTP speedup, a build with the draft-mtp speculative path.

Quick Start

llama.cpp / llama-server

llama-server \
  --model Qwen3.8-27B-Uncensored-NVFP4-MTP.gguf \
  --mmproj mmproj-BF16.gguf \
  --ctx-size 229376 \
  --flash-attn on \
  -ctk q8_0 -ctv q8_0 \
  --spec-type draft-mtp \
  --spec-draft-n-max 3 \
  --spec-draft-p-split 0.2 \
  --temp 1.0 --top-p 0.95 --top-k 20

MTP acceptance is hardware and prompt dependent. Sweep --spec-draft-n-max from 1 to 6 and keep whatever is fastest on your box. Drop the --mmproj line (or pass --no-mmproj) for a text-only server with a little more VRAM headroom.

LM Studio

Load the file and chat. Two things to know:

  • Thinking is on by default. Give it a generous max-tokens budget or the whole budget can be spent inside the hidden reasoning block and the visible answer comes back empty. Add /no_think to a message for direct replies.
  • MTP speculation works here too, on by default. LM Studio (runtime 2.29.1+) detects the embedded MTP head and self-speculates automatically. Measured on an RTX 5090 (greedy, structured output): speculation forced off 76 tok/s, default load 95 tok/s, tuned 103 tok/s. Tune it at load time with lms load ... --speculative-draft-mtp --speculative-draft-max-tokens 3 (or the speculative decoding settings in the UI); sweep max-tokens for your hardware.

Performance (RTX 5090, single-stream)

Measured with the llama.cpp draft-mtp path, q8_0 KV cache, flash attention on. Single-run figures, not a formal benchmark.

Context VRAM Notes
131,072 ~27.5 GiB most headroom
229,376 ~29.2 GiB recommended full-speed daily
262,144 ~30.7 GiB full native context, lean desktop only

MTP self-speculation reaches high draft acceptance on structured output (counting, code, repetitive passages) and lifts decode well above the no-speculation baseline. Acceptance and speedup drop on high-entropy creative prose, as expected.


What is NVFP4?

NVFP4 is NVIDIA's 4-bit floating-point weight format for Blackwell tensor cores: 16-element blocks, each with an FP8 (E4M3) block scale plus a global tensor scale. It keeps more of the weight distribution than integer 4-bit and runs on the FP4 tensor cores, so the whole backbone stays dense at 4-bit on a Blackwell card.


How it was made

  1. Started from the JonathanColetti/Qwen3.8-27B-Uncensored BF16 checkpoint.
  2. Converted to a BF16 GGUF parent (MTP head preserved as blk.64).
  3. Quantized the transformer backbone to NVFP4 with an importance matrix, keeping lm_head, embeddings, and the MTP head in BF16.
  4. The vision projector is the base model's own BF16 projector (the abliteration never touched the vision tower), verified working against this quant (image in, correct description out).

No weights were trained or fine-tuned. The abliteration and all behavior come from the base model; quantization is a transformation only.


Notes

  • Uncensored / abliterated. Refusal behavior is inherited from the base model (Heretic abliteration), not from this quant. You are responsible for how you use it.
  • Blackwell only for the native FP4 path.

Acknowledgements

Apache-2.0, same as every upstream artifact. "Qwen" is a trademark of Alibaba, used only to identify the upstream model; this repo is not affiliated with or endorsed by Alibaba.

Quantized locally by BennyDaBall.

Follow me on X @BennyDaBall_OG !

README history 4 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-26README: measured MTP numbers for both runtimes, LM Studio tuning flags, think...abb5b8d7.1 KB
    Loading...
  2. 2026-08-21Upload README.md with huggingface_hubd8ad61f6.3 KB
    Loading...
  3. 2026-08-21Upload README.md with huggingface_hubc0b672a6.1 KB
    Loading...
  4. 2026-08-21Upload README.md with huggingface_hubee48eaa5.3 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration