license: apache-2.0
base_model: hotdogs/Qwen3.8-27B-abliterated
base_model_relation: quantized
library_name: gguf
pipeline_tag: text-generation
tags:
- gguf
- qwen3.8
- abliterated
- uncensored
- imatrix
- llama.cpp
- 16gb-vram
Qwen3.8-27B abliterated: Q3_K_XL for 16 GB GPUs
What this is
This is Qwen3.8-27B with its refusal behavior removed (abliterated), compressed to fit entirely on a 16 GB graphics card
with room left for a useful context window, the vision projector, and the Windows desktop.
It combines three pieces of other people's work:
- The model: hotdogs/Qwen3.8-27B-abliterated, Qwen's 27B dense
model with the "refusal direction" projected out of its weights. Abliteration is a weight edit, not a fine-tune. The model's
knowledge is largely untouched, and hotdogs reports MMLU going from 0.8388 to 0.8342 on the full-precision model. - The compression recipe: Unsloth's UD-Q3_K_XL layout. Unsloth's
"Dynamic" quants are not a single format; each of the model's ~500 weight tensors gets its own precision, from 2-bit to 8-bit,
depending on how sensitive it is. This file copies Unsloth's per-tensor choices exactly and uses Unsloth's own importance matrix,
but applies them to the abliterated weights instead of stock Qwen. - The vision encoder (optional): Blackfrost-AI's Q8_0 projector, which has been tested with this model and works.
Unsloth publishes Q3_K_XL only for stock Qwen3.8, and most abliterated GGUFs use plain llama.cpp quant mixes. This
file was made to get both: the abliterated model with Unsloth's layout.
Why 16 GB
The file is 13.25 GB, sized so that the model weights, a 32K context and the vision projector all stay in VRAM on a
16 GB card, with about 1.5–2 GB left for the Windows display. Nothing spills into shared system memory, which would
slow generation sharply.
The VRAM budget at the settings below, tested on an RTX 5070 Ti (16 GB):
| Component | Approx. size |
|---|---|
| Model weights | 13.25 GB |
| KV cache, 32K context at Q4_0 | ~0.6 GB |
| Vision projector (Q8_0, optional) | ~0.63 GB |
| Linear-attention state, compute buffers | ~0.9 GB |
| Display / desktop | remainder |
Qwen3.8-27B is a hybrid model: only 16 of its 64 layers use full attention, so its KV cache is about a quarter the size
of a conventional 27B model's. That is what makes long context practical on 16 GB.
Files
| File | Size | Notes |
|---|---|---|
Qwen3.8-27B-abliterated-Q3_K_XL.gguf |
13.25 GB | Main model, text only. Includes the MTP layer. |
qwen38-hotdogs-tensor-types.txt |
4 KB | The exact per-tensor recipe used, for reproducing or tweaking. |
For vision, put mmproj-Qwen3.8-27B-ABLITERATED-Q8_0.gguf
(629 MB, Blackfrost-AI) in the same folder as the model. LM Studio will pick it up automatically. The settings below were
tested with this projector loaded, and image input works.
hf download Blackfrost-AI/Qwen3.8-27B-ABLITERATED-GGUF mmproj-Qwen3.8-27B-ABLITERATED-Q8_0.gguf --local-dir <your model folder>
Settings (LM Studio / Bionic)
These are the exact settings this model was run with on a 16 GB RTX 5070 Ti, with the Q8_0 vision projector loaded.
Context and performance
| Setting | Value |
|---|---|
| Context Length | 32000 |
| GPU Offload | 65 (all layers) |
| CPU Thread Pool Size | 12 |
| Evaluation Batch Size | 2048 |
| Physical Batch Size | 512 |
| Max Concurrent Predictions | 1 |
| Flash Attention | On |
| Automatic Optimize Based on Hardware | Off |
Memory
| Setting | Value |
|---|---|
| Offload KV Cache to GPU Memory | On |
| Unified KV Cache | On |
| Context Checkpoints | 32 |
| K Cache Quantization Type | Q4_0 |
| V Cache Quantization Type | Q4_0 |
| Keep Model in Memory | Off |
| Try mmap() | Off |
Speculative decoding
| Setting | Value |
|---|---|
| Speculative Decoding | Off |
Reasoning and generation
| Setting | Value |
|---|---|
| Enable Thinking | On |
| Reasoning Budget | 8192 |
| Reasoning Budget Message | "I have to answer now." |
| Limit Response Length | Off |
| Context Overflow | Truncate Middle |
Sampling (as tested)
| Setting | Value |
|---|---|
| Temperature | 1 |
| Top K | 40 |
| Top P | 0.95 |
| Min P | 0.05 |
| Repeat Penalty | 1.1 |
| Presence Penalty | Off |
| System Prompt | (empty) |
These are LM Studio's defaults, apart from the temperature. They work, but they are not what Qwen recommends for Qwen3.8.
For Qwen's own values:
| Setting | Thinking mode | Non-thinking mode |
|---|---|---|
| Temperature | 1.0 | 0.7 |
| Top K | 20 | 20 |
| Top P | 0.95 | 0.8 |
| Min P | 0 (off) | 0 (off) |
| Repeat Penalty | 1.0 (off) | 1.0 (off) |
| Presence Penalty | 0 (off) | 1.5 |
The repeat penalty matters most here: 1.1 discourages repeating tokens the model legitimately needs, such as variable names in code.
Two settings that matter on 16 GB
- Speculative Decoding (MTP) must be off. The file contains Qwen3.8's MTP layer, and LM Studio may turn MTP on by
default. It needs 1–2 GB of extra VRAM, enough to push this setup into shared memory and slow it down sharply.
Turn it on only if you have the headroom, and use 2 draft tokens rather than 3. - Max Concurrent Predictions should be 1. Qwen3.8's linear-attention layers keep a fixed state per concurrent sequence,
roughly 150 MB each, so 4 slots cost about 0.5 GB more than 1 for no benefit in a single chat.
Tips
- Long outputs: at 32K context, a single very long reply (thousands of words), plus the reasoning budget and an agent's
system prompt, can hit the limit and get cut off. In an agent like Bionic, this shows up as an "Unterminated string in JSON"
error. Ask for long documents in parts (outline first, then one section at a time). - More context: without the vision projector, there is room for a longer context. This was run at 32K, so if you go
higher, watch Task Manager's "Shared GPU memory" and step back if it starts to fill. - Other runtimes: any recent llama.cpp-based runtime works. For llama-server, the equivalent is
-c 32000 -ngl 99 -fa on -ctk q4_0 -ctv q4_0 -np 1 --mmproj mmproj-Qwen3.8-27B-ABLITERATED-Q8_0.gguf --temp 1.0 --top-k 20 --top-p 0.95 --min-p 0
(Qwen's recommended sampling).
How it was made
- Source:
Qwen3.8-27B-abliterated-mtp-f16.gguffrom
hotdogs/Qwen3.8-27B-abliterated-MTP-GGUF, full precision. - Layout: the per-tensor quant types were read from Unsloth's
Qwen3.8-27B-UD-Q3_K_XL.ggufwithgguf_dump
and copied 1:1 usingllama-quantize --tensor-type-file. - Importance matrix: Unsloth's
imatrix_unsloth.gguf. - Extra bits: about 100 MB was added on top of Unsloth's 13.15 GB layout to fill a 13.25 GB budget. The upgrades follow
llama.cpp's own Q3_K_M priorities, one quant step each:ffn_downin layers 0–7attn_vin the full-attention layers- one
attn_output
llama-quantize --imatrix imatrix_unsloth.gguf --tensor-type-file qwen38-hotdogs-tensor-types.txt \
--token-embedding-type q3_k --output-tensor-type q5_k \
Qwen3.8-27B-abliterated-mtp-f16.gguf Qwen3.8-27B-abliterated-Q3_K_XL.gguf Q3_K_M
No benchmarks have been run on this file. This is not an official Unsloth quant. Unsloth's published quality numbers
apply to their own file, not to this one. The extra-bit choices are a heuristic, not measured.
Credits
- Qwen: base model (Apache-2.0)
- hotdogs: abliteration and the f16 source (Apache-2.0)
- Unsloth: UD-Q3_K_XL tensor layout and importance matrix (Apache-2.0)
- Blackfrost-AI: Q8_0 vision projector (Apache-2.0)
Note
This model has had refusal behavior removed and will comply with requests the original model would decline.
You are responsible for how you use it.