language:
- en
- zh
license: apache-2.0
library_name: llama.cpp
base_model: - JonathanColetti/Qwen3.8-27B-Uncensored
- Qwen/Qwen3.8-27B
base_model_relation: quantized
pipeline_tag: text-generation
tags: - unsloth
- dynamic-3.0
- ud3
- gguf
- qwen
- uncensored
- mtp
- speculative-decoding
- 16gb-vram
- llama.cpp
- text-generation
- conversational
Qwen3.8-27B-Uncensored (Unsloth Dynamic 3.0 UD3-GGUF)
An advanced, hybrid mixed-precision Unsloth Dynamic 3.0 (UD3) quantization of Qwen3.8-27B-Uncensored featuring a verified, pinned Q8_0 Multi-Token Prediction (MTP) draft head (Layer 64).
Engineered specifically to fit a complete 27-Billion parameter uncensored reasoning model plus a massive 128K context window directly inside consumer 16GB VRAM GPUs (such as the NVIDIA RTX 4070 Ti SUPER, RTX 4080, RTX 3090, and RTX 4090) as well as Apple Silicon Macs and Linux workstations.
Demystifying the "2-Bit" Label: The Mixed-Precision Architecture
Why is this labeled "2-bit" on Hugging Face?
Hugging Face automatically buckets this repository under its "2-bit" filter because the base ftype identifier isQ2_K(~2.7 bits per weight on non-critical MLP blocks).However, this is NOT a degraded uniform 2-bit model.
Standard uniform 2-bit quants (IQ2_Mat 10.6 GB) compress all 64 layers equally, resulting in severe degradation of reasoning and vocabulary.This build (
UD-Q2_K_XLat 12.1 GB on disk) injects 1.5 GB of extra high-precision tensor data into the most critical neural paths, delivering near-4-bit reasoning accuracy with the memory footprint of a 2-bit model.
Precision Allocation Breakdown
| # | Precision | Target Layers & Tensors | Purpose |
|---|---|---|---|
| 1 | Q8_0 / F32 | Input embeddings (token_embd), Output logits, all 17 Layer-64 NextN draft heads |
Zero vocabulary loss and intact MTP speculative speed |
| 2 | Q4_K | Core attention projections (Layers 18 to 28: attn_k, attn_v, attn_o) |
Full 4-bit reasoning fidelity on deep logic layers |
| 3 | Q3_K | Intermediate self-attention scoring matrices | Optimal balance between memory and attention scoring |
| 4 | IQ2_M | Bulk feed-forward network (FFN/MLP) blocks | Maximum compression on noise-resilient weights |
Key Advantages Over Standard Quantizations
| Feature | Standard IQ2_M |
Standard Q4_K_M |
This Build: UD3-Q2_K_XL |
|---|---|---|---|
| File Size on Disk | 10.6 GB | 16.8 GB | 12.1 GB (11.23 GB raw) |
| Quantization Method | Uniform 2-bit | Uniform 4-bit | Dynamic Layer-Importance Mix |
| Token Vocabulary | Degraded | Baseline | Max Precision (Q8_0 / F32) |
| MTP Draft Head | Missing / Fused | Missing / Fused | Pinned Q8_0 (Layer 64 Verified) |
| 16GB VRAM + 128K Context | Fits (quality loss) | Out of Memory | Fits Comfortably (~14.2 GB) |
| Reasoning Quality | Degraded (+0.70 PPL) | Baseline (+0.02 PPL) | Near-4-Bit Quality (~0.12 PPL) |
🚀 Beginner-Friendly Setup Guides (Pick Your Operating System)
Click on your operating system below for an exact, step-by-step walkthrough.
💾 How to Use a Different Drive (D:, E:, G:, etc.) & Check Free Space (Click to open)
If your main C: drive is running low on storage, you can easily save and run the 12.1 GB model on a secondary hard drive or NVMe SSD (such as D:, E:, or G:).
Method 1: Check Available Drive Letters via Graphical Interface (GUI)
- On your keyboard, press the Windows Key + E to open File Explorer.
- In the left navigation pane, click on This PC.
- Under the Devices and drives section, look at all your connected drives.
- Note which drive letter has at least 20 GB of free space (e.g.
D:,E:, orG:).
Method 2: Check Available Drive Letters via PowerShell
- In PowerShell, type this command and press Enter:
Get-Volume | Select-Object DriveLetter, FileSystemLabel, @{Name="FreeSpaceGB";Expression={[math]::round($_.SizeRemaining/1GB,2)}}
- You will see a clean list of every drive letter and exactly how many gigabytes of free space remain on each.
How to Apply Your Chosen Drive Letter
Whenever you see C:\QwenModel in the instructions below, simply replace the letter C with your target drive letter.
For example, if using drive D::
- Switch to drive D in PowerShell: Type
D:and press Enter. - Create your folder:
mkdir QwenModel; cd QwenModel. - In
docker-compose.yml, change"C:/QwenModel:/models:ro"to"D:/QwenModel:/models:ro". - In direct launch commands, change
"C:\QwenModel\..."to"D:\QwenModel\...".
🪟 Windows Setup Guide (Click to open)
Follow these steps on Windows 10 or Windows 11:
Step 1: Open PowerShell
- On your keyboard, press the Windows Key + R at the same time to open the Run window.
- In the box, type:
powershelland press Enter (or click OK). - A blue or black terminal window will appear.
Step 2: Create a Dedicated Folder
(Note: If you want to use a different drive such as D: or G:, see the Drive Selection guide above).
Copy and paste this command into your PowerShell window, then press Enter:
mkdir C:\QwenModel; cd C:\QwenModel
Step 3: Download the Model Files
Copy and paste these commands into PowerShell to download the model directly:
# Install the official downloader (if not already installed)
pip install -U huggingface_hub
# Download the model and vision files into your folder
huggingface-cli download DavidrPatton/Qwen3.8-27B-Uncensored-UD3-GGUF Qwen3.8-27B-Uncensored-UD-Q2_K_XL.gguf --local-dir C:\QwenModel
huggingface-cli download DavidrPatton/Qwen3.8-27B-Uncensored-UD3-GGUF mmproj-BF16.gguf --local-dir C:\QwenModel
(Alternative: You can also click the Files and versions tab at the top of this Hugging Face page and click the download button next to each file, then move them into your folder).
Step 4: Run the Server (Choose Option A or Option B)
Option A: 1-Click Launch with Docker (Recommended)
If you have Docker Desktop installed:
- In
C:\QwenModel, create a text file nameddocker-compose.yml. - Paste the following text into the file and save it (update
C:/QwenModelif using another drive):
services:
qwen38-server:
image: ghcr.io/ggml-org/llama.cpp:server-cuda
container_name: qwen38-server
restart: unless-stopped
ports:
- "6969:6969"
volumes:
- "C:/QwenModel:/models:ro"
command: >
--model "/models/Qwen3.8-27B-Uncensored-UD-Q2_K_XL.gguf"
--mmproj "/models/mmproj-BF16.gguf"
--n-gpu-layers 99
--ctx-size 131072
--cache-type-k q8_0
--cache-type-v q4_0
--temp 1.0
--top-p 0.95
--top-k 20
--min-p 0.05
--dry-multiplier 0.6
--dry-base 1.75
--dry-allowed-length 2
--xtc-probability 0.1
--flash-attn on
--spec-type draft-mtp
--spec-draft-n-max 2
-b 4096
-ub 1024
--cont-batching
--parallel 1
--reasoning-preserve
--reasoning-budget 1024
--metrics
--host 0.0.0.0
--port 6969
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: all
capabilities: [gpu]
- In your PowerShell window inside
C:\QwenModel, run:
docker compose up -d
Option B: Run Directly on Windows (No Docker)
- Download the pre-built Windows CUDA ZIP from llama.cpp Releases (look for
llama-b*-bin-win-cuda-cu12.4-x64.zip). - Extract the files into
C:\QwenModel. - In your PowerShell window, run:
.\llama-server.exe `
-m "C:\QwenModel\Qwen3.8-27B-Uncensored-UD-Q2_K_XL.gguf" `
--mmproj "C:\QwenModel\mmproj-BF16.gguf" `
-ngl 99 `
-c 131072 `
--cache-type-k q8_0 `
--cache-type-v q4_0 `
--temp 1.0 `
--top-p 0.95 `
--top-k 20 `
--min-p 0.05 `
--dry-multiplier 0.6 `
--dry-base 1.75 `
--dry-allowed-length 2 `
--xtc-probability 0.1 `
--flash-attn on `
--spec-type draft-mtp `
--spec-draft-n-max 2 `
-b 4096 `
-ub 1024 `
--cont-batching `
--reasoning-preserve `
--reasoning-budget 1024 `
--port 6969
Step 5: Start Chatting!
Open your web browser (Chrome, Edge, Firefox) and go to:
👉 http://localhost:6969
You will see the interactive chat window. Type a message and watch the model reason inside <think> tags and generate code!
🍎 macOS Setup Guide (Apple Silicon M1 / M2 / M3 / M4) (Click to open)
Follow these steps on Apple Silicon Macs (16GB+ Unified Memory recommended):
Step 1: Open Terminal
- On your Mac keyboard, press Command (⌘) + Space to open Spotlight.
- Type
Terminaland press Return.
Step 2: Create a Folder for Your Models
Copy and paste this into Terminal and press Return:
mkdir -p ~/QwenModel && cd ~/QwenModel
Step 3: Download the Model Files
Run this in Terminal to download the files:
curl -L -O https://huggingface.co/DavidrPatton/Qwen3.8-27B-Uncensored-UD3-GGUF/resolve/main/Qwen3.8-27B-Uncensored-UD-Q2_K_XL.gguf
curl -L -O https://huggingface.co/DavidrPatton/Qwen3.8-27B-Uncensored-UD3-GGUF/resolve/main/mmproj-BF16.gguf
Step 4: Install and Run llama.cpp with Apple Metal GPU
- Install llama.cpp using Homebrew:
brew install llama.cpp
- Start the server with full Apple Metal GPU acceleration:
llama-server \
-m ~/QwenModel/Qwen3.8-27B-Uncensored-UD-Q2_K_XL.gguf \
--mmproj ~/QwenModel/mmproj-BF16.gguf \
-ngl 99 \
-c 32768 \
--temp 1.0 \
--top-p 0.95 \
--top-k 20 \
--min-p 0.05 \
--flash-attn on \
--spec-type draft-mtp \
--spec-draft-n-max 2 \
--reasoning-preserve \
--reasoning-budget 1024 \
--port 6969
Step 5: Start Chatting!
Open Safari or Chrome and navigate to:
👉 http://localhost:6969
🐧 Linux Setup Guide (Ubuntu / Debian / Arch) (Click to open)
Follow these steps on Linux with an NVIDIA GPU:
Step 1: Open Terminal
Press Ctrl + Alt + T on your keyboard to open the terminal.
Step 2: Create Folder & Download Files
mkdir -p ~/QwenModel && cd ~/QwenModel
# Download model weights
pip install -U huggingface_hub
huggingface-cli download DavidrPatton/Qwen3.8-27B-Uncensored-UD3-GGUF Qwen3.8-27B-Uncensored-UD-Q2_K_XL.gguf --local-dir ~/QwenModel
huggingface-cli download DavidrPatton/Qwen3.8-27B-Uncensored-UD3-GGUF mmproj-BF16.gguf --local-dir ~/QwenModel
Step 3: Run with Docker Compose
- Create
docker-compose.ymlin~/QwenModel:
services:
qwen38-server:
image: ghcr.io/ggml-org/llama.cpp:server-cuda
container_name: qwen38-server
restart: unless-stopped
ports:
- "6969:6969"
volumes:
- "$HOME/QwenModel:/models:ro"
command: >
--model "/models/Qwen3.8-27B-Uncensored-UD-Q2_K_XL.gguf"
--mmproj "/models/mmproj-BF16.gguf"
--n-gpu-layers 99
--ctx-size 131072
--cache-type-k q8_0
--cache-type-v q4_0
--temp 1.0
--top-p 0.95
--top-k 20
--min-p 0.05
--dry-multiplier 0.6
--dry-base 1.75
--dry-allowed-length 2
--xtc-probability 0.1
--flash-attn on
--spec-type draft-mtp
--spec-draft-n-max 2
-b 4096
-ub 1024
--cont-batching
--parallel 1
--reasoning-preserve
--reasoning-budget 1024
--metrics
--host 0.0.0.0
--port 6969
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: all
capabilities: [gpu]
- Start the container:
docker compose up -d
Step 4: Start Chatting!
Open your browser and navigate to http://localhost:6969.
🔌 Connect to AI Coding Editors (ZCode, Cursor, Zed) (Click to open)
Because the server provides an OpenAI-compatible API endpoint at http://localhost:6969/v1, you can connect your favorite coding editors:
In ZCode
- Open ZCode Settings -> Add Custom Model.
- Provider: OpenAI-compatible.
- Base URL:
http://127.0.0.1:6969/v1 - API Key:
local(or leave blank). - Model Name:
Qwen3.8-27B-Uncensored - Context Window:
131072
In Cursor
- Open Cursor Settings -> Models.
- Under OpenAI API Key, type:
local. - Under Base URL, type:
http://localhost:6969/v1. - Add Model Name:
Qwen3.8-27B-Uncensored.
In Zed
- In Zed Settings (
settings.json), add a custom endpoint:
{
"language_models": {
"openai": {
"api_url": "http://localhost:6969/v1",
"available_models": [
{
"name": "Qwen3.8-27B-Uncensored",
"max_tokens": 131072
}
]
}
}
}
📊 Performance Benchmarks (Tested on RTX 4070 Ti SUPER 16GB)
- Generation Speed: ~58 to 64 tokens/second (with MTP speculative decoding enabled).
- Prompt Prefill Speed: 2,500+ tokens/second (
-b 4096 -ub 1024batch evaluation). - VRAM Allocation:
- Base Model Weights: 11.8 GB
- 128K Context Buffer (
q8_0Keys /q4_0Values): 2.4 GB - Total GPU VRAM Footprint: ~14.2 GB (Cleanly fits inside 16GB VRAM with zero PCIe swapping).
🙏 Credits & Heritage
- Base Uncensored Weights: Created by Jonathan Coletti (
JonathanColetti/Qwen3.8-27B-Uncensored) using Heretic orthogonal abliteration at BF16. - Base Architecture: Developed by the Qwen Team / Alibaba Cloud (
Qwen/Qwen3.8-27B). - Dynamic 3.0 Quantization: Layer importance mapping, MTP draft pinning, and local quantization build engineered by David Patton.