license: apache-2.0
library_name: llama.cpp
tags:
- gguf
- llama.cpp
- qwen
- qwen3.8
- text-generation
- conversational
pipeline_tag: text-generation
Qwen3.8 27B Uncensored — GGUF
A collection of GGUF quantizations of Qwen3.8 27B Uncensored, prepared for use with llama.cpp and other GGUF-compatible inference engines.
The goal of this repository is to make multiple quantization levels publicly available so users can compare quality, memory usage, inference speed, and compression across different quantization formats.
Status: Active quantization and benchmarking project. Additional quantizations and benchmark results will be added over time.
Model Information
- Architecture: Qwen3.8 27B
- Parameter count: ~27B
- Format: GGUF
- Original precision: F16
- Quantization: Multiple
llama.cppquantization formats - Primary inference engine: llama.cpp
- Repository type: Community GGUF quantizations
Available Models
| File | Quantization | Approx. Size | Intended Use |
|---|---|---|---|
model-f16.gguf |
F16 | ~52 GB | Maximum-quality reference |
model-q8_0.gguf |
Q8_0 | ~26 GB | Very high quality |
model-q6_k.gguf |
Q6_K | ~21–22 GB | Excellent quality |
model-q5_k_m.gguf |
Q5_K_M | ~18–19 GB | High quality / good compression |
model-q5_k_s.gguf |
Q5_K_S | ~18 GB | High quality |
model-q4_k_m.gguf |
Q4_K_M | ~16 GB | Recommended general-purpose quant |
model-q4_k_s.gguf |
Q4_K_S | ~15 GB | Smaller Q4 option |
model-q3_k_l.gguf |
Q3_K_L | ~14 GB | More aggressive compression |
model-q3_k_m.gguf |
Q3_K_M | ~13 GB | Low-memory option |
model-q3_k_s.gguf |
Q3_K_S | ~12 GB | Aggressive compression |
model-q2_k.gguf |
Q2_K | ~11 GB | Extreme compression |
Additional IQ/imatrix-based variants may be added as the project progresses.
Note: File sizes are approximate. Actual sizes depend on tensor composition and the exact model/quantizer version.
Recommended Quantization
For most users, start with:
Q4_K_M
model-q4km.gguf
This provides a strong balance between:
- Model quality
- Memory requirements
- Inference speed
- Disk size
The Q4_K_M version produced from the F16 source is approximately 16 GiB.
If your hardware has more memory available, consider Q5_K_M, Q6_K, or Q8_0.
If memory is limited, Q3/Q2 variants may be useful, although quality degradation becomes increasingly noticeable at lower bitrates.
Quantization Method
The quantizations are generated from the original F16 GGUF rather than repeatedly quantizing an already-quantized model.
Conceptually:
F16 GGUF
│
├── Q8_0
├── Q6_K
├── Q5_K_M
├── Q5_K_S
├── Q4_K_M
├── Q4_K_S
├── Q3_K_L
├── Q3_K_M
├── Q3_K_S
└── Q2_K
This repository uses llama.cpp for GGUF quantization.
Example: Running with llama.cpp
Install/build llama.cpp:
git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp
cmake -B build
cmake --build build --config Release -j
Then run a model:
./build/bin/llama-cli \
-m model-q4km.gguf \
-p "Hello! Explain what you can do."
For a server:
./build/bin/llama-server \
-m model-q4km.gguf \
--host 0.0.0.0 \
--port 8080
The exact command-line options may vary with your installed version of llama.cpp.
Choosing a Quant
A simple rule of thumb:
More VRAM/RAM
│
▼
F16
Q8_0
Q6_K
Q5_K_M
Q4_K_M ← Recommended starting point
Q3_K_M
Q2_K
│
▼
Less VRAM/RAM
The best quantization depends on your hardware and workload.
For a fair comparison, users are encouraged to test the same:
- Prompt
- Context length
- GPU offload configuration
- Number of GPU layers
- Batch size
- Sampling parameters
when comparing variants.
Benchmarking
Benchmarking is an ongoing part of this project.
Future benchmark results will compare:
- Model size
- RAM usage
- VRAM usage
- Prompt processing speed
- Token generation speed
- Context length
- Output quality
- Quantization degradation
Benchmark results will be added here as they are measured.
Important
No quality or performance ranking should be considered authoritative until it has been experimentally measured.
Quantization Details
The initial Q4_K_M quantization was generated from the F16 source using llama.cpp.
Source model size:
~52,115 MiB
Q4_K_M output:
~16,021 MiB
The reported quantization density was approximately:
4.92 BPW
This corresponds to roughly a 69% reduction in model storage compared with the original F16 GGUF.
Intended Use
These GGUF files are intended for:
- Local LLM inference
- llama.cpp
- GGUF-compatible applications
- Hardware/inference benchmarking
- Quantization experiments
- Comparing memory/performance tradeoffs
Limitations
Quantization is inherently lossy.
Lower-bit quantizations generally reduce memory requirements but can introduce greater quality degradation.
Results can also vary depending on:
- Prompt type
- Context length
- Sampling configuration
- Inference engine
- Hardware
- Quantization method
- Calibration/imatrix data
Therefore, users should evaluate the quantization that best fits their particular workload.
Disclaimer
This repository contains community-generated GGUF quantizations.
The original model's license, terms of use, safety policies, and attribution requirements remain applicable. Users should review the original model documentation and license before using or redistributing these files.
This repository does not claim ownership of the underlying model.
Credits
Original Model
Qwen3.8 27B Uncensored
Please refer to the original model release for:
- Model architecture
- Training information
- License
- Intended use
- Safety considerations
- Original model documentation
Quantization
Quantizations generated using:
llama.cpp
https://github.com/ggml-org/llama.cpp
Contributing
If you test one of the quantizations, useful feedback includes:
- Quantization used
- Hardware
- RAM/VRAM
- Context length
- Tokens/second
- Prompt processing speed
- Backend/inference engine
- Observed quality differences
Pull requests and benchmark contributions are welcome.
Changelog
2026-08-17
- Added F16 GGUF
- Added Q4_K_M GGUF
- Began generating additional quantization variants
- Started public benchmarking/quantization comparison project
More quantizations and benchmark results will be added as testing progresses.