base_model: philbert440/Qwen3.8-27B-Uncensored-Aggressive
library_name: llama.cpp
pipeline_tag: text-generation
tags:
- gguf
- qwen
- qwen3.8
- llama.cpp
- quantized
- speculative-decoding
- mtp
Qwen3.8-27B-Uncensored-Aggressive · i1 IQ4_XS Smaller GGUF
A compact llama.cpp GGUF build of:
philbert440/Qwen3.8-27B-Uncensored-Aggressive
This repository is optimized for fitting Qwen3.8-27B into lower-VRAM systems while retaining as much quality as possible.
The main model uses a custom mixed quantization:
- base quantization:
IQ4_XS - FFN projections:
IQ3_S - selected attention tensors:
Q5_K output.weight:Q6_K
A separate Q4_0 MTP draft model is included for optional speculative decoding.
Files
| File | Purpose | Required |
|---|---|---|
Qwen3.8-27B-Uncensored-Aggressive-i1-IQ4_XS-Smaller.gguf |
Main 27B model | Yes |
mtp-Qwen3.8-27B-Uncensored-Aggressive-Q4_0.gguf |
MTP speculative-decoding draft model | Optional |
Main model
Qwen3.8-27B-Uncensored-Aggressive-i1-IQ4_XS-Smaller.gguf
Approximate characteristics:
- size: ~13 GB
- effective weight density: ~3.97 BPW
- 64 transformer blocks
- designed for
llama.cpp - MTP is stored separately rather than bundled into the target GGUF
MTP draft model
mtp-Qwen3.8-27B-Uncensored-Aggressive-Q4_0.gguf
This is the updated Q4_0 MTP sidecar.
This MTP was quantized using --pure, with both output.weight and token_embd.weight explicitly forced to Q4_0. This avoids the default promotion of large vocabulary-related tensors to higher-bit types and reduces the MTP VRAM/storage footprint.
The MTP model is optional. The main GGUF works normally without it.
Quick start
Main model only
llama-server \
-m Qwen3.8-27B-Uncensored-Aggressive-i1-IQ4_XS-Smaller.gguf \
-ngl 999 \
--jinja
Main model + MTP speculative decoding
llama-server \
-m Qwen3.8-27B-Uncensored-Aggressive-i1-IQ4_XS-Smaller.gguf \
-md mtp-Qwen3.8-27B-Uncensored-Aggressive-Q4_0.gguf \
-ngl 999 \
--spec-type draft-mtp \
--spec-draft-n-max 3 \
--jinja
Speculative-decoding performance depends on:
- llama.cpp version
- GPU / CPU
- context length
- GPU offload configuration
- split mode
- MTP acceptance rate
For memory-constrained GPUs, loading the MTP sidecar may reduce the amount of VRAM available for KV cache.
Vision support
The upstream Qwen3.8 model family is multimodal, but this repository currently contains the main GGUF and MTP sidecar only.
Image input in llama.cpp additionally requires a compatible mmproj GGUF.
In other words:
- text inference: supported by the main GGUF
- MTP speculative decoding: supported with the included MTP sidecar
- image input: requires a compatible
mmprojfile in addition to the main model
The MTP file itself is unrelated to the vision projector.
Main-model quantization
A model-specific importance matrix was used:
mradermacher/Qwen3.8-27B-Uncensored-Aggressive-i1-GGUF
Importance-matrix file:
Qwen3.8-27B-Uncensored-Aggressive.imatrix.gguf
The main model was quantized directly from the BF16 GGUF with:
llama-quantize \
--imatrix Qwen3.8-27B-Uncensored-Aggressive.imatrix.gguf \
--tensor-type ffn_down=iq3_s \
--tensor-type ffn_up=iq3_s \
--tensor-type ffn_gate=iq3_s \
Qwen3.8-27B-Uncensored-Aggressive-BF16-noMTP.gguf \
Qwen3.8-27B-Uncensored-Aggressive-i1-IQ4_XS-Smaller.gguf \
IQ4_XS
Main-model tensor layout
Validated tensor distribution:
| Type | Count |
|---|---|
| F32 | 353 |
| IQ3_S | 192 |
| IQ4_XS | 241 |
| Q5_K | 64 |
| Q6_K | 1 |
| Total | 851 |
Additional checks:
- transformer blocks:
blk.0throughblk.63 - all 192
ffn_down,ffn_gate, andffn_upprojection tensors areIQ3_S - no BF16/F16 weight tensors remain in the final main GGUF
output.weightisQ6_K
MTP Q4_0 quantization
The MTP sidecar is generated separately from the same BF16 source model and then quantized directly from its BF16 MTP GGUF.
The current MTP build uses:
llama-quantize \
--pure \
--output-tensor-type q4_0 \
--token-embedding-type q4_0 \
Qwen3.8-27B-Uncensored-Aggressive-MTP-BF16.gguf \
mtp-Qwen3.8-27B-Uncensored-Aggressive-Q4_0.gguf \
Q4_0
Important details:
- quantization starts from the BF16 MTP source
- no
--allow-requantizeis used --pureprevents the normal mixed-type promotion strategyoutput.weightis explicitly forced toQ4_0- token embeddings are explicitly forced to
Q4_0 - the goal is to reduce MTP storage and VRAM usage while preserving speculative-decoding usefulness
Because MTP draft tokens are verified by the main model, MTP quantization primarily affects draft acceptance rate and performance, not the validity of unverified draft tokens.
Quality evaluation
The main quantized model was compared directly against the BF16 source using the same llama.cpp runtime and decoding configuration.
MTP was disabled during these quality measurements.
WikiText-2 perplexity
| Model | PPL |
|---|---|
| BF16 | 6.3640 |
| Quant | 6.5848 |
Difference:
- absolute ΔPPL:
+0.2208 - relative PPL increase: approximately
+3.47%
KLD / probability divergence
32-chunk comparison:
| Metric | Result |
|---|---|
| Mean KLD | 0.059965 |
| Median KLD | 0.024371 |
| 95th percentile KLD | 0.173735 |
| 99th percentile KLD | 0.578934 |
| Same top-token winner | ~90.68% |
Task-level regression
All task benchmarks below used deterministic decoding with reasoning disabled.
| Benchmark | BF16 | Quant | Quant - BF16 |
|---|---|---|---|
| IFEval instruction loose | 90.89% | 90.89% | +0.00 pp |
| IFEval instruction strict | 88.73% | 88.61% | -0.12 pp |
| IFEval prompt loose | 86.51% | 86.14% | -0.37 pp |
| IFEval prompt strict | 83.92% | 83.36% | -0.55 pp |
| GSM8K-CoT flexible extract | 83.32% | 90.37% | +7.05 pp |
| GSM8K-CoT strict match | 79.23% | 88.78% | +9.55 pp |
| GPQA Diamond CoT flexible extract | 13.13% | 14.65% | +1.52 pp |
Sample counts:
- IFEval: 541
- GSM8K-CoT: 1,319
- GPQA Diamond CoT: 198
Interpretation
The quantized model does not show a consistent task-level regression relative to BF16.
IFEval is effectively unchanged, while the deterministic GSM8K generation run scored higher for the quantized model.
The GSM8K increase should not be interpreted as evidence that quantization inherently improves mathematical reasoning. Quantization can alter greedy generation trajectories, which can change the final answer even when the underlying model is slightly less precise at the logit level.
SHA256 checksums
Main model
c9fdab970822cb72bc2d585b73bec24a5fd68fda1dfa976af3f6e9dd47f8bd1f Qwen3.8-27B-Uncensored-Aggressive-i1-IQ4_XS-Smaller.gguf
MTP Q4_0
e10857664938dfc50240ec16313b8766e9432c5975a9db386bb1a278720242c4 mtp-Qwen3.8-27B-Uncensored-Aggressive-Q4_0.gguf
Verify locally with:
sha256sum Qwen3.8-27B-Uncensored-Aggressive-i1-IQ4_XS-Smaller.gguf
sha256sum mtp-Qwen3.8-27B-Uncensored-Aggressive-Q4_0.gguf
Source model
Upstream model:
philbert440/Qwen3.8-27B-Uncensored-Aggressive
Please refer to the upstream repository for:
- model architecture
- original model behavior
- licensing
- tokenizer / chat-template details
- upstream documentation
Credits
- Source model:
philbert440/Qwen3.8-27B-Uncensored-Aggressive - Importance matrix:
mradermacher/Qwen3.8-27B-Uncensored-Aggressive-i1-GGUF - Quantization and runtime:
llama.cpp