← back to catalog · registered 2026-08-24 11:02

aldenw/Qwen3.8-27B-Uncensored-Aggressive-i1-IQ4_XS-Smaller-GGUF

aldenw Qwen 27B GGUF second-order 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/aldenw%2FQwen3.8-27B-Uncensored-Aggressive-i1-IQ4_XS-Smaller-GGUF"
Response includes
  • classification m8
  • files 4
  • hub_downloads_all_time 7,069
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
7K
904 last 30d - stable
Likes
0
Model age
6w ago
created 2026-08-24
Downloads over time
Now7.2K→from988↑633%
6753.1K5.5K7.9K988 on Aug 267.2K on Oct 11AugSepOct
Aug 26 → Oct 11 · 47 snapshots · spans 46 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

Quantizations
Q4
Tags
llama.cpp gguf qwen qwen3.8 qwen35 quantized imatrix speculative-decoding mtp uncensored abliterated text-generation

Related

Total size
14.0 GB
Files
4
Quantizations
2
Registered
2026-08-24 11:02
Last updated on HF
2026-08-25 08:38

Files by quantization

Q4 1 file 1.56 GB
mtp-Qwen3.8-27B-Uncensored-Aggressive-Q4_0.gguf 1.56 GB e1085766 download
Auxiliary files 3 files 12.4 GB
Qwen3.8-27B-Uncensored-Aggressive-i1-IQ4_XS-Smaller.gguf 12.4 GB c9fdab97 download
README.md 7.41 KB 3cc9bc42 download
.gitattributes 1.66 KB 0bca0dd6 download

README current version from Hugging Face


base_model: philbert440/Qwen3.8-27B-Uncensored-Aggressive
library_name: llama.cpp
pipeline_tag: text-generation
tags:

  • gguf
  • qwen
  • qwen3.8
  • llama.cpp
  • quantized
  • speculative-decoding
  • mtp

Qwen3.8-27B-Uncensored-Aggressive · i1 IQ4_XS Smaller GGUF

A compact llama.cpp GGUF build of:

philbert440/Qwen3.8-27B-Uncensored-Aggressive

This repository is optimized for fitting Qwen3.8-27B into lower-VRAM systems while retaining as much quality as possible.

The main model uses a custom mixed quantization:

  • base quantization: IQ4_XS
  • FFN projections: IQ3_S
  • selected attention tensors: Q5_K
  • output.weight: Q6_K

A separate Q4_0 MTP draft model is included for optional speculative decoding.


Files

File Purpose Required
Qwen3.8-27B-Uncensored-Aggressive-i1-IQ4_XS-Smaller.gguf Main 27B model Yes
mtp-Qwen3.8-27B-Uncensored-Aggressive-Q4_0.gguf MTP speculative-decoding draft model Optional

Main model

Qwen3.8-27B-Uncensored-Aggressive-i1-IQ4_XS-Smaller.gguf

Approximate characteristics:

  • size: ~13 GB
  • effective weight density: ~3.97 BPW
  • 64 transformer blocks
  • designed for llama.cpp
  • MTP is stored separately rather than bundled into the target GGUF

MTP draft model

mtp-Qwen3.8-27B-Uncensored-Aggressive-Q4_0.gguf

This is the updated Q4_0 MTP sidecar.

This MTP was quantized using --pure, with both output.weight and token_embd.weight explicitly forced to Q4_0. This avoids the default promotion of large vocabulary-related tensors to higher-bit types and reduces the MTP VRAM/storage footprint.

The MTP model is optional. The main GGUF works normally without it.


Quick start

Main model only

llama-server \
  -m Qwen3.8-27B-Uncensored-Aggressive-i1-IQ4_XS-Smaller.gguf \
  -ngl 999 \
  --jinja

Main model + MTP speculative decoding

llama-server \
  -m Qwen3.8-27B-Uncensored-Aggressive-i1-IQ4_XS-Smaller.gguf \
  -md mtp-Qwen3.8-27B-Uncensored-Aggressive-Q4_0.gguf \
  -ngl 999 \
  --spec-type draft-mtp \
  --spec-draft-n-max 3 \
  --jinja

Speculative-decoding performance depends on:

  • llama.cpp version
  • GPU / CPU
  • context length
  • GPU offload configuration
  • split mode
  • MTP acceptance rate

For memory-constrained GPUs, loading the MTP sidecar may reduce the amount of VRAM available for KV cache.


Vision support

The upstream Qwen3.8 model family is multimodal, but this repository currently contains the main GGUF and MTP sidecar only.

Image input in llama.cpp additionally requires a compatible mmproj GGUF.

In other words:

  • text inference: supported by the main GGUF
  • MTP speculative decoding: supported with the included MTP sidecar
  • image input: requires a compatible mmproj file in addition to the main model

The MTP file itself is unrelated to the vision projector.


Main-model quantization

A model-specific importance matrix was used:

mradermacher/Qwen3.8-27B-Uncensored-Aggressive-i1-GGUF

Importance-matrix file:

Qwen3.8-27B-Uncensored-Aggressive.imatrix.gguf

The main model was quantized directly from the BF16 GGUF with:

llama-quantize \
  --imatrix Qwen3.8-27B-Uncensored-Aggressive.imatrix.gguf \
  --tensor-type ffn_down=iq3_s \
  --tensor-type ffn_up=iq3_s \
  --tensor-type ffn_gate=iq3_s \
  Qwen3.8-27B-Uncensored-Aggressive-BF16-noMTP.gguf \
  Qwen3.8-27B-Uncensored-Aggressive-i1-IQ4_XS-Smaller.gguf \
  IQ4_XS

Main-model tensor layout

Validated tensor distribution:

Type Count
F32 353
IQ3_S 192
IQ4_XS 241
Q5_K 64
Q6_K 1
Total 851

Additional checks:

  • transformer blocks: blk.0 through blk.63
  • all 192 ffn_down, ffn_gate, and ffn_up projection tensors are IQ3_S
  • no BF16/F16 weight tensors remain in the final main GGUF
  • output.weight is Q6_K

MTP Q4_0 quantization

The MTP sidecar is generated separately from the same BF16 source model and then quantized directly from its BF16 MTP GGUF.

The current MTP build uses:

llama-quantize \
  --pure \
  --output-tensor-type q4_0 \
  --token-embedding-type q4_0 \
  Qwen3.8-27B-Uncensored-Aggressive-MTP-BF16.gguf \
  mtp-Qwen3.8-27B-Uncensored-Aggressive-Q4_0.gguf \
  Q4_0

Important details:

  • quantization starts from the BF16 MTP source
  • no --allow-requantize is used
  • --pure prevents the normal mixed-type promotion strategy
  • output.weight is explicitly forced to Q4_0
  • token embeddings are explicitly forced to Q4_0
  • the goal is to reduce MTP storage and VRAM usage while preserving speculative-decoding usefulness

Because MTP draft tokens are verified by the main model, MTP quantization primarily affects draft acceptance rate and performance, not the validity of unverified draft tokens.


Quality evaluation

The main quantized model was compared directly against the BF16 source using the same llama.cpp runtime and decoding configuration.

MTP was disabled during these quality measurements.

WikiText-2 perplexity

Model PPL
BF16 6.3640
Quant 6.5848

Difference:

  • absolute ΔPPL: +0.2208
  • relative PPL increase: approximately +3.47%

KLD / probability divergence

32-chunk comparison:

Metric Result
Mean KLD 0.059965
Median KLD 0.024371
95th percentile KLD 0.173735
99th percentile KLD 0.578934
Same top-token winner ~90.68%

Task-level regression

All task benchmarks below used deterministic decoding with reasoning disabled.

Benchmark BF16 Quant Quant - BF16
IFEval instruction loose 90.89% 90.89% +0.00 pp
IFEval instruction strict 88.73% 88.61% -0.12 pp
IFEval prompt loose 86.51% 86.14% -0.37 pp
IFEval prompt strict 83.92% 83.36% -0.55 pp
GSM8K-CoT flexible extract 83.32% 90.37% +7.05 pp
GSM8K-CoT strict match 79.23% 88.78% +9.55 pp
GPQA Diamond CoT flexible extract 13.13% 14.65% +1.52 pp

Sample counts:

  • IFEval: 541
  • GSM8K-CoT: 1,319
  • GPQA Diamond CoT: 198

Interpretation

The quantized model does not show a consistent task-level regression relative to BF16.

IFEval is effectively unchanged, while the deterministic GSM8K generation run scored higher for the quantized model.

The GSM8K increase should not be interpreted as evidence that quantization inherently improves mathematical reasoning. Quantization can alter greedy generation trajectories, which can change the final answer even when the underlying model is slightly less precise at the logit level.


SHA256 checksums

Main model

c9fdab970822cb72bc2d585b73bec24a5fd68fda1dfa976af3f6e9dd47f8bd1f  Qwen3.8-27B-Uncensored-Aggressive-i1-IQ4_XS-Smaller.gguf

MTP Q4_0

e10857664938dfc50240ec16313b8766e9432c5975a9db386bb1a278720242c4  mtp-Qwen3.8-27B-Uncensored-Aggressive-Q4_0.gguf

Verify locally with:

sha256sum Qwen3.8-27B-Uncensored-Aggressive-i1-IQ4_XS-Smaller.gguf
sha256sum mtp-Qwen3.8-27B-Uncensored-Aggressive-Q4_0.gguf

Source model

Upstream model:

philbert440/Qwen3.8-27B-Uncensored-Aggressive

Please refer to the upstream repository for:

  • model architecture
  • original model behavior
  • licensing
  • tokenizer / chat-template details
  • upstream documentation

Credits

  • Source model: philbert440/Qwen3.8-27B-Uncensored-Aggressive
  • Importance matrix: mradermacher/Qwen3.8-27B-Uncensored-Aggressive-i1-GGUF
  • Quantization and runtime: llama.cpp

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-25Restructure model card around quantization design3fe11518 KB
    Loading...
  2. 2026-08-25Remove not-for-all-audiences tag7a684758.7 KB
    Loading...
  3. 2026-08-25Improve model card metadata and licensing7e111fb8.7 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration