← back to catalog · registered 2026-08-22 13:56

end9214/qwen3.8-27B-uncensored-gguf

end9214 Qwen 27B GGUF 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/end9214%2Fqwen3.8-27B-uncensored-gguf"
Response includes
  • classification m-uncensored
  • files 13
  • hub_downloads_all_time 5,288
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
5K
1K last 30d - stable
Likes
2
Model age
7w ago
created 2026-08-17
Downloads over time
Now5.5K→from1.8K↑204%
1.6K3K4.4K5.8K1.8K on Aug 195.5K on Oct 11AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Metadata

License
apache-2.0
Quantizations
F16
Tags
llama.cpp gguf qwen qwen3.8 text-generation conversational license:apache-2.0 endpoints_compatible region:us
Total size
213 GB
Files
13
Quantizations
2
Registered
2026-08-22 13:56
Last updated on HF
2026-08-17 07:23

Files by quantization

F16 1 file 50.9 GB
model-f16.gguf 50.9 GB b6121e2d download
Auxiliary files 12 files 162 GB
model-q8_0.gguf 27.1 GB 623d243c download
model-q6_k.gguf 20.9 GB e661e973 download
model-q5_k_m.gguf 18.2 GB 4957abea download
model-q5_k_s.gguf 17.7 GB 246a6484 download
model-q4km.gguf 15.7 GB b58e9a78 download
model-q4_k_s.gguf 14.7 GB 0e574c9b download
model-q3_k_l.gguf 13.6 GB e37c3eb9 download
model-q3_k_m.gguf 12.6 GB f9391949 download
model-q3_k_s.gguf 11.4 GB 52bfede3 download
model-q2_k.gguf 10.1 GB 19797814 download
README.md 6.97 KB 879772f2 download
.gitattributes 2.05 KB 539262c7 download

README current version from Hugging Face


license: apache-2.0
library_name: llama.cpp
tags:

  • gguf
  • llama.cpp
  • qwen
  • qwen3.8
  • text-generation
  • conversational
    pipeline_tag: text-generation

Qwen3.8 27B Uncensored — GGUF

A collection of GGUF quantizations of Qwen3.8 27B Uncensored, prepared for use with llama.cpp and other GGUF-compatible inference engines.

The goal of this repository is to make multiple quantization levels publicly available so users can compare quality, memory usage, inference speed, and compression across different quantization formats.

Status: Active quantization and benchmarking project. Additional quantizations and benchmark results will be added over time.

Model Information

  • Architecture: Qwen3.8 27B
  • Parameter count: ~27B
  • Format: GGUF
  • Original precision: F16
  • Quantization: Multiple llama.cpp quantization formats
  • Primary inference engine: llama.cpp
  • Repository type: Community GGUF quantizations

Available Models

File Quantization Approx. Size Intended Use
model-f16.gguf F16 ~52 GB Maximum-quality reference
model-q8_0.gguf Q8_0 ~26 GB Very high quality
model-q6_k.gguf Q6_K ~21–22 GB Excellent quality
model-q5_k_m.gguf Q5_K_M ~18–19 GB High quality / good compression
model-q5_k_s.gguf Q5_K_S ~18 GB High quality
model-q4_k_m.gguf Q4_K_M ~16 GB Recommended general-purpose quant
model-q4_k_s.gguf Q4_K_S ~15 GB Smaller Q4 option
model-q3_k_l.gguf Q3_K_L ~14 GB More aggressive compression
model-q3_k_m.gguf Q3_K_M ~13 GB Low-memory option
model-q3_k_s.gguf Q3_K_S ~12 GB Aggressive compression
model-q2_k.gguf Q2_K ~11 GB Extreme compression

Additional IQ/imatrix-based variants may be added as the project progresses.

Note: File sizes are approximate. Actual sizes depend on tensor composition and the exact model/quantizer version.

Recommended Quantization

For most users, start with:

Q4_K_M

model-q4km.gguf

This provides a strong balance between:

  • Model quality
  • Memory requirements
  • Inference speed
  • Disk size

The Q4_K_M version produced from the F16 source is approximately 16 GiB.

If your hardware has more memory available, consider Q5_K_M, Q6_K, or Q8_0.

If memory is limited, Q3/Q2 variants may be useful, although quality degradation becomes increasingly noticeable at lower bitrates.

Quantization Method

The quantizations are generated from the original F16 GGUF rather than repeatedly quantizing an already-quantized model.

Conceptually:

F16 GGUF
   │
   ├── Q8_0
   ├── Q6_K
   ├── Q5_K_M
   ├── Q5_K_S
   ├── Q4_K_M
   ├── Q4_K_S
   ├── Q3_K_L
   ├── Q3_K_M
   ├── Q3_K_S
   └── Q2_K

This repository uses llama.cpp for GGUF quantization.

Example: Running with llama.cpp

Install/build llama.cpp:

git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp
cmake -B build
cmake --build build --config Release -j

Then run a model:

./build/bin/llama-cli \
    -m model-q4km.gguf \
    -p "Hello! Explain what you can do."

For a server:

./build/bin/llama-server \
    -m model-q4km.gguf \
    --host 0.0.0.0 \
    --port 8080

The exact command-line options may vary with your installed version of llama.cpp.

Choosing a Quant

A simple rule of thumb:

More VRAM/RAM
      │
      ▼
    F16
    Q8_0
    Q6_K
    Q5_K_M
    Q4_K_M  ← Recommended starting point
    Q3_K_M
    Q2_K
      │
      ▼
Less VRAM/RAM

The best quantization depends on your hardware and workload.

For a fair comparison, users are encouraged to test the same:

  • Prompt
  • Context length
  • GPU offload configuration
  • Number of GPU layers
  • Batch size
  • Sampling parameters

when comparing variants.

Benchmarking

Benchmarking is an ongoing part of this project.

Future benchmark results will compare:

  • Model size
  • RAM usage
  • VRAM usage
  • Prompt processing speed
  • Token generation speed
  • Context length
  • Output quality
  • Quantization degradation

Benchmark results will be added here as they are measured.

Important

No quality or performance ranking should be considered authoritative until it has been experimentally measured.

Quantization Details

The initial Q4_K_M quantization was generated from the F16 source using llama.cpp.

Source model size:

~52,115 MiB

Q4_K_M output:

~16,021 MiB

The reported quantization density was approximately:

4.92 BPW

This corresponds to roughly a 69% reduction in model storage compared with the original F16 GGUF.

Intended Use

These GGUF files are intended for:

  • Local LLM inference
  • llama.cpp
  • GGUF-compatible applications
  • Hardware/inference benchmarking
  • Quantization experiments
  • Comparing memory/performance tradeoffs

Limitations

Quantization is inherently lossy.

Lower-bit quantizations generally reduce memory requirements but can introduce greater quality degradation.

Results can also vary depending on:

  • Prompt type
  • Context length
  • Sampling configuration
  • Inference engine
  • Hardware
  • Quantization method
  • Calibration/imatrix data

Therefore, users should evaluate the quantization that best fits their particular workload.

Disclaimer

This repository contains community-generated GGUF quantizations.

The original model's license, terms of use, safety policies, and attribution requirements remain applicable. Users should review the original model documentation and license before using or redistributing these files.

This repository does not claim ownership of the underlying model.

Credits

Original Model

Qwen3.8 27B Uncensored

Please refer to the original model release for:

  • Model architecture
  • Training information
  • License
  • Intended use
  • Safety considerations
  • Original model documentation

Quantization

Quantizations generated using:

llama.cpp

https://github.com/ggml-org/llama.cpp

Contributing

If you test one of the quantizations, useful feedback includes:

  • Quantization used
  • Hardware
  • RAM/VRAM
  • Context length
  • Tokens/second
  • Prompt processing speed
  • Backend/inference engine
  • Observed quality differences

Pull requests and benchmark contributions are welcome.

Changelog

2026-08-17

  • Added F16 GGUF
  • Added Q4_K_M GGUF
  • Began generating additional quantization variants
  • Started public benchmarking/quantization comparison project

More quantizations and benchmark results will be added as testing progresses.

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-17Update README.md11c45e27 KB
    Loading...
  2. 2026-08-17Create README.mdbfe10ac6.8 KB
    Loading...

Discussions 1 thread

  1. 2026-08-18Seems to work, thanks!closed2 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration