← back to catalog · registered 2026-08-22 13:56

Bucoid/Qwen3.8-27B-Uncensored-IQ4-XS-MTP-16GB-VRAM-GGUF

Bucoid Qwen 27B GGUF 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/Bucoid%2FQwen3.8-27B-Uncensored-IQ4-XS-MTP-16GB-VRAM-GGUF"
Response includes
  • classification m-uncensored
  • files 3
  • hub_downloads_all_time 140,971
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
141K
25K last 30d - stable
Likes
60
Model age
8w ago
created 2026-08-16
Downloads over time
Now143.1K→from8.6K↑1,564%
052.3K104.6K156.8K8.6K on Aug 19143.1K on Oct 11AugSepOct
Aug 19 → Oct 11 · 49 snapshots · spans 53 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Quantizations
IQ4
Tags
gguf base_model:Qwen/Qwen3.8-27B base_model:quantized:Qwen/Qwen3.8-27B license:apache-2.0 endpoints_compatible region:us imatrix conversational

Related

Total size
13.0 GB
Files
3
Quantizations
2
Registered
2026-08-22 13:56
Last updated on HF
2026-08-18 11:52

Files by quantization

IQ4 1 file 13.0 GB
Qwen3.8-27B-Uncensored-IQ4_XS_4BPW.gguf 13.0 GB 5e7f62e8 download
Auxiliary files 2 files 5.13 KB
README.md 3.57 KB 2a43b33f download
.gitattributes 1.56 KB 7a6126eb download

README current version from Hugging Face


license: apache-2.0
base_model:

  • Qwen/Qwen3.8-27B

此模型可能短期更新,我发布了一个基于Heretic Arbitrary-Rank Ablation的无审查版本模型,性能更好且体积更小
链接:
https://huggingface.co/Bucoid/Qwen3.8-27B-Heretic-Ara-IQ4-XS-16GB-VRAM-GGUF
这个模型可能过一段时间我会更新让他不那么菜,如果你需要无审查版本的模型,基于下载这个Ara的

This model may receive short-term updates.
I have released an uncensored version based on Heretic Arbitrary-Rank Ablation,
which offers better performance and a smaller file size.
Link: https://huggingface.co/Bucoid/Qwen3.8-27B-Heretic-Ara-IQ4-XS-16GB-VRAM-GGUF
This model may be updated in a while to make it less underwhelming
If you need an uncensored version, please download this Ara-based one instead.

Qwen3.8-27B Uncensored IQ4_XS 量化模型(适配 16GB 显存)
本模型基于 Qwen3.8-27B Uncensored 进行 IQ4_XS 量化(4‑bit),文件体积为 12.9 GiB,专为 16GB 显存 的显卡优化,在保持较低困惑度的同时,兼顾推理速度和显存占用。

与同体积的 UD_IQ3_K_XL(12.5 GiB)量化方案进行了全面对比,评估指标如下。

📊 量化质量对比

评估指标 IQ4_XS (本模型) UD_IQ3_K_XL (对比)
文件大小 12.9 GB 12.5 GB
量化精度 IQ4_XS (4‑bit) UD_IQ3_K_XL (约 3‑bit?)
量化模型困惑度 (Mean PPL) 7.1481 ± 0.0465 7.1117 ± 0.0459
与基座模型 PPL 相关性 99.28% 99.31%
平均 KL 散度 (Mean KLD) 0.03268 ± 0.00030 0.03130 ± 0.00032
最大 KL 散度 (Max KLD) 16.017(更小) 21.409
99.9% KL 分位数 1.075 1.219
Top‑1 一致率 (Same top p) 91.655% ± 0.072% 92.419% ± 0.069%
平均概率变化 (Mean Δp) -0.343% ± 0.013%(更接近 0) -0.738% ± 0.013%
RMS 概率变化 (RMS Δp) 4.986% ± 0.039%(更小) 5.120% ± 0.046%

在不启用MTP的情况下可以做到16GiB净空VRAM(不作为Windows的显示显卡)的情况下110k上下文

开启MTP大概80k上下文。


license: apache-2.0
base_model:

  • Qwen/Qwen3.8-27B

Qwen3.8-27B Uncensored IQ4_XS Quantized Model (Optimized for 16GB VRAM)
This model is based on Qwen3.8-27B Uncensored and quantized with IQ4_XS (4‑bit), with a file size of 12.9 GiB. It is tailored for GPUs with 16GB VRAM, balancing low perplexity, inference speed, and memory usage.

We conducted a comprehensive comparison against the UD_IQ3_K_XL quantization scheme (12.5 GiB, roughly 3‑bit) of the same model size. The evaluation metrics are as follows.

📊 Quantization Quality Comparison

Metric IQ4_XS (this model) UD_IQ3_K_XL (baseline)
File size 12.9 GB 12.5 GB
Quantization precision IQ4_XS (4‑bit) UD_IQ3_K_XL (~3‑bit)
Mean perplexity (quantized) 7.1481 ± 0.0465 7.1117 ± 0.0459
Correlation with base model PPL 99.28% 99.31%
Mean KL divergence 0.03268 ± 0.00030 0.03130 ± 0.00032
Maximum KL divergence 16.017 (lower) 21.409
99.9% KL quantile 1.075 1.219
Top‑1 agreement rate 91.655% ± 0.072% 92.419% ± 0.069%
Mean probability change (Mean Δp) -0.343% ± 0.013% (closer to 0) -0.738% ± 0.013%
RMS probability change (RMS Δp) 4.986% ± 0.039% (lower) 5.120% ± 0.046%

With MTP disabled, the model can achieve ~110k context length while keeping ~16 GiB free VRAM (when not used as the primary display GPU on Windows). With MTP enabled, the context length is around 80k.

README history 6 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-18Update README.md5ae30aa3.6 KB
    Loading...
  2. 2026-08-18Update README.md23eaf183.6 KB
    Loading...
  3. 2026-08-18Update README.md33c62de3.6 KB
    Loading...
  4. 2026-08-16Update README.md3cbc4cc2.8 KB
    Loading...
  5. 2026-08-16Update README.mdd1d9a542.5 KB
    Loading...
  6. 2026-08-16initial commit4352ee828 B
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration