← back to catalog · registered 2026-08-22 13:56

vcruz305/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-GGUF

vcruz305 Qwen 27B GGUF second-order 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/vcruz305%2FQwen3.8-27B-AEON-ULTIMATE-UNCENSORED-GGUF"
Response includes
  • classification m8
  • files 12
  • hub_downloads_all_time 41,737
  • author_summary 7 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
42K
4K last 30d - cooling
Likes
32
Model age
8w ago
created 2026-08-15
Downloads over time
Now42.8K→from8K↑438%
6.2K19.6K32.9K46.3K8K on Aug 1742.8K on Oct 11AugSepOct
Aug 17 → Oct 11 · 49 snapshots · spans 55 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh
Quantizations
F16 Q2_K Q3_K Q4 Q4_K Q5_K Q6_K Q8_0
Tags
gguf qwen qwen3.8 llama.cpp uncensored abliterated aeon mtp text-generation conversational en zh

Related

Total size
115 GB
Files
12
Quantizations
9
Registered
2026-08-22 13:56
Last updated on HF
2026-08-16 23:22

Files by quantization

Q8_0 2 files 30.0 GB
Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-Q8_0.gguf 27.1 GB 5de61c48 download
mtp-Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-Q8_0.gguf 2.95 GB c0eeee94 download
Q6_K 1 file 20.9 GB
Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-Q6_K.gguf 20.9 GB c68230d9 download
Q5_K 1 file 18.2 GB
Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-Q5_K_M.gguf 18.2 GB 42a2c535 download
Q4_K 1 file 15.7 GB
Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-Q4_K_M.gguf 15.7 GB d4396fed download
Q3_K 1 file 12.6 GB
Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-Q3_K_M.gguf 12.6 GB 04019115 download
Q2_K 1 file 10.1 GB
Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-Q2_K.gguf 10.1 GB b573e26d download
F16 1 file 5.54 GB
mtp-Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-F16.gguf 5.54 GB 764404d0 download
Q4 1 file 1.87 GB
mtp-Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-Q4_0.gguf 1.87 GB 993f410a download
Auxiliary files 3 files 24.9 KB
chat_template.jinja 18.8 KB a921aafd download
README.md 3.85 KB 7e7f1ea5 download
.gitattributes 2.23 KB c7fc9d99 download

README current version from Hugging Face


language:

  • en
  • zh
    license: apache-2.0
    library_name: gguf
    pipeline_tag: text-generation
    base_model: AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16
    base_model_relation: quantized
    quantized_by: vcruz305
    tags:
  • gguf
  • qwen
  • qwen3.8
  • llama.cpp
  • uncensored
  • abliterated
  • aeon
  • mtp

Qwen3.8-27B AEON Ultimate Uncensored GGUF

llama.cpp K-quants of AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16, an abliterated BF16 of Qwen/Qwen3.8-27B.

This is not official Qwen. It is also not vcruz305/Qwen3.8-27B-GGUF (official trunk) or vcruz305/Qwen3.8-27B-Uncensored-GGUF (orcarouter FP8).

MTP is baked into every Q2–Q8 file (866 tensors, qwen35.nextn_predict_layers=1, +~0.24 GiB vs the old trunk-only files). You do not need a second GGUF for draft-mtp.

What is in these files

27B dense hybrid-attention (qwen35). 64 language-trunk blocks plus 1 nextn/MTP block. Converted from the AEON BF16 master (no --no-mtp). Pair a separate mmproj if you need vision.

The source removed the refusal direction. These GGUFs inherit that.

Chat template

Official 3.8 jinja wraps empty <think> blocks and breaks multi-turn agents. These GGUFs bake a fixed template. Use --jinja.

llama-server -m Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-Q4_K_M.gguf --jinja --reasoning-format deepseek

Files

File Quant Bytes GiB Notes
Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-Q2_K.gguf Q2_K 10864592928 10.12 live; MTP baked in
Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-Q3_K_M.gguf Q3_K_M 13500737568 12.57 live; MTP baked in
Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-Q4_K_M.gguf Q4_K_M 16810715168 15.66 live; MTP baked in
Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-Q5_K_M.gguf Q5_K_M 19535702048 18.20 live; MTP baked in
Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-Q6_K.gguf Q6_K 22431000608 20.89 live; MTP baked in; largest full-GPU on 24GB Turing
Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-Q8_0.gguf Q8_0 29047085088 27.05 live; MTP baked in; will not -ngl 99 on 24GB

How to run MTP

Needs a llama.cpp build that understands Qwen3.5 nextn / draft-mtp. --parallel 1 is required.

llama-server \
  -m Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-Q4_K_M.gguf \
  --spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-p-min 0.7 \
  --parallel 1 --jinja --reasoning-format deepseek \
  -a qwen38-27b-aeon --host 127.0.0.1 --port 8085 \
  -ngl 99 -fa on -b 512 -ub 512 -c 32768

Optional extra mtp-* sidecars are still in the repo if you want a separate -md draft. They are not required for the files above.

Download

Use hf_xet. Do not git clone. Grab only the file that exists.

export HF_XET_HIGH_PERFORMANCE=1
hf download vcruz305/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-GGUF \
  --local-dir Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-GGUF \
  --include "Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-Q4_K_M.gguf"

Q6_K is the largest file that still full-offloads a 24GB Turing card. Q8_0 does not.

Intended use

Local llama.cpp of the AEON uncensored 27B trunk + native MTP. Out of scope: treating this as official Qwen, or as a drop-in for the official-trunk GGUF pack.

Source

Credits

Abliteration / BF16: AEON-7. Base: Qwen / Alibaba. GGUF pack: Victor Cruz (vcruz305).

README history 8 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-16Upload README.md with huggingface_hubee09a863.8 KB
    Loading...
  2. 2026-08-16Upload README.md with huggingface_hube5098e94.6 KB
    Loading...
  3. 2026-08-16Upload README.md with huggingface_hubd05e6c83.3 KB
    Loading...
  4. 2026-08-16Upload README.md with huggingface_hub8e1de783.4 KB
    Loading...
  5. 2026-08-15Upload README.md with huggingface_hub05beaac3.4 KB
    Loading...
  6. 2026-08-15Upload README.md with huggingface_hub0d5be4a3.4 KB
    Loading...
  7. 2026-08-15Upload README.md with huggingface_hub772b2033.7 KB
    Loading...
  8. 2026-08-15Upload README.md with huggingface_hube21db323.4 KB
    Loading...

Discussions 1 thread

  1. 2026-08-17Qwen3.8-AEON (The Challenger)closed2 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration