← back to catalog · registered 2026-08-22 13:56

vcruz305/Qwen3.8-27B-Uncensored-GGUF

vcruz305 Qwen 27B GGUF second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/vcruz305%2FQwen3.8-27B-Uncensored-GGUF"
Response includes
  • classification m8
  • files 2
  • author_summary 7 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
3
Model age
8w ago
created 2026-08-15
Downloads over time
Now0→from0↑0%
00110 on Aug 190 on Oct 11AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh
Tags
gguf qwen qwen3.8 llama.cpp uncensored abliterated text-generation en zh base_model:orcarouter/Qwen3.8-27B-Uncensored-FP8 base_model:quantized:orcarouter/Qwen3.8-27B-Uncensored-FP8 license:apache-2.0

Related

Total size
0 B
Files
2
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-15 19:55

Files by quantization

Auxiliary files 2 files 5.58 KB
README.md 4.10 KB b12aca93 download
.gitattributes 1.48 KB a6344aac download

README current version from Hugging Face


language:

  • en
  • zh
    license: apache-2.0
    library_name: gguf
    pipeline_tag: text-generation
    base_model: orcarouter/Qwen3.8-27B-Uncensored-FP8
    base_model_relation: quantized
    quantized_by: vcruz305
    tags:
  • gguf
  • qwen
  • qwen3.8
  • llama.cpp
  • uncensored
  • abliterated

Qwen3.8-27B Uncensored GGUF

Standalone llama.cpp K-quants of orcarouter/Qwen3.8-27B-Uncensored-FP8, a community abliterated block-FP8 of Qwen/Qwen3.8-27B.

This is not official Qwen. It is also not vcruz305/Qwen3.8-27B-GGUF — that pack is the official BF16 trunk.

What is in these files

27B dense hybrid-attention (qwen35). 64 language-trunk blocks (0–63). Hidden 5120, FFN 17408. Native context 262,144.

MTP / nextn is omitted (--no-mtp). Speculative decode does not make the model smarter; the extra head steals KV on 12–24 GB cards. Need vision? Pair a separate mmproj. Need MTP? Use another pack.

The source checkpoint had the refusal direction removed (abliteration). These GGUFs inherit that behavior.

Chat template

Official 3.8 jinja wraps every assistant turn in <think>…</think> even when reasoning is empty, then opens another <think> on generate. That truncates multi-turn agents.

These GGUFs bake a fixed template. Use --jinja. A standalone chat_template.jinja ships in the repo if an older copy is still on disk.

llama-server -m Qwen3.8-27B-Uncensored-Q4_K_M.gguf --jinja --reasoning-format deepseek

Files

One file per quant. Byte / GiB filled in when the ladder lands.

File Quant Bytes GiB Notes
Qwen3.8-27B-Uncensored-Q2_K.gguf Q2_K TBD TBD 12GB start
Qwen3.8-27B-Uncensored-Q3_K_M.gguf Q3_K_M TBD TBD 16GB
Qwen3.8-27B-Uncensored-Q4_K_M.gguf Q4_K_M TBD TBD 24GB start — default
Qwen3.8-27B-Uncensored-Q5_K_M.gguf Q5_K_M TBD TBD 24GB comfortable
Qwen3.8-27B-Uncensored-Q6_K.gguf Q6_K TBD TBD Largest full-GPU on 24GB Turing
Qwen3.8-27B-Uncensored-Q8_0.gguf Q8_0 TBD TBD 32GB+; will not -ngl 99 on 24GB

Download

Use hf_xet. Do not git clone.

export HF_XET_HIGH_PERFORMANCE=1
hf download vcruz305/Qwen3.8-27B-Uncensored-GGUF \
  --local-dir Qwen3.8-27B-Uncensored-GGUF \
  --include "Qwen3.8-27B-Uncensored-Q4_K_M.gguf"

Change --include for the quant you want.

How to run

Needs llama.cpp new enough for qwen35 (Gated DeltaNet hybrid).

24GB (default Q4_K_M):

llama-server \
  -m Qwen3.8-27B-Uncensored-GGUF/Qwen3.8-27B-Uncensored-Q4_K_M.gguf \
  -a qwen38-27b-unc \
  --host 127.0.0.1 --port 8085 \
  -ngl 99 -c 32768 -np 1 --jinja --reasoning-format deepseek

Q6_K is the largest file that still full-offloads 24GB Turing. Q8_0 does not (-ngl 99 will not fit).

Intended use

Local llama.cpp serving of the uncensored 27B trunk: research, red-team, and unfiltered generation in a setting you control.

Out of scope: treating this as official Qwen or as a drop-in for vcruz305/Qwen3.8-27B-GGUF; deploying to end users without your own filters; any use that breaks Apache-2.0 or the law.

Bias, risks, limitations

Safety alignment was removed at the source. The model will answer requests the official 27B would refuse. It still carries the bias and failure modes of Qwen3.8-27B, plus K-quant error. These files are language-only (no vision tower, no MTP).

Source

Credits

Abliteration and FP8: orcarouter. Base model: Qwen / Alibaba. GGUF pack: Victor Cruz (vcruz305).

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-15Upload README.md with huggingface_hub2fe68794.1 KB
    Loading...
  2. 2026-08-15Upload README.md with huggingface_hubb88dd8d4.1 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration