← back to catalog · registered 2026-08-26 12:02

guideboardlabs/SuperQwen3.8-27B-abliterated-Q3-DOWN-XS-GGUF

guideboardlabs Qwen 27B GGUF second-order 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/guideboardlabs%2FSuperQwen3.8-27B-abliterated-Q3-DOWN-XS-GGUF"
Response includes
  • classification m8
  • files 3
  • hub_downloads_all_time 437
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
437
89 last 30d - stable
Likes
0
Model age
6w ago
created 2026-08-26

Training datasets

Corpora the author lists in the model card. Datasets tracked in our /datasets catalog carry a category badge linking to the workflow stage. Others open on Hugging Face.

Downloads over time
Now447→from0↑0%
01643284920 on Aug 26447 on Oct 11447 on Oct 10AugSepOct
Aug 26 → Oct 11 · 47 snapshots · spans 46 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
gguf qwen qwen3.5 quantization amd vulkan rDNA1 Radeon llama.cpp local-llm abliterated dataset:Jiunsong/SuperQwen3.8-27b-abliterated

Related

Total size
7.72 GB
Files
3
Quantizations
1
Registered
2026-08-26 12:02
Last updated on HF
2026-08-26 11:19

Files by quantization

Auxiliary files 3 files 7.72 GB
SuperQwen3.8-27B-abliterated-Q3-DOWN-XS.gguf 7.72 GB 47c47d10 download
README.md 4.15 KB 3dcfd8c2 download
.gitattributes 1.56 KB c23359a8 download

README current version from Hugging Face


license: apache-2.0
tags:

  • qwen
  • qwen3.5
  • gguf
  • quantization
  • amd
  • vulkan
  • rDNA1
  • Radeon
  • llama.cpp
  • local-llm
  • abliterated
    base_model:
  • Jiunsong/SuperQwen3.8-27b-abliterated
    datasets:
  • Jiunsong/SuperQwen3.8-27b-abliterated

SuperQwen3.8-27B (abliterated) — Q3-DOWN-XS GGUF (fully resident on 8 GB RDNA1)

A 7.73 GiB GGUF of the abliterated SuperQwen3.8-27B, quantized with the exact same Q3-DOWN-XS recipe as the base model — proving the recipe generalizes to other weight sets. Fully resident in 8 GB VRAM, zero CPU spill, measured on an AMD Radeon RX 5700 XT (RDNA1, gfx1010, 8 GB).

This is the "recipe generalizes" proof. Same tensor-type map, same imatrix, same flags as guideboardlabs/Qwen3.8-27B-Q3-DOWN-XS-GGUF — applied to a different weight set (abliterated). It runs at the same speed and scores higher on the same capability harness.

File Size bpw Decode Agon Sprint (/26)
SuperQwen3.8-27B-abliterated-Q3-DOWN-XS.gguf 7.73 GiB ~3.2 24.40 tok/s 23/26

The honest comparison: on the same 8-task / 26-point Agon Sprint (medium reasoning, seed 42), the abliterated Q3-DOWN-XS scores 23/26 vs the base Q3-DOWN-XS's 21/26 — a +2 net gain. The gain is concentrated in coding (String Cleaner 4/5 vs base's 0/5); the cost is one pure-reasoning task (Logical Deduction 2/2 → 0/2). So it's a trade, not a strict upgrade: better at code/tool-following, slightly worse at one logical-reasoning task. Speed is identical (24.40 vs 24.48 tok/s, within 0.3%).

Hardware measured on: RX 5700 XT (8 GB, RDNA1) + Ryzen 5 1600, llama.cpp Vulkan build.


Why "Q3-DOWN-XS" (the recipe)

The model is a per-tensor-tuned mix of stock quant types, chosen by testing every combination on the GPU. "DOWN-XS" = the FFN down projections use IQ2_XS while the rest uses a higher-quality mix. This is the same recipe as the base model — the point is that it transfers to other weights unchanged.


How to run it (the recipe)

1. Serve it with stock llama.cpp (Vulkan)

GGML_VK_DISABLE_HOST_VISIBLE_VIDMEM=1 llama-server \
  -m SuperQwen3.8-27B-abliterated-Q3-DOWN-XS.gguf \
  -c 16384 -np 1 -ngl 999 \
  -ctk q8_0 -ctv q8_0 -fa on \
  --jinja --reasoning off --reasoning-format none \
  --port 8101 --host 0.0.0.0

2. The ONE environment variable that matters

GGML_VK_DISABLE_HOST_VISIBLE_VIDMEM=1

This is the single biggest measured win: +48.6% decode throughput. Without it, part of the hot working set gets allocated in host-visible (PCIe-accessible) VRAM, and the card thrashes through it at ~13 GB/s instead of the resident ~300 GB/s. The fix forces full residency.

3. Set the GPU to COMPUTE DPM profile

# as root
rocm-smi --setperflevel 5        # COMPUTE
# or via sysfs (RX 5700 XT):
echo COMPUTE | sudo tee /sys/class/drm/card0/device/power_dpm_force_performance_level

4. Keep it fully resident (-ngl 999)

Offload every layer. Do NOT let layers spill to CPU — a single CPU layer's latency lands on top of the GPU's, killing throughput.


Why this is reproducible (no custom kernel)

  • 100% stock quant types (IQ2_XS, IQ4_XS, IQ2_XXS, IQ1_M, IQ3_S, F32...) — verified, zero custom/RDNA tensor types.
  • Stock llama.cpp — the fork is a superset; upstream runs this file unchanged.
  • Stock env var + stock AMD setting + stock server flags.

An average user with plain llama.cpp + this file + the env var + COMPUTE DPM gets the same result. That's the point of the recipe.


Build / provenance

  • Base: Jiunsong/SuperQwen3.8-27b-abliterated (safetensors, qwen3_5 arch)
  • Converted to BF16 GGUF (text-only, --no-mtp), then quantized with llama.cpp llama-quantize using the exact same q3-down-xs.tensor-types.txt map + base imatrix as the base model.
  • Capability: Agon Sprint 8-task / 26-point harness, medium reasoning, seed 42, same flags for both models.
  • Speed: frozen protocol llama-bench -p 512 -n 128 -r 5, Vulkan0, same build/GPU.

License

Apache-2.0. Model: SuperQwen3.8-27B-abliterated (Apache-2.0). See original for details.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-26Add SuperQwen3.8-27B abliterated Q3-DOWN-XS GGUF (recipe generalizes)46476724.2 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration