← back to catalog · registered 2026-08-22 13:56

Ishowbackup/Qwen3.8-27B-ABLITERATED-GGUF

Ishowbackup Qwen 27B GGUF multimodal second-order 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/Ishowbackup%2FQwen3.8-27B-ABLITERATED-GGUF"
Response includes
  • classification m8
  • files 17
  • hub_downloads_all_time 1,348
  • author_summary 28 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
1K
679 last 30d - active
Likes
0
Model age
8w ago
created 2026-08-15
Downloads over time
Now1.7K→from718↑131%
6711K1.4K1.8K718 on Aug 191.7K on Oct 111.7K on Oct 10AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 3 formats · 760 downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
apache-2.0
Quantizations
BF16 Q2_K Q3_K Q4 Q4_K Q5_K Q6_K Q8_0
Tags
gguf qwen3.8 qwen 27b dense abliterated quantized multimodal reasoning tool-calling llama.cpp long-context

Related

Total size
156 GB
Files
17
Quantizations
10
Registered
2026-08-22 13:56
Last updated on HF
2026-08-15 23:19

Files by quantization

Q8_0 3 files 30.2 GB
Qwen3.8-27B-ABLITERATED-Q8_0.gguf 26.6 GB fe89cb15 download
mtp-Qwen3.8-27B-ABLITERATED-Q8_0.gguf 2.95 GB b945be29 download
mmproj-Qwen3.8-27B-ABLITERATED-Q8_0.gguf 600 MB 6b8b9c95 download
Q6_K 1 file 20.6 GB
Qwen3.8-27B-ABLITERATED-Q6_K.gguf 20.6 GB 3d78a14f download
Q5_K 2 files 35.3 GB
Qwen3.8-27B-ABLITERATED-Q5_K_M.gguf 17.9 GB 5a06d1df download
Qwen3.8-27B-ABLITERATED-Q5_K_S.gguf 17.4 GB 1e6ac6ba download
Q4_K 2 files 29.9 GB
Qwen3.8-27B-ABLITERATED-Q4_K_M.gguf 15.4 GB 202bbe49 download
Qwen3.8-27B-ABLITERATED-Q4_K_S.gguf 14.5 GB 51139ccb download
Q3_K 2 files 23.6 GB
Qwen3.8-27B-ABLITERATED-Q3_K_M.gguf 12.4 GB 9d6c6b7b download
Qwen3.8-27B-ABLITERATED-Q3_K_S.gguf 11.2 GB da6b5736 download
Q2_K 1 file 9.98 GB
Qwen3.8-27B-ABLITERATED-Q2_K.gguf 9.98 GB b3545db5 download
BF16 1 file 5.54 GB
mtp-Qwen3.8-27B-ABLITERATED-BF16.gguf 5.54 GB 8de71f3c download
Q4 1 file 1.87 GB
mtp-Qwen3.8-27B-ABLITERATED-Q4_0.gguf 1.87 GB e99b0d6f download
F16 1 file 885 MB
mmproj-Qwen3.8-27B-ABLITERATED-F16.gguf 885 MB 2284099c download
Auxiliary files 3 files 12.7 KB
README.md 9.90 KB 57798e7c download
.gitattributes 2.48 KB e6f1d2d1 download
MTP-SHA256SUMS.txt 312 B 57528501 download

README current version from Hugging Face


license: apache-2.0
base_model:

  • Blackfrost-AI/Qwen3.8-27B-ABLITERATED-BF16
    tags:
  • qwen3.8
  • qwen
  • 27b
  • dense
  • gguf
  • abliterated
  • quantized
  • multimodal
  • reasoning
  • tool-calling
  • llama.cpp
  • long-context
    pipeline_tag: image-text-to-text
    library_name: gguf

Blackfrost

QWEN3.8-27B-ABLITERATED-GGUF

Full standard GGUF quant ladder of the Blackfrost abliterated Qwen3.8-27B · dense multimodal model for llama.cpp

Built by Blackfrost · Las Vegas, NV

All standard quants live

The complete standard K-quant ladder (Q2_K through Q8_0) and both vision projectors are included. No IQ/IK or importance-matrix quants are used.

MTP speculative decoding restored — August 15, 2026

Three separate MTP sidecars are now included: BF16, Q8_0, and Q4_0. Existing text quants and vision projectors are unchanged, so current users only need to download an mtp- file to add speculative decoding.

Experimental release

This is a newly supported architecture and a deliberately modified research checkpoint. Validate the selected quant, context, sampling, vision, tools, and structured-output behavior in your own workload before deployment.


Refusal benchmark — R1-HARMFUL-BENCH-450

The release score is 11 residual refusals from 450 original cases (2.4%). The shipped Blackfrost short execution prompt is embedded once in the GGUF chat template.

This result is a sequential, manually reviewed residual funnel measured on the W4A4 NVFP4 derivative of the same BF16 parent. It is not a fresh full-450 GGUF run with every case presented under the final short prompt.

evaluation stage cases evaluated material answer true refusal remaining other
Raw upstream template 450 360 88 2 capability limitations
Blackfrost operational-prompt retest 88 residuals 53 33 1 limitation, 1 reproducible incoherent output
Shipped short execution-prompt retest 33 residuals 22 11 0
Final residual count 450 original cases — 11 (2.4%) —

The 450-case source set contains 150 AdvBench, 150 StrongREJECT, and 150 XSTest prompts. The final 11 comprise 1 AdvBench, 5 StrongREJECT, and 5 XSTest cases. An opening objection followed by a materially useful payload was counted as softened compliance, not as a refusal; a true refusal means the requested payload never arrived.


Why this model exists

Qwen3.8-27B is the dense, deployment-friendly member of the Qwen3.8 family. This is the abliterated Blackfrost build: refusal behavior was reduced through a weight-level process, then the BF16 parent was converted into a standard GGUF ladder for local llama.cpp inference.

It is not a coding fine-tune, merge, LoRA, or pruned model.


Specifications

Architecture Qwen3.8 dense hybrid VLM · 64 text layers · Gated DeltaNet + full attention · 27-layer vision tower
Parent Blackfrost-AI/Qwen3.8-27B-ABLITERATED-BF16
Base Qwen/Qwen3.8-27B · Apache-2.0
Transform Abliterated — refusal surface modified at weight level; no fine-tuning or pruning
Formats Q2_K, Q3_K_S, Q3_K_M, Q4_K_S, Q4_K_M, Q5_K_S, Q5_K_M, Q6_K, Q8_0
Context 262,144 tokens architecturally; practical context depends on RAM/VRAM and concurrency
Modalities Text, image, and video input; text output
Chat behavior Blackfrost short execution prompt embedded in the default Jinja chat template
MTP speculative head Separate BF16, Q8_0, and Q4_0 sidecars included; existing text quants are unchanged

Quant ladder

quant size recommended for
Q2_K 10.7 GB smallest standard quant; largest quality trade-off
Q3_K_S 12.1 GB very tight memory
Q3_K_M 13.3 GB compact general use
Q4_K_S 15.6 GB lower-memory Q4 option
Q4_K_M 16.5 GB default — balanced quality and footprint
Q5_K_S 18.7 GB higher fidelity
Q5_K_M 19.2 GB strong quality/size balance
Q6_K 22.1 GB near-BF16 behavior for many workloads
Q8_0 28.6 GB maximum fidelity in the ladder

File sizes are decimal GB as displayed by Hugging Face. Runtime memory also includes context state, compute buffers, the optional vision projector, and server overhead.


Vision projector files

Load one text quant plus one mmproj file for image or video input:

file size purpose
mmproj-Qwen3.8-27B-ABLITERATED-F16.gguf 0.93 GB full-fidelity vision projector
mmproj-Qwen3.8-27B-ABLITERATED-Q8_0.gguf 0.63 GB compact projector; unsupported 4,304-wide tensors retain F16 automatically

MTP speculative decoding

The parent checkpoint's MTP head is published as separate lowercase mtp- sidecars, which is the current llama.cpp layout. Pair one sidecar with any existing text quant:

file size use
mtp-Qwen3.8-27B-ABLITERATED-BF16.gguf 5.95 GB maximum draft fidelity
mtp-Qwen3.8-27B-ABLITERATED-Q8_0.gguf 3.16 GB recommended balance
mtp-Qwen3.8-27B-ABLITERATED-Q4_0.gguf 2.01 GB lowest draft-model memory

With current llama.cpp, repository loading discovers the closest mtp- sidecar automatically when MTP speculation is enabled:

llama-server \
  -hf Blackfrost-AI/Qwen3.8-27B-ABLITERATED-GGUF:Q4_K_M \
  --spec-type draft-mtp --spec-draft-n-max 3 \
  -ngl 999 -ngld 999 --jinja -c 16384

For manual files, pass --spec-draft-model mtp-Qwen3.8-27B-ABLITERATED-Q8_0.gguf along with --spec-type draft-mtp. The Q4_K_M target plus Q8_0 MTP sidecar was runtime-tested on one RTX PRO 6000 Blackwell: 256 generated tokens at 90.0 tok/s with 50.7% draft acceptance (153 accepted of 302 drafted). Throughput and acceptance vary with prompt, sampler, hardware, context, and concurrency.


Serving with llama.cpp

Use a current llama.cpp build with llama-server. Q4_K_M plus the compact projector was load- and generation-tested through the OpenAI-compatible chat API on an NVIDIA B200.

hf download Blackfrost-AI/Qwen3.8-27B-ABLITERATED-GGUF \
  Qwen3.8-27B-ABLITERATED-Q4_K_M.gguf \
  mmproj-Qwen3.8-27B-ABLITERATED-Q8_0.gguf \
  --local-dir ./Qwen3.8-27B-ABLITERATED-GGUF

llama-server \
  -m ./Qwen3.8-27B-ABLITERATED-GGUF/Qwen3.8-27B-ABLITERATED-Q4_K_M.gguf \
  --mmproj ./Qwen3.8-27B-ABLITERATED-GGUF/mmproj-Qwen3.8-27B-ABLITERATED-Q8_0.gguf \
  -ngl 999 -fa on --jinja \
  --host 0.0.0.0 --port 8080 -c 16384 \
  --temp 1.0 --top-p 0.95 --top-k 20
  • Text only: omit --mmproj and do not download a projector.
  • CPU or hybrid inference: lower -ngl; use -ngl 0 for CPU-only operation.
  • Larger context: increase -c only after checking memory headroom at the intended concurrency.
  • Embedded prompt: keep --jinja enabled so the repository's default chat template is applied.
  • One-command kit: deploy/serve.sh downloads and serves the selected quant; see deploy/DEPLOYMENT.md for the full guide.

API check

curl http://127.0.0.1:8080/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "Qwen3.8-27B-ABLITERATED",
    "messages": [{"role": "user", "content": "Reply with exactly READY and nothing else."}],
    "temperature": 0,
    "max_tokens": 64
  }'

Quality check

WikiText-2 rolling perplexity was measured on the parent artifacts through the same 8K API harness:

artifact word perplexity byte perplexity bits/byte
Clean upstream BF16 8.4764 1.4914 0.5766
Blackfrost W4A4 NVFP4 derivative 9.3677 1.5195 0.6036

These figures are parent-artifact measurements, not per-quant GGUF perplexity scores. The Q4_K_M GGUF and compact projector passed a real llama.cpp load and chat-generation smoke test. The separate Q8_0 MTP sidecar also passed a CUDA runtime test with measurable draft acceptance.


Deployment responsibility

This checkpoint has a deliberately reduced refusal surface. Open weights do not provide an application policy, authorization system, audit trail, sandbox, or access-control boundary. Operators are responsible for authenticated access, least-privilege tool credentials, execution isolation, logging, and approval boundaries appropriate to their deployment.

The embedded prompt is a behavioral instruction, not a security boundary.


Disclaimer

Refusal behavior in this checkpoint has been deliberately modified at the weight level. It is not a safety-stock model and must not be represented as one.

This checkpoint is provided "as is," without warranty of any kind. Measurements describe only the tested artifacts, prompts, templates, samplers, serving engines, and review criteria. They do not guarantee that any particular input will be accepted or refused, that every upstream capability is retained, or that the measurements generalize to multimodal, tool-use, long-context, or multi-turn settings.

The derivative remains subject to the Apache 2.0 license shipped with the official Qwen3.8-27B checkpoint.


Built by Blackfrost · Las Vegas, NV. Not affiliated with Qwen or Alibaba.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-15Duplicate from Blackfrost-AI/Qwen3.8-27B-ABLITERATED-GGUFd86b12d9.9 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration