← back to catalog · registered 2026-08-24 17:02

Jiunsong/SuperQwen3.8-27b-abliterated-GGUF

Jiunsong Qwen 27B GGUF multimodal second-order 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/Jiunsong%2FSuperQwen3.8-27b-abliterated-GGUF"
Response includes
  • classification m8
  • files 7
  • hub_downloads_all_time 17,837
  • author_summary 35 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
18K
2K last 30d - stable
Likes
30
Model age
6w ago
created 2026-08-24
Downloads over time
Now19.4K→from5.4K↑257%
4.7K10.1K15.4K20.8K5.4K on Aug 2619.4K on Oct 11AugSepOct
Aug 26 → Oct 11 · 47 snapshots · spans 46 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 2 formats · 3K downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
apache-2.0
Languages
en ko
Quantizations
Q4 Q4_K
Tags
gguf qwen3.8 qwen3.5 q4-k-m multimodal reasoning tool-calling long-context uncensored abliterated obliteratus speculative-decoding

Related

Total size
17.0 GB
Files
7
Quantizations
4
Registered
2026-08-24 17:02
Last updated on HF
2026-08-25 02:23

Files by quantization

Q4_K 1 file 15.4 GB
SuperQwen3.8-27b-abliterated-Q4_K_M.gguf 15.4 GB 8edcf493 download
Q4 1 file 1.56 GB
mtp-SuperQwen3.8-27b-abliterated-Q4_0.gguf 1.56 GB 3f346731 download
Q8_0 1 file 600 MB
mmproj-SuperQwen3.8-27b-abliterated-Q8_0.gguf 600 MB ead8e9f2 download
Auxiliary files 4 files 25.5 KB
LICENSE 11.3 KB f938136e download
README.md 10.6 KB 4c4a7875 download
SHA256SUMS.json 1.95 KB 18f19627 download
.gitattributes 1.72 KB 27fdcd9b download

README current version from Hugging Face


license: apache-2.0
library_name: gguf
pipeline_tag: image-text-to-text
base_model: Jiunsong/SuperQwen3.8-27b-abliterated
base_model_relation: quantized
tags:

  • qwen3.8
  • qwen3.5
  • multimodal
  • image-text-to-text
  • reasoning
  • tool-calling
  • long-context
  • uncensored
  • abliterated
  • supertune
  • gguf
  • llama.cpp
  • q4-k-m
  • speculative-decoding
  • mtp
  • dgx-spark
    language:
  • en
  • ko

SuperQwen3.8 27B — BF16, NVFP4, GGUF and MLX

SuperQwen3.8-27b-abliterated-GGUF

The portable one-box edition: imatrix Q4_K_M, native Qwen MTP, and a separate multimodal projector.

C1 decode
Spec decode
Vision
Context
License

One model. Four native releases.

BF16 · NVFP4 · GGUF · MLX 4-bit

Choose your build

Release Best for Size / precision Runtime
BF16 Maximum fidelity and further tuning ~52 GB · BF16 Transformers / vLLM
NVFP4 Fast single-DGX-Spark serving ~19.2 GiB · W4A4 G16 vLLM
GGUF — this repo Portable one-box inference + native MTP ~17.6 GiB runtime set llama.cpp
MLX 4-bit Apple Silicon ~15.0 GiB · affine 4-bit MLX

[!NOTE]
An expanded post-release Obliteratus and cross-format audit is in progress. Existing
benchmark claims remain bound to the hashed evidence listed below and will be replaced,
not extrapolated, when the wider independent suite completes.

This is the llama.cpp-compatible release of
Jiunsong/SuperQwen3.8-27b-abliterated,
pinned to BF16 revision 84efc06504219387b308f15561ddfc8a2b880966 and ultimately to
Qwen/Qwen3.8-27B@1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0.
It keeps text generation, the model's native MTP draft head, and the vision projector in
separate files so each runtime component can use the precision that fits its job.

Release highlights

Verified release value
Hardware 1× NVIDIA DGX Spark / GB10, one request at a time (C1)
Target weights Importance-matrix-guided Q4_K_M, 15.41 GiB
Speculative draft Native Qwen MTP Q4_0, selected at K=3
Vision Mixed Q8_0/F16/F32 mmproj, directly tested with an image
C1 decode without speculation 12.4955 tok/s mean, p256/n512, 3 trials
C1 decode with MTP K=3 22.1 tok/s median, n512, 3 trials
OpenAI API fixed generation 23.535 tok/s end to end, 256/256 completion tokens
Functional gates Chat, reasoning, tool call, benign-sensitive, bounded reasoning, vision: PASS
Long-context gate 46,235 prompt tokens, hidden needle retrieved exactly

Why this release

  • Runs on one Spark. The three runtime artifacts total about 17.56 GiB and the
    verified server reported 20,760 MiB of GPU memory at a 65,536-token context allocation.
  • C1 is the headline. The promoted number is a single interactive request, not a
    concurrent aggregate relabeled as single-user speed.
  • Speculation was searched, not guessed. K=1, 3, 5, and 8 were measured; K=3 was
    fastest and then repeated three times at 512 generated tokens.
  • Multimodal remains real. The separate projector identified the llama in the
    release image gate, and the model emitted a valid forced function call.
  • Reasoning terminates. The embedded bounded template returned the one-line answer
    42 in three tokens, while a separate reasoning gate correctly produced 1517.

Files and integrity

File Size SHA-256 Role
SuperQwen3.8-27b-abliterated-Q4_K_M.gguf 16,547,401,056 B 8edcf493410cd049abf25c3c5ddd0c28dde0b22d44278bfb899f82683b1e4e43 851-tensor target
mtp-SuperQwen3.8-27b-abliterated-Q4_0.gguf 1,680,272,384 B 3f346731cb86f664351b41fde590b9cecd4f66473a73e8e81bb33fdc2d918392 18-tensor MTP draft
mmproj-SuperQwen3.8-27b-abliterated-Q8_0.gguf 629,247,680 B ead8e9f2839910775e3f871109a6a84392fc7c9c8e00e34997bb3819c18bbfe8 334-tensor vision projector

The target uses the standard Q4_K_M mix: 433 Q4_K tensors, 65 Q6_K tensors, and 353
small or structurally sensitive F32 tensors. The MTP file contains 10 Q4_0 tensors and
8 F32 tensors. The mmproj contains 83 Q8_0, 27 F16, and 224 F32 tensors; unsupported
4304-column vision matrices are deliberately preserved instead of being force-packed.

Target quantization used the Unsloth Qwen3.8 importance matrix with SHA-256
0ee5b10bd0c2fa2127c6f4b43dbfe1efd71e383b63217af9dade1de36599f1c1.
Conversion and validation used llama.cpp revision
b3c3b96a139d4ef1bdec926ac17aa040981cfc5d compiled for CUDA architecture 121a.

Serving on one DGX Spark

llama-server \
  -m SuperQwen3.8-27b-abliterated-Q4_K_M.gguf \
  -md mtp-SuperQwen3.8-27b-abliterated-Q4_0.gguf \
  --spec-type draft-mtp \
  --spec-draft-n-max 3 \
  --spec-draft-ngl 999 \
  --spec-draft-type-k q8_0 \
  --spec-draft-type-v q8_0 \
  -mm mmproj-SuperQwen3.8-27b-abliterated-Q8_0.gguf \
  -ngl 999 \
  -c 65536 \
  -ctk q8_0 \
  -ctv q8_0 \
  -fa on \
  -np 1 \
  --host 0.0.0.0 \
  --port 8891 \
  -a SuperQwen3.8-27b-abliterated-GGUF \
  --jinja \
  --reasoning-format deepseek

Use --spec-type none and omit -md for the non-speculative baseline. Reduce -c if
you prefer a smaller KV allocation, or increase it only after measuring memory and
retrieval behavior for your workload.

Measured C1 performance

The non-speculative baseline used llama-bench with p256, n512, three repetitions,
Q8_0 KV, Flash Attention, and all layers on the GB10 GPU:

Trial C1 decode
1 12.5248 tok/s
2 12.4931 tok/s
3 12.4685 tok/s
Mean 12.4955 tok/s

Mean p256 prompt processing was 805.15 tok/s.

The MTP search used the same model, GPU, Q8_0 KV, deterministic prompt, and fixed
256-token generation:

Profile Decode
No speculation 12.3 tok/s
MTP K=1 19.1 tok/s
MTP K=3 21.0 tok/s
MTP K=5 20.2 tok/s
MTP K=8 19.2 tok/s

The selected K=3 profile was then repeated with fixed 512-token generation:

Trial C1 decode
1 22.2 tok/s
2 22.0 tok/s
3 22.1 tok/s
Median 22.1 tok/s

Native MTP speculation

Across the OpenAI-compatible functional suite, K=3 generated 606 draft tokens and the
target accepted 375: 61.88% aggregate acceptance. On the fixed 256-token throughput
request, 160 of 282 drafts were accepted and server-side decode measured 25.36 tok/s;
including request overhead, the client observed 23.535 tok/s.

incoai/Qwen3.8-27B-DFlash2-GGUF
was checked rather than advertised on assumption. The exact 1,143,006,752-byte Q4_K_M
file (18a380…0594) failed direct draft loading on the pinned llama.cpp build with
expected 81 tensors, got 58. DFlash is therefore not enabled in this release.

Functional release gates

Gate Result
/v1/models and ordinary chat PASS
Arithmetic reasoning (37 × 41 = 1517) PASS
Forced lookup_weather function call with valid JSON arguments PASS
Benign-sensitive Ubuntu owner-recovery guidance PASS
Bounded one-line reasoning (17 + 25 = 42) PASS
Vision: identify the llama logo PASS
Fixed 256-token API generation PASS

The BF16 source release carries the broader refusal, capability, tool, vision, and
36-case overthinking evidence. These GGUF gates independently confirm that the converted
runtime still exercises each critical path.

Verified long context

A 46,235-token prompt placed ORCHID-7291 in the middle of neutral filler. The K=3
server returned the key exactly. Prompt processing measured 689.05 tok/s and the full
request completed in 67.58 seconds. The GGUF retains the model's 262,144-token training
metadata, but this card promotes only the context length directly exercised here.

Other formats

Limitations

  • Q4_K_M can regress tasks outside the measured gates; use BF16 when maximum fidelity
    matters more than memory and speed.
  • Abliteration reduces a measured refusal direction. It does not make every answer
    correct, harmless, or suitable for every deployment.
  • MTP gains depend on prompt distribution and acceptance. Re-measure your own workload.
  • The 46K result is a retrieval gate, not a claim of perfect recall for every long task.
  • Speed depends on hardware, llama.cpp revision, KV precision, context, and sampling.

Evidence identities

Evidence SHA-256
C1 p256/n512 llama-bench JSON a0631cc5e7d211225faeba25bd43264ad3cfc64da50dabaf45d012be299c610b
OpenAI API functional and long-context gates 1ea93778c2ff5eafa8e1882145b87cf5011ad016700915cc58b389ffddcd5e48
BF16 source-weight hash verification f258ecce502083514b0e865bebbff417692c8338d1c81cbaf5b6fcb27c7e244a

SHA256SUMS.json covers every published artifact and release-support file. The private
repository, public main, and v1.0.0 tag are independently verified before release.

License

Apache-2.0, following the upstream Qwen3.8 release.

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-25Corrected verified SuperQwen3.8-27b-abliterated-GGUF release9c9ca067.2 KB
    Loading...
  2. 2026-08-24Refresh model card and release navigation9f4d70b10.6 KB
    Loading...
  3. 2026-08-24Publish verified SuperQwen3.8 abliterated GGUF release0d680e49.2 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration