license: apache-2.0
library_name: gguf
pipeline_tag: image-text-to-text
base_model: Jiunsong/SuperQwen3.8-27b-abliterated
base_model_relation: quantized
tags:
- qwen3.8
- qwen3.5
- multimodal
- image-text-to-text
- reasoning
- tool-calling
- long-context
- uncensored
- abliterated
- supertune
- gguf
- llama.cpp
- q4-k-m
- speculative-decoding
- mtp
- dgx-spark
language: - en
- ko
SuperQwen3.8-27b-abliterated-GGUF
The portable one-box edition: imatrix Q4_K_M, native Qwen MTP, and a separate multimodal projector.
One model. Four native releases.
Choose your build
| Release | Best for | Size / precision | Runtime |
|---|---|---|---|
| BF16 | Maximum fidelity and further tuning | ~52 GB · BF16 | Transformers / vLLM |
| NVFP4 | Fast single-DGX-Spark serving | ~19.2 GiB · W4A4 G16 | vLLM |
| GGUF — this repo | Portable one-box inference + native MTP | ~17.6 GiB runtime set | llama.cpp |
| MLX 4-bit | Apple Silicon | ~15.0 GiB · affine 4-bit | MLX |
[!NOTE]
An expanded post-release Obliteratus and cross-format audit is in progress. Existing
benchmark claims remain bound to the hashed evidence listed below and will be replaced,
not extrapolated, when the wider independent suite completes.
This is the llama.cpp-compatible release ofJiunsong/SuperQwen3.8-27b-abliterated,
pinned to BF16 revision 84efc06504219387b308f15561ddfc8a2b880966 and ultimately toQwen/Qwen3.8-27B@1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0.
It keeps text generation, the model's native MTP draft head, and the vision projector in
separate files so each runtime component can use the precision that fits its job.
Release highlights
| Verified release value | |
|---|---|
| Hardware | 1× NVIDIA DGX Spark / GB10, one request at a time (C1) |
| Target weights | Importance-matrix-guided Q4_K_M, 15.41 GiB |
| Speculative draft | Native Qwen MTP Q4_0, selected at K=3 |
| Vision | Mixed Q8_0/F16/F32 mmproj, directly tested with an image |
| C1 decode without speculation | 12.4955 tok/s mean, p256/n512, 3 trials |
| C1 decode with MTP K=3 | 22.1 tok/s median, n512, 3 trials |
| OpenAI API fixed generation | 23.535 tok/s end to end, 256/256 completion tokens |
| Functional gates | Chat, reasoning, tool call, benign-sensitive, bounded reasoning, vision: PASS |
| Long-context gate | 46,235 prompt tokens, hidden needle retrieved exactly |
Why this release
- Runs on one Spark. The three runtime artifacts total about 17.56 GiB and the
verified server reported 20,760 MiB of GPU memory at a 65,536-token context allocation. - C1 is the headline. The promoted number is a single interactive request, not a
concurrent aggregate relabeled as single-user speed. - Speculation was searched, not guessed. K=1, 3, 5, and 8 were measured; K=3 was
fastest and then repeated three times at 512 generated tokens. - Multimodal remains real. The separate projector identified the llama in the
release image gate, and the model emitted a valid forced function call. - Reasoning terminates. The embedded bounded template returned the one-line answer
42in three tokens, while a separate reasoning gate correctly produced1517.
Files and integrity
| File | Size | SHA-256 | Role |
|---|---|---|---|
SuperQwen3.8-27b-abliterated-Q4_K_M.gguf |
16,547,401,056 B | 8edcf493410cd049abf25c3c5ddd0c28dde0b22d44278bfb899f82683b1e4e43 |
851-tensor target |
mtp-SuperQwen3.8-27b-abliterated-Q4_0.gguf |
1,680,272,384 B | 3f346731cb86f664351b41fde590b9cecd4f66473a73e8e81bb33fdc2d918392 |
18-tensor MTP draft |
mmproj-SuperQwen3.8-27b-abliterated-Q8_0.gguf |
629,247,680 B | ead8e9f2839910775e3f871109a6a84392fc7c9c8e00e34997bb3819c18bbfe8 |
334-tensor vision projector |
The target uses the standard Q4_K_M mix: 433 Q4_K tensors, 65 Q6_K tensors, and 353
small or structurally sensitive F32 tensors. The MTP file contains 10 Q4_0 tensors and
8 F32 tensors. The mmproj contains 83 Q8_0, 27 F16, and 224 F32 tensors; unsupported
4304-column vision matrices are deliberately preserved instead of being force-packed.
Target quantization used the Unsloth Qwen3.8 importance matrix with SHA-2560ee5b10bd0c2fa2127c6f4b43dbfe1efd71e383b63217af9dade1de36599f1c1.
Conversion and validation used llama.cpp revisionb3c3b96a139d4ef1bdec926ac17aa040981cfc5d compiled for CUDA architecture 121a.
Serving on one DGX Spark
llama-server \
-m SuperQwen3.8-27b-abliterated-Q4_K_M.gguf \
-md mtp-SuperQwen3.8-27b-abliterated-Q4_0.gguf \
--spec-type draft-mtp \
--spec-draft-n-max 3 \
--spec-draft-ngl 999 \
--spec-draft-type-k q8_0 \
--spec-draft-type-v q8_0 \
-mm mmproj-SuperQwen3.8-27b-abliterated-Q8_0.gguf \
-ngl 999 \
-c 65536 \
-ctk q8_0 \
-ctv q8_0 \
-fa on \
-np 1 \
--host 0.0.0.0 \
--port 8891 \
-a SuperQwen3.8-27b-abliterated-GGUF \
--jinja \
--reasoning-format deepseek
Use --spec-type none and omit -md for the non-speculative baseline. Reduce -c if
you prefer a smaller KV allocation, or increase it only after measuring memory and
retrieval behavior for your workload.
Measured C1 performance
The non-speculative baseline used llama-bench with p256, n512, three repetitions,
Q8_0 KV, Flash Attention, and all layers on the GB10 GPU:
| Trial | C1 decode |
|---|---|
| 1 | 12.5248 tok/s |
| 2 | 12.4931 tok/s |
| 3 | 12.4685 tok/s |
| Mean | 12.4955 tok/s |
Mean p256 prompt processing was 805.15 tok/s.
The MTP search used the same model, GPU, Q8_0 KV, deterministic prompt, and fixed
256-token generation:
| Profile | Decode |
|---|---|
| No speculation | 12.3 tok/s |
| MTP K=1 | 19.1 tok/s |
| MTP K=3 | 21.0 tok/s |
| MTP K=5 | 20.2 tok/s |
| MTP K=8 | 19.2 tok/s |
The selected K=3 profile was then repeated with fixed 512-token generation:
| Trial | C1 decode |
|---|---|
| 1 | 22.2 tok/s |
| 2 | 22.0 tok/s |
| 3 | 22.1 tok/s |
| Median | 22.1 tok/s |
Native MTP speculation
Across the OpenAI-compatible functional suite, K=3 generated 606 draft tokens and the
target accepted 375: 61.88% aggregate acceptance. On the fixed 256-token throughput
request, 160 of 282 drafts were accepted and server-side decode measured 25.36 tok/s;
including request overhead, the client observed 23.535 tok/s.
incoai/Qwen3.8-27B-DFlash2-GGUF
was checked rather than advertised on assumption. The exact 1,143,006,752-byte Q4_K_M
file (18a380…0594) failed direct draft loading on the pinned llama.cpp build withexpected 81 tensors, got 58. DFlash is therefore not enabled in this release.
Functional release gates
| Gate | Result |
|---|---|
/v1/models and ordinary chat |
PASS |
Arithmetic reasoning (37 × 41 = 1517) |
PASS |
Forced lookup_weather function call with valid JSON arguments |
PASS |
| Benign-sensitive Ubuntu owner-recovery guidance | PASS |
Bounded one-line reasoning (17 + 25 = 42) |
PASS |
| Vision: identify the llama logo | PASS |
| Fixed 256-token API generation | PASS |
The BF16 source release carries the broader refusal, capability, tool, vision, and
36-case overthinking evidence. These GGUF gates independently confirm that the converted
runtime still exercises each critical path.
Verified long context
A 46,235-token prompt placed ORCHID-7291 in the middle of neutral filler. The K=3
server returned the key exactly. Prompt processing measured 689.05 tok/s and the full
request completed in 67.58 seconds. The GGUF retains the model's 262,144-token training
metadata, but this card promotes only the context length directly exercised here.
Other formats
- Original BF16 weights:
Jiunsong/SuperQwen3.8-27b-abliterated - Single-DGX-Spark NVFP4:
Jiunsong/SuperQwen3.8-27b-abliterated-NVFP4-DGX-Spark - Apple Silicon MLX 4-bit:
Jiunsong/SuperQwen3.8-27b-abliterated-MLX-4bit
Limitations
- Q4_K_M can regress tasks outside the measured gates; use BF16 when maximum fidelity
matters more than memory and speed. - Abliteration reduces a measured refusal direction. It does not make every answer
correct, harmless, or suitable for every deployment. - MTP gains depend on prompt distribution and acceptance. Re-measure your own workload.
- The 46K result is a retrieval gate, not a claim of perfect recall for every long task.
- Speed depends on hardware, llama.cpp revision, KV precision, context, and sampling.
Evidence identities
| Evidence | SHA-256 |
|---|---|
| C1 p256/n512 llama-bench JSON | a0631cc5e7d211225faeba25bd43264ad3cfc64da50dabaf45d012be299c610b |
| OpenAI API functional and long-context gates | 1ea93778c2ff5eafa8e1882145b87cf5011ad016700915cc58b389ffddcd5e48 |
| BF16 source-weight hash verification | f258ecce502083514b0e865bebbff417692c8338d1c81cbaf5b6fcb27c7e244a |
SHA256SUMS.json covers every published artifact and release-support file. The private
repository, public main, and v1.0.0 tag are independently verified before release.
License
Apache-2.0, following the upstream Qwen3.8 release.