license: apache-2.0
library_name: transformers
pipeline_tag: image-text-to-text
base_model: Qwen/Qwen3.8-27B
base_model_relation: finetune
tags:
- qwen3.8
- qwen3.5
- multimodal
- image-text-to-text
- reasoning
- tool-calling
- long-context
- uncensored
- abliterated
- supertune
- vllm
- nvfp4
- w4a4
- quantized
- dgx-spark
- speculative-decoding
language: - en
- ko
SuperQwen3.8-27b-abliterated-NVFP4-DGX-Spark
A single-DGX-Spark release tuned for fast C1 decode without trading away the model's multimodal, tool, reasoning, or long-context behavior.
This is the one-box performance edition ofJiunsong/SuperQwen3.8-27b-abliterated.
It fits and serves on one NVIDIA DGX Spark with TP=1. The release combines true
compressed-tensors NVFP4 W4A4, group size 16 with the model's own MTP draft head at
K=5. It does not need a second Spark for the measured result.
Release highlights
| Verified release value | |
|---|---|
| Hardware | 1× NVIDIA DGX Spark / GB10, TP=1 |
| C1 decode | 27.1270 tok/s median at p256; trials 26.6789 / 27.1270 / 27.3181 |
| Speculative decoding | Built-in Qwen MTP, K=5, TRITON draft attention |
| Checkpoint format | NVFP4 W4A4 G16; 5 packed shards + protected BF16 MTP shard; about 19.2 GiB |
| Serving memory profile | 0.82 GPU utilization target, FP8 KV cache, max 262,144 tokens |
| Quality gates | Capability 7/8 (paired-parent floor), tool PASS, vision PASS |
| Behavior gate | Benign-sensitive refusal 0/8; no forced refusal in the final suite |
| Bounded reasoning | 36/36 PASS across default, low, medium, and xhigh |
| Native context | 250,046 actual prompt tokens, hidden needle retrieved, K=5 |
Why this release
- Fast single-stream decode on one Spark. The number above is C1, not a concurrent
aggregate presented as single-user speed. - Quality-selected speculation. K=5 is the fastest candidate that completed the
full paired capability, tool, vision, refusal, and 36-case reasoning release gate. - Quantized where it pays, protected where it matters. Vision, MTP, conv1d, and
lm_headremain on their verified protected paths instead of being blindly packed. - Reasoning that terminates. The bounded template fixes the common pattern where a
correct answer is reconsidered, repeated, or talked out of existence. - Still multimodal and tool-capable. This is an image-text-to-text checkpoint, not a
text-only repack.
Measured C1 performance
The benchmark follows the post-first-token decode contract used byMiaAI-Lab/sparkDash at commitbf2709a80ef25d0e1a6ee41efec4c9b8042a5b8b:
- one request at a time (C1)
- 256-token target prompt class; 292 prompt tokens after chat formatting
- fixed 512-token generation
- thinking disabled for the throughput lane
- decode window from first visible reasoning/content token to the last
| Trial | C1 decode | TTFT |
|---|---|---|
| 1 | 26.6789 tok/s | 0.460 s |
| 2 | 27.1270 tok/s | 0.464 s |
| 3 | 27.3181 tok/s | 0.458 s |
| Median | 27.1270 tok/s | 0.460 s |
Only C1 is promoted here because it matches the intended interactive, single-user Spark
deployment. Concurrent aggregate numbers are deliberately not used as the headline.
Quantization and integrity
| Component | Precision / treatment |
|---|---|
| Eligible transformer linear weights and activations | NVFP4 W4A4, group size 16 |
| Vision tower | Protected, exact |
| MTP draft head | Protected BF16 shard, exact |
conv1d paths |
Protected, exact |
lm_head |
BF16, exact |
| Serving KV cache | FP8 in the measured profile |
The provenance chain is pinned toQwen/Qwen3.8-27B@1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0.
The packed artifact bytes are inherited from the fully verifiedNVFP4-2xDGX
release; this repository changes the single-node launch profile, bounded-reasoning
template, and release evidence, not the packed model tensors.
Why K=5 MTP
The native Qwen MTP head shares the target model's tokenizer and architecture. K=5 was
selected after K-depth measurement and then independently subjected to the complete
release suite. The final server reported healthy accepted drafts during both ordinary
generation and the 250K retrieval gate.
The incoai/Qwen3.8-27B-DFlash2
checkpoint was also tested rather than assumed to work. On the measured vLLM backport,
both the SuperQwen and official NVFP4 A/B runs drafted tokens but accepted 0. The
SuperQwen DFlash lane measured only 8.1125 tok/s C1. MLX-LM 0.31.3 also could not
load it as a conventional causal draft. DFlash 2 is therefore not enabled or advertised
as an acceleration path in this release.
Bounded reasoning
The installed template is revision 2:
- unspecified effort defaults to bounded
medium - explicit
lowandxhighboth carry repeat/restart stop guards - the template embedded in
tokenizer_config.jsonexactly matcheschat_template.jinja
All nine deterministic cases passed at every effort level:
| Effort | Result |
|---|---|
| default | 9/9 |
| low | 9/9 |
| medium | 9/9 |
| xhigh | 9/9 |
| Total | 36/36 |
Explicit controls remain available through chat_template_kwargs, for example:
extra_body={"chat_template_kwargs": {"enable_thinking": True, "reasoning_effort": "xhigh"}}
Verified native context
At K=5 and the official 262,144-token native limit, a 250,046-token prompt completed
end to end and returned the hidden needle exactly. The run took 803.03 seconds. This is a
real retrieval gate, not a tokenizer-only context claim. It does not imply perfect recall
on every possible 250K task.
Serving on one DGX Spark
The bundled launcher defaults to TP=1, K=5 MTP, FP8 KV, TRITON attention, eager mode,
and the 262,144-token native window:
bash repro/scripts/serve_superqwen38_single_dgx.sh \
/path/to/SuperQwen3.8-27b-abliterated-NVFP4-DGX-Spark
The exact benchmark runtime was vLLM0.25.2.dev0+g752a3a504.d20260714 on the DGX Spark image lineageghcr.io/anemll/dspark-vllm-gx10:0.1.1, with Qwen3.8 support built from vLLM revision3406ec1dae9916f920b90f0dbf90dcf54923d042. Override the image explicitly when needed:
QWEN38_VLLM_IMAGE=your-compatible-vllm-image \
bash repro/scripts/serve_superqwen38_single_dgx.sh /path/to/model
The OpenAI-compatible API is exposed on port 8888 by default. The served model ID isSuperQwen3.8-27b-abliterated-NVFP4-DGX-Spark.
Other formats
- Original BF16 weights:
Jiunsong/SuperQwen3.8-27b-abliterated - Apple Silicon MLX 4-bit:
Jiunsong/SuperQwen3.8-27b-abliterated-MLX-4bit
— measured 29.960 tok/s C1 median on the release Mac - Two independent DGX replicas:
Jiunsong/SuperQwen3.8-27b-abliterated-NVFP4-2xDGX
Limitations
- NVFP4 compression can regress workloads outside the measured suites.
- Abliteration reduces a measured refusal direction; it does not make every answer
correct, harmless, or appropriate for every deployment. - The DFlash 2 result is specific to the tested checkpoint and runtime; it may change in
a future implementation with verified non-zero acceptance. - Speed is hardware, prompt, runtime, and sampling dependent. Re-measure your workload.
- The 250K result is a retrieval check and should not be read as universal 250K accuracy.
Evidence identities
| Evidence | SHA-256 |
|---|---|
| C1 + capability/tool/vision/refusal/overthinking release gate | 497e8c054ad8ee8f3a42a966ec4adc2c0a3ca7f63778e6dfc4d9d13cf9316457 |
| Native 250K retrieval gate | 64d962a633782d20f749e74af03753bb28d3d65c3ba6de6a5757994c0d2ae5f8 |
| Bounded template v2 | f8035177dc3ffcccb94281f180247a0ca1f0bffc56080b61a2f31f1304ae5cd3 |
| Bounded-template installation report | 10040d14cdd5bafded7de93f520bb5f223f823655033ef05f2d3dbc19609898a |
| DFlash 2 compatibility decision | 1dbc0b2ffa67595ead86e6f7fd00c5479b60520431c105956774023468d33e85 |
| DFlash 2 SuperQwen C1 run | cea99b86ab35dd3654ef27e760bf0583a563cc55e9f73523335291b25db7a1e2 |
| DFlash 2 official-NVFP4 post-run metrics | a58237f09ed133217e4a9fc28ea71d3503f3a58285a7adb2c3bcb7a3d2896054 |
Every repository file, including the packed weight shards inherited from the source
release, is covered by SHA256SUMS.json and is verified once while private, again after
public visibility, and again at the v1.0.0 tag.
License
Apache-2.0, following the upstream Qwen3.8 release.