← back to catalog · registered 2026-08-24 15:02

Jiunsong/SuperQwen3.8-27b-abliterated-NVFP4-DGX-Spark

Jiunsong Qwen 24B multimodal
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/Jiunsong%2FSuperQwen3.8-27b-abliterated-NVFP4-DGX-Spark"
Response includes
  • classification m1
  • files 23
  • hub_downloads_all_time 6,830
  • author_summary 35 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
7K
909 last 30d - stable
Likes
30
Model age
6w ago
created 2026-08-24
Downloads over time
Now7.1K→from845↑738%
5332.9K5.3K7.7K845 on Aug 267.1K on Oct 11AugSepOct
Aug 26 → Oct 11 · 47 snapshots · spans 46 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en ko
Quantizations
BF16
Tags
transformers safetensors qwen3_5 image-text-to-text qwen3.8 qwen3.5 multimodal reasoning tool-calling long-context 262k-context uncensored

Related

Total size
19.1 GB
Files
23
Quantizations
2
Registered
2026-08-24 15:02
Last updated on HF
2026-08-25 02:10

Files by quantization

BF16 1 file 810 MB
model-mtp-bf16.safetensors 810 MB 2a83d7c1 download
Auxiliary files 22 files 18.4 GB
model-00004-of-00005.safetensors 4.66 GB d0deb4f8 download
model-00002-of-00005.safetensors 4.65 GB 64853350 download
model-00003-of-00005.safetensors 4.65 GB 2c5415ac download
model-00001-of-00005.safetensors 2.37 GB df446302 download
model-00005-of-00005.safetensors 2.03 GB f8fc4537 download
tokenizer.json 19.1 MB 06b95093 download
model.safetensors.index.json 272 KB 5c418317 download
abliteration_config.json 48.7 KB 9621fc9f download
tokenizer_config.json 16.1 KB 75525de3 download
config.json 15.3 KB 851f2840 download
LICENSE 11.3 KB f938136e download
SHA256SUMS.json 9.91 KB bb78df57 download
README.md 9.45 KB 0fd1f6e8 download
chat_template.jinja 8.88 KB 2f59b5d5 download
abliteration_verification.json 6.84 KB a9ea67b1 download
quantization_verification.json 3.82 KB 0de4d0bd download
quantization_provenance.json 2.88 KB 9a548b7a download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.16 KB 33818c7f download
overthinking_template_report.json 640 B da11047b download
recipe.yaml 256 B 3bbc8398 download
generation_config.json 214 B a7aa1d4c download

README current version from Hugging Face


license: apache-2.0
library_name: transformers
pipeline_tag: image-text-to-text
base_model: Qwen/Qwen3.8-27B
base_model_relation: finetune
tags:

  • qwen3.8
  • qwen3.5
  • multimodal
  • image-text-to-text
  • reasoning
  • tool-calling
  • long-context
  • uncensored
  • abliterated
  • supertune
  • vllm
  • nvfp4
  • w4a4
  • quantized
  • dgx-spark
  • speculative-decoding
    language:
  • en
  • ko

SuperQwen3.8-27b-abliterated-NVFP4-DGX-Spark

A single-DGX-Spark release tuned for fast C1 decode without trading away the model's multimodal, tool, reasoning, or long-context behavior.

C1 decode
Precision
Spec decode
Overthinking
Context
License

This is the one-box performance edition of
Jiunsong/SuperQwen3.8-27b-abliterated.
It fits and serves on one NVIDIA DGX Spark with TP=1. The release combines true
compressed-tensors NVFP4 W4A4, group size 16 with the model's own MTP draft head at
K=5. It does not need a second Spark for the measured result.

Release highlights

Verified release value
Hardware 1× NVIDIA DGX Spark / GB10, TP=1
C1 decode 27.1270 tok/s median at p256; trials 26.6789 / 27.1270 / 27.3181
Speculative decoding Built-in Qwen MTP, K=5, TRITON draft attention
Checkpoint format NVFP4 W4A4 G16; 5 packed shards + protected BF16 MTP shard; about 19.2 GiB
Serving memory profile 0.82 GPU utilization target, FP8 KV cache, max 262,144 tokens
Quality gates Capability 7/8 (paired-parent floor), tool PASS, vision PASS
Behavior gate Benign-sensitive refusal 0/8; no forced refusal in the final suite
Bounded reasoning 36/36 PASS across default, low, medium, and xhigh
Native context 250,046 actual prompt tokens, hidden needle retrieved, K=5

Why this release

  • Fast single-stream decode on one Spark. The number above is C1, not a concurrent
    aggregate presented as single-user speed.
  • Quality-selected speculation. K=5 is the fastest candidate that completed the
    full paired capability, tool, vision, refusal, and 36-case reasoning release gate.
  • Quantized where it pays, protected where it matters. Vision, MTP, conv1d, and
    lm_head remain on their verified protected paths instead of being blindly packed.
  • Reasoning that terminates. The bounded template fixes the common pattern where a
    correct answer is reconsidered, repeated, or talked out of existence.
  • Still multimodal and tool-capable. This is an image-text-to-text checkpoint, not a
    text-only repack.

Measured C1 performance

The benchmark follows the post-first-token decode contract used by
MiaAI-Lab/sparkDash at commit
bf2709a80ef25d0e1a6ee41efec4c9b8042a5b8b:

  • one request at a time (C1)
  • 256-token target prompt class; 292 prompt tokens after chat formatting
  • fixed 512-token generation
  • thinking disabled for the throughput lane
  • decode window from first visible reasoning/content token to the last
Trial C1 decode TTFT
1 26.6789 tok/s 0.460 s
2 27.1270 tok/s 0.464 s
3 27.3181 tok/s 0.458 s
Median 27.1270 tok/s 0.460 s

Only C1 is promoted here because it matches the intended interactive, single-user Spark
deployment. Concurrent aggregate numbers are deliberately not used as the headline.

Quantization and integrity

Component Precision / treatment
Eligible transformer linear weights and activations NVFP4 W4A4, group size 16
Vision tower Protected, exact
MTP draft head Protected BF16 shard, exact
conv1d paths Protected, exact
lm_head BF16, exact
Serving KV cache FP8 in the measured profile

The provenance chain is pinned to
Qwen/Qwen3.8-27B@1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0.
The packed artifact bytes are inherited from the fully verified
NVFP4-2xDGX
release; this repository changes the single-node launch profile, bounded-reasoning
template, and release evidence, not the packed model tensors.

Why K=5 MTP

The native Qwen MTP head shares the target model's tokenizer and architecture. K=5 was
selected after K-depth measurement and then independently subjected to the complete
release suite. The final server reported healthy accepted drafts during both ordinary
generation and the 250K retrieval gate.

The incoai/Qwen3.8-27B-DFlash2
checkpoint was also tested rather than assumed to work. On the measured vLLM backport,
both the SuperQwen and official NVFP4 A/B runs drafted tokens but accepted 0. The
SuperQwen DFlash lane measured only 8.1125 tok/s C1. MLX-LM 0.31.3 also could not
load it as a conventional causal draft. DFlash 2 is therefore not enabled or advertised
as an acceleration path in this release.

Bounded reasoning

The installed template is revision 2:

  • unspecified effort defaults to bounded medium
  • explicit low and xhigh both carry repeat/restart stop guards
  • the template embedded in tokenizer_config.json exactly matches chat_template.jinja

All nine deterministic cases passed at every effort level:

Effort Result
default 9/9
low 9/9
medium 9/9
xhigh 9/9
Total 36/36

Explicit controls remain available through chat_template_kwargs, for example:

extra_body={"chat_template_kwargs": {"enable_thinking": True, "reasoning_effort": "xhigh"}}

Verified native context

At K=5 and the official 262,144-token native limit, a 250,046-token prompt completed
end to end and returned the hidden needle exactly. The run took 803.03 seconds. This is a
real retrieval gate, not a tokenizer-only context claim. It does not imply perfect recall
on every possible 250K task.

Serving on one DGX Spark

The bundled launcher defaults to TP=1, K=5 MTP, FP8 KV, TRITON attention, eager mode,
and the 262,144-token native window:

bash repro/scripts/serve_superqwen38_single_dgx.sh \
  /path/to/SuperQwen3.8-27b-abliterated-NVFP4-DGX-Spark

The exact benchmark runtime was vLLM
0.25.2.dev0+g752a3a504.d20260714 on the DGX Spark image lineage
ghcr.io/anemll/dspark-vllm-gx10:0.1.1, with Qwen3.8 support built from vLLM revision
3406ec1dae9916f920b90f0dbf90dcf54923d042. Override the image explicitly when needed:

QWEN38_VLLM_IMAGE=your-compatible-vllm-image \
  bash repro/scripts/serve_superqwen38_single_dgx.sh /path/to/model

The OpenAI-compatible API is exposed on port 8888 by default. The served model ID is
SuperQwen3.8-27b-abliterated-NVFP4-DGX-Spark.

Other formats

Limitations

  • NVFP4 compression can regress workloads outside the measured suites.
  • Abliteration reduces a measured refusal direction; it does not make every answer
    correct, harmless, or appropriate for every deployment.
  • The DFlash 2 result is specific to the tested checkpoint and runtime; it may change in
    a future implementation with verified non-zero acceptance.
  • Speed is hardware, prompt, runtime, and sampling dependent. Re-measure your workload.
  • The 250K result is a retrieval check and should not be read as universal 250K accuracy.

Evidence identities

Evidence SHA-256
C1 + capability/tool/vision/refusal/overthinking release gate 497e8c054ad8ee8f3a42a966ec4adc2c0a3ca7f63778e6dfc4d9d13cf9316457
Native 250K retrieval gate 64d962a633782d20f749e74af03753bb28d3d65c3ba6de6a5757994c0d2ae5f8
Bounded template v2 f8035177dc3ffcccb94281f180247a0ca1f0bffc56080b61a2f31f1304ae5cd3
Bounded-template installation report 10040d14cdd5bafded7de93f520bb5f223f823655033ef05f2d3dbc19609898a
DFlash 2 compatibility decision 1dbc0b2ffa67595ead86e6f7fd00c5479b60520431c105956774023468d33e85
DFlash 2 SuperQwen C1 run cea99b86ab35dd3654ef27e760bf0583a563cc55e9f73523335291b25db7a1e2
DFlash 2 official-NVFP4 post-run metrics a58237f09ed133217e4a9fc28ea71d3503f3a58285a7adb2c3bcb7a3d2896054

Every repository file, including the packed weight shards inherited from the source
release, is covered by SHA256SUMS.json and is verified once while private, again after
public visibility, and again at the v1.0.0 tag.

License

Apache-2.0, following the upstream Qwen3.8 release.

README history 5 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-25Corrected verified SuperQwen3.8-27b-abliterated-NVFP4-DGX-Spark release3138ac66.8 KB
    Loading...
  2. 2026-08-24Refresh model card and release navigation225e22311 KB
    Loading...
  3. 2026-08-24Fix canonical provenance and add GGUF releasedfc55b79.6 KB
    Loading...
  4. 2026-08-24Install verified single-DGX C1 release card and evidencef9c203e9.5 KB
    Loading...
  5. 2026-08-24Duplicate from Jiunsong/SuperQwen3.8-27b-abliterated-NVFP4-2xDGX859ad547.7 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration