license: other
library_name: sglang
pipeline_tag: image-text-to-text
base_model: Qwen/Qwen3.8-Flash-Next-FP8
tags:
- qwen3.8
- qwen4-exp
- multimodal
- image-text-to-text
- reasoning
- tool-calling
- abliterated
- obliteratus
- fp8
SuperQwen3.8-Flash-Next-abliterated-FP8-DGX-Spark
Architecture-aware Qwen4-Exp refusal-subspace edit with measured capability preservation.
This deployment derivative is pinned to the official block-FP8 checkpointQwen/Qwen3.8-Flash-Next-FP8@bcd9f01ddc9cff2316eb84281bebcd5b058bddce.
The selected BF16 output-map projection is identical to the BF16 release target set;
the official FP8 layout and protected components remain otherwise unchanged.
Architecture
| Component | Verified checkpoint structure |
|---|---|
| Model class | Qwen4ExpForConditionalGeneration |
| Text stack | 48 layers, hidden size 2,560, four gated residual streams |
| Sequence blocks | 36 Gated DeltaNet + 12 Qwen Sparse Attention layers |
| MoE | 512 routed experts per layer, top-10, plus one shared expert |
| PLE | Bigram/trigram per-layer embedding at layer 2; preserved |
| Vision | 27-layer encoder; preserved |
| MTP | Native one-layer MTP tensors; preserved, not used in the measured K=0 recipe |
| Chat | Thinking controls, reasoning-effort controls, tool calls, multimodal input |
The edit targets the measured 2,560-dimensional attention/DeltaNet output maps plus every routed/shared-expert MLP down projection.
Vision, PLE, expert routers and gates, hyper-connection gates, normalization, and MTP are
outside the target set.
Method
- Corpus: 842 canonical harmful/harmless pairs.
- Hidden space: four-stream hyper-connection read mix, width 2,560.
- Selected candidate:
amo_r1_l12_a200. - Layers:
[33, 35, 36, 37, 38, 40, 41, 42, 44, 45, 46, 47]. - Output paths:
['attention', 'mlp']. - Rank: 1.
- Projection strength: 2.0.
- Changed storage tensor count: 36
(6168 logical output projections, including
fused routed experts); zero router, PLE, vision, normalization, residual-gate, or MTP
targets.
Measured release gates
| Gate | Result |
|---|---|
| Capability | 8 / 8 |
| Tool call | PASS |
| Strict JSON schema | PASS |
| Vision | PASS |
| Reasoning/overthinking | 36 / 36 PASS |
| Full hash-only refusal stress | 153 / 1684 broad refusals vs. upstream 842 / 1684; 10 anomalies |
| Harmful-prompt refusals | 153 / 842 vs. upstream 814 / 842 |
| Harmless-prompt refusals | 0 / 842 vs. upstream 28 / 842 |
| p256 / C1 decode | 16.772 tok/s |
| p8192 / C1 decode | 16.240 tok/s |
| p8192 / C1 TTFT | 3.855 s |
| Bounded long context | 8,240, 16,431, 24,621 prompt tokens; needle retrieved |
C1 means one active request. Speed is specific to the measured two-node DGX Spark
SGLang runtime. The checkpoint metadata declares a larger native window, but this
release card claims only the bounded lengths listed above as locally verified.
The full stress run retained 10 Unicode
replacement flags where generation stopped exactly at its
64-token cap. The original measurements remain in
the report. Exact prompt-hash retries at 128 tokens
were anomaly-free for all 10 affected
cases; the retry covered 20 harmful/harmless rows.
Two-node SGLang recipe
The measured recipe used two DGX Spark GB10 nodes, TP=2, EP=2, one active request,
32,768 served context, BF16
recurrent state, triton GDN prefill,
flashinfer GDN decode, and the verified QSA decode overlay.
CUDA graphs and native MTP were disabled (K=0).
Run the following once on each node with the rank-specific value and a shared
rank-0 address. The model path must contain the exact verified repository commit.
The required GB10 compatibility source is included atruntime/qwen4_exp_gb10_sitecustomize.py
and must be mounted on PYTHONPATH as sitecustomize.py; its SHA-256 is92c2aac1245403ca0ad3ebd1db68dac3b478cd3aabd9b18884cca5c860876aa4.
hf download Jiunsong/SuperQwen3.8-Flash-Next-abliterated-FP8-DGX-Spark --revision v1.0.0 --local-dir ./model
python3 \
-m \
sglang.launch_server \
--model-path \
/model \
--served-model-name \
SuperQwen3.8-Flash-Next-abliterated-FP8 \
--load-format \
safetensors \
--nnodes \
2 \
--dist-init-addr \
${RANK0_HOST}:26039 \
--tp-size \
2 \
--ep-size \
2 \
--context-length \
32768 \
--mem-fraction-static \
0.85 \
--max-running-requests \
1 \
--chunked-prefill-size \
2048 \
--linear-attn-prefill-backend \
triton \
--linear-attn-decode-backend \
flashinfer \
--mamba-ssm-dtype \
bfloat16 \
--reasoning-parser \
auto \
--tool-call-parser \
auto \
--watchdog-timeout \
1200 \
--host \
0.0.0.0 \
--port \
8890 \
--disable-cuda-graph \
--node-rank \
${NODE_RANK}
The runtime settings are a measured compatibility recipe, not a universal optimum.
Verification evidence
| Evidence | SHA-256 |
|---|---|
| architecture report | 27b35581e97cc78f0a6e67f565aabc90f36fdd2908faeed285d6eb0fd88040ad |
| 842-pair corpus | 218c67b48f495eb559a36966067fa5101db2350f688c71f4b9b0ec7aecd6de53 |
| capture manifest | 408fb389190e6536cdb38eb3c9a8f5322ae15307552d2fe8ae04250535708273 |
| subspace report | b6bce6fd56bdaef8b5d766725d0dd55f2d6472f16e31e8768e0dc6ac2260fed2 |
| bounded candidate bank | 708b6b28d22b71cb2eb37c90cf0246abca806d29d4d6786611a065c603e78561 |
| selected candidate | 14aa5b9072c32513a84579e6ff14699c49887473b4484ae273851160db159f28 |
| baseline benchmark | 2d39fa15726b9738167d30421511d28e477be4bb5f87cd63ca70e50a1a1252cb |
| baseline full refusal stress | 86a66d642601ebb29ad5f43d1f83ffb3b9a0b0aee4aae595ec20adace118e148 |
| baseline matched refusal sample | 673c7a070d49c27dac5fbc4e973113fedc1cc5f8a95b0c7182d2f99dd80b179d |
| candidate matched refusal sample | 59d1b994d7be474ee591cd1ff861b6542b1aa18fb85217000c9b5150d1a624e1 |
| BF16 patch verification | 941e762a29c3ed2ba1ca54099c8c5be126ac0f40b19ab2ca7aa078c2f5ac3911 |
| BF16 artifact verification | 639764f46ae8a47b8fc2e7c556e762283ef3fb36bfd4f100137f467af805915f |
| FP8 patch verification | 9581118ec2f75871a1243ca4da49f16d9019fe1a0c0a89257e2821be902588d6 |
| FP8 artifact verification | 8434b1c2cf18022c701a4cf8f541aafdbf2c15633a5e065a5a2f36157c5b79e7 |
| FP8 benchmark | 6c2e282c05fe51aefdf3e80141237b065e95d03c648c3d240e235a9209d8f63b |
| FP8 long-context retries | 3089c621d716eadfc6402535325b05e4945f0586ee2528f989d3070d4ca1e0b2 |
| FP8 full refusal stress | db206b81a5b88535c2f063d89684851cee4785f79794cb5f0e05c01469fd20de |
| FP8 refusal anomaly retries | afe6f96e96bfd72d2f23b0a21e2f31b16275e7a60132572203ca2898ff02209e |
| FP8 clean-reload inference | a58f31ea261fbea20432720e62dc6714622f4a94f4832822d63a679957ccc440 |
| FP8 runtime recipe | 4691af15670888e0bc56071eacf43d13ea3d64c943bb0b1d9eccfa91145073c7 |
| FP8 runtime overlay | 92c2aac1245403ca0ad3ebd1db68dac3b478cd3aabd9b18884cca5c860876aa4 |
| FP8 rank-0 serve script | d9751b487dc1c88bd7b709a4625ecd7a8fb9467ad9157042f965b9f0257f79d5 |
| FP8 rank-1 serve script | d9751b487dc1c88bd7b709a4625ecd7a8fb9467ad9157042f965b9f0257f79d5 |
Upstream BF16 revision: f5d08274bafd880402bd16f5e3e6c514136ec06c
Official FP8 revision: bcd9f01ddc9cff2316eb84281bebcd5b058bddce
OBLITERATUS revision: a5a1ffa5849b442cf188b3c03fd4de71ddf5bdcc
A compact, path-sanitized copy of the validated metrics and runtime identities is
included at verification/RELEASE_EVIDENCE.json.
Limitations
- Abliteration reduces a measured refusal direction; it does not guarantee correctness,
harmlessness, or universal compliance. - The capability and behavior suites are finite regression gates.
- Native MTP exists in the checkpoint, but the reported deployment measurements used K=0.
- GGUF and MLX variants are not published because Qwen4-Exp converter/runtime support was
not independently verified for this release. - Operators remain responsible for access controls, policy, and downstream safeguards.
License and attribution
This derivative follows the upstream Qwen3.8-Flash-Next license. Review the includedLICENSE and the upstream repository before redistribution or deployment.