license: other
license_name: swift-open-license-1.0
license_link: LICENSE
base_model: jessedye90/Swift-1.5-Qwen3.8-Flash-Next-W4A16-GB10
base_model_relation: finetune
pipeline_tag: image-text-to-text
library_name: transformers
tags:
- abliterated
- uncensored
- qwen3_8
- qwen4_exp
- moe
- w4a16
- gptq
- fp8
- mtp
- vllm
- dgx-spark
- gb10
extra_gated_prompt: >-
This model has had its safety refusals removed (abliteration). It will comply with harmful requests that the
original model refuses. It is published for research into refusal mechanisms, alignment and red-teaming. You are
responsible for how you use it and for everything it generates, and you must comply with the Swift Open License
v1.0, the Qwen Community License 1.0 and applicable law.
qwen3.8-flash-next-swift-uncensored
jessedye90/Swift-1.5-Qwen3.8-Flash-Next-W4A16-GB10
(UkisAI's Swift 1.5 fine-tune of Qwen3.8-Flash-Next, INT4 W4A16, packaged for one DGX Spark) with OrcaRouter's
refusal direction projected out of its weights. The edit was made inside the checkpoint's own INT4 and FP8 grids,
so the file layout, size, serving recipe and speed are the same as the base model's.
Safety alignment is removed. This model will comply with harmful, unethical or illegal requests that the base
model refuses. It is released for research (refusal mechanisms, interpretability, red-teaming, robustness
evaluation). Do not deploy it to end users without your own safety layer. You alone are responsible for its use
and outputs.
Not self-contained, same as the base: the 51B-parameter n-gram embedding (PLE) table is served from
Saren/Qwen3.8-Flash-Next-ple-table-fp8
(revision50511b0a41aa1d34b8beb7e5d4bb06a0b650dc14). The abliteration does not touch that table.
How it was made
- Direction.
orcarouter/Qwen3.8-Flash-Next-Uncensored
is an Arditi-style abliteration ofQwen/Qwen3.8-Flash-Next.
For every residual-writing tensor W, OrcaRouter's weight is W − r rᵀW, so (original − abliterated) is rank one.
Its top singular vector is the refusal direction r. Across all 101 tensors compared (attention and DeltaNet output
projections, shared and MTP experts,ple.value_proj,embed_tokens), the diff is rank one (≥ 99.3 % of its
energy), the scale is 1.000 ± 0.001, and every tensor uses the same r (|cos| ≥ 0.99997). The unit vector isabliteration/refusal_direction.safetensors. - Transfer. Swift 1.5 keeps the base model's refusal component. Its weights carry the same share along r as
Qwen's (2.04 % vs 2.11 % ono_proj), and the per-row components agree at cos 0.92–0.95. - Edit inside the quantization. 25,186 tensors in 147 residual writers: 13
o_proj, 36linear_attn.out_proj
and 24,577 routeddown_proj(all GPTQ INT4); 48 FP8-blockshared_expert.down_proj; and 512 BF16 MTP experts.
Plain dequantize → project → re-round leaves 92–98 % of the refusal component, because the edit (~2 % of the
weight) is smaller than one INT4 step. Each row therefore starts from round-to-nearest and flips the cheapest
near-tie elements to the other bracketing grid point until the row's component along r is cancelled. The
residual left is < 1 % (per tensor inabliteration-report.json).
All scales, zero-points and other tensors are byte-identical to the base. - One deliberate difference from OrcaRouter.
embed_tokensandple.value_projare not edited. With
them edited, the model failed our strict input-validation coding task (0/8 in four runs, base 8/8). A control
with the same rounding noise along a random direction scored 8/8 twice. Leaving the token-embedding path alone
restored the task and still removed refusals (results below).
Scripts: abliteration/. They cover range-fetching tensors, direction recovery, the quantized edit, and the
refusal evaluation.
Measured results
Measured 2026-10-04 on one DGX Spark (GB10) with the UltraFast vLLM recipe, the same build as the base model's
production config. The control is the unmodified base model, measured in the same session.
| base (Swift 1.5 GB10) | this model | |
|---|---|---|
| AdvBench harmful prompts refused (100, greedy, thinking off) | 100 % | 0 % |
| XSTest safe prompts refused (100) | 5 % | 0 % |
| Quality set without code_gen (bug-find, reasoning, JSON, SQL, tool call, 31.7k needle) | 12/12 (earlier runs) | 12/12 |
| code_gen (strict ISO-8601 parser, 3 runs) | 8, 8, 8 / 8 | 8, 8, 6 / 8 |
| LiveCodeBench v6 sample (6 problems) | 6/6 | 6/6 |
| Single-stream decode, 256 tokens | 55 tok/s (base card) | 56 tok/s |
Claude Code and Codex CLI sessions with tool use ran correctly against it.
Limits: small samples. Refusal is classified by refusal phrases in the first 400 characters. AIME, long-context
needles beyond 31.7k and the 524k window were not re-measured on this build. Abliteration can shift other
behaviours that these tests do not cover.
Running it
Exactly as the base model (its card):
the dime-online/qwen3.8-Flash-DGX-UltraFast recipe
at commit 0c391a3, with serving/swift-prod/env (524k) or serving/swift-va/env (262k). Point MODEL_DIR at
this repo and TABLE_DIR at the PLE table.
License
Distributed under the same terms as the base model:
- UkisAI's Swift Contribution is under the Swift Open License v1.0 (
LICENSE), including its
commercial-use limitation. Commercial use by an entity with US$1,000,000 or more in annual gross revenue needs a
separate Swift Enterprise License from UkisAI (ukisai.com/contact). - The Base Model, Qwen3.8-Flash-Next (Copyright (c) 2026 Qwen), is under the Qwen Community License 1.0
(LICENSE-QWEN). Its terms apply. NOTICEcarries UkisAI's attribution and the change notices, including the one for this abliteration.
"UkisAI", "Swift" and "Qwen" are used only to say where this model comes from. This release is not made or
endorsed by UkisAI, Qwen, OrcaRouter, Intel, NVIDIA or the authors of the recipe and tools above.