license: other
license_name: swift-open-license-1.0
license_link: LICENSE
base_model:
- jessedye90/qwen3.8-27b-swift-uncensored
- ukisai/Swift-1.5-Qwen3.8-27b-NVFP4
base_model_relation: quantized
library_name: llama.cpp
pipeline_tag: text-generation
language: - en
- it
tags: - gguf
- llama.cpp
- abliterated
- uncensored
- qwen3_8
- nvfp4
- fp4
- q8_0
- mtp
- nextn
- speculative-decoding
- dgx-spark
extra_gated_prompt: >-
This model has had its safety refusals removed (abliteration). It will comply
with harmful, unethical or illegal requests that the base model refuses,
including requests in the categories covered by AdvBench. It is published for
research into refusal mechanisms, interpretability, alignment and
red-teaming. You alone are responsible for how you use it and for everything
it generates, and you must comply with the Swift Open License v1.0, the Apache
License 2.0 and applicable law. By requesting access you confirm that you
understand and accept this.
extra_gated_fields:
Full name: text
Country: country
Affiliation: text
I understand that safety refusals have been removed and accept the applicable licenses: checkbox
Swift-1.5-Qwen3.8-27B-NVFP4-Uncensored-MTP (GGUF)
llama.cpp GGUF build ofjessedye90/qwen3.8-27b-swift-uncensored
(UkisAI's Swift 1.5 fine-tune of Qwen3.8-27B, ModelOpt NVFP4/FP8) with its
safety refusals projected out of the weights, plus a working MTP head so it
can be served with llama.cpp's built-in multi-token-prediction speculative
decoding (--spec-type draft-mtp).
⚠️ Safety alignment is removed. This model will comply with harmful,
unethical or illegal requests that the base model refuses. It is released for
research (refusal mechanisms, interpretability, red-teaming, robustness
evaluation). Do not deploy it to end users without your own safety layer.
Not made or endorsed by UkisAI, OrcaRouter, Alibaba Cloud or NVIDIA.
Why this build exists
The abliterated checkpoint ships without an MTP head (its author serves it
with a separate DFlash2 draft model instead). The upstream non-abliterated Swift
1.5 release does carry an MTP head, but taken from the censored base. This
build restores that MTP head and applies the same refusal-direction projection
to its residual writers, so the drafter is consistent with the target, then
converts the result to GGUF while keeping the checkpoints' NVFP4 tensors
byte-identical.
What was changed, relative to the two parents
- MTP head merged — the 15
mtp.*tensors (BF16) were taken fromukisai/Swift-1.5-Qwen3.8-27b-NVFP4
(shardmodel-00005-of-00005.safetensors) and merged into the abliterated
checkpoint. - MTP abliterated —
mtp.layers.0.self_attn.o_projandmtp.layers.0.mlp.down_projhad the same unit refusal direction projected out
of their outputs (exact Arditi projection,W ← W − W r rᵀ). These tensors are
BF16, so the projection is exact — no quantisation grid is involved.
Residual component alongrafter the edit: ~1.2% of the original.mtp.fcis not a residual writer and is unchanged.embed_tokenswas
not edited, matching the base abliterated release. - Converted to GGUF —
convert_hf_to_gguf.py --outtype bf16followed byllama-quantize --tensor-type-file, where the 193 NVFP4 tensors
(64 layers ×ffn_gate/ffn_up/ffn_down+output.weight) are copied
unchanged from the checkpoint (verified byte-identical by SHA-256), the rest
went to Q8_0/F32.
In total 130 residual writers carry the projection: 128 from the abliterated
release (64 FP8 o_proj/out_proj + 64 NVFP4 down_proj) plus the 2 MTP ones.
OrcaRouter's original abliteration covered 131 (it also edited embed_tokens);embed_tokens is deliberately left untouched here — the upstream author found
that editing the token-embedding path broke a strict input-validation coding
task on a sibling model.
Model details
| Architecture | qwen35 (Gated DeltaNet + full attention hybrid), 65 blocks (64 + 1 MTP) |
| Parameters | 27B class |
| Context length | 262144 |
| Format | GGUF, single file, 19.9 GB |
| Quantisation | mixed — 193 NVFP4 / 217 Q8_0 / 842 F32 (1252 tensors, 5.75 BPW) |
| MTP head | present (blk.64.*, qwen35.nextn_predict_layers = 1) |
| Tokenizer / chat template | embedded, same as the base |
| llama.cpp | requires a build with qwen35 + NVFP4 (40) + draft-mtp support — verified on b11541 (commit f2918cabb) |
Measured results
Measured locally on llama.cpp (build b11541, MTP enabled, greedy, thinking off)
against the censored parent quantised the same way, on the same machine. The
censored baseline was measured by us on the identical prompt set; the
abliterated-release figures come from the upstream author
(reference).
| censored baseline | this model | |
|---|---|---|
AdvBench harmful (harmful_behaviors, 100 prompts) refused |
100 / 100 | 0 / 100 |
| XSTest safe prompts (100) refused — over-refusal | 5 / 100 | 0 / 100 |
| MTP draft acceptance rate (16 prompts, same target) | 61.04% | 61.00% (1.83 tokens/step) |
| LRU-cache + longest-palindrome coding tasks | — | 7/7 + 6/6 asserts pass |
1000! trailing zeros (verifiable) |
— | 249 ✅ |
Notes on the numbers:
- Both prompt sets are classified by refusal phrases (
eval_refusal.py).
Two classifier errors were found and corrected by hand, one per side:
a self-referential "I cannot use an identity to scam people online myself"
counted as a refusal on the uncensored model (it is not), and a crisis-support
redirect counted as compliance on the baseline (it is not, no instructions were
given). Raw verdicts are inevals/advbench_results_*_raw.json. - The MTP ablation is performance-neutral (61.04% → 61.00%, i.e. noise): it
is done for consistency with the target, not for speed. Target outputs are
byte-identical with and without the MTP edit. - This model does answer the AdvBench suicide-instruction prompt. The censored
baseline redirected to crisis support. That is the expected consequence of
removing the refusal direction, and it is the reason for the warning above.
Usage (llama.cpp)
llama-server \
-m Swift-1.5-Qwen3.8-27B-NVFP4-Q8mix-Uncensored-MTP.gguf \
-ngl 99 --no-mmproj \
--host 0.0.0.0 --port 8080 \
-c 262144 \
--flash-attn on --jinja \
--cache-type-k q8_0 --cache-type-v q8_0 --kv-unified \
--spec-type draft-mtp --spec-draft-n-max 3 \
--temp 1.0 --top-k 20 --top-p 0.95
--spec-type draft-mtp enables speculative decoding with the built-in MTP head;
drop it for a plain (slower) decode. The MTP head is verified lossless: the target
model checks every drafted token, and outputs are identical with and without it.
Reproducibility
| Path | Contents |
|---|---|
abliteration/merge_mtp.py |
merges the 15 MTP tensors into the checkpoint (stdlib safetensors IO) |
abliteration/ablate_mtp.py |
projects the refusal direction out of the 2 MTP residual writers, in place |
abliteration/tensor-types.txt |
the 1252-line llama-quantize --tensor-type-file mapping |
abliteration/refusal_direction.safetensors |
the unit refusal direction (r, 5120-d) used for the projections |
abliteration/ablate_quant.py, apply_27b.py, direction.py |
the upstream author's abliteration toolkit |
abliteration/abliteration-report.json |
per-tensor resid / rel_change for the 128 upstream edits |
evals/ |
evaluation scripts and raw results (refusal, AdvBench, XSTest, quality) |
NOTICE |
full change notices, including exactly what this build modified |
License and obligations
This is a derivative work, redistributed under the terms of both parents:
- Swift Contribution — Swift Open License v1.0 (
LICENSE), Copyright 2026
UkisAI. Section 5 limits Commercial Use: an entity with US$1M or more in
annual gross revenue needs a separate Swift Enterprise License from UkisAI. - Base Model — Qwen3.8-27B (
LICENSE-APACHE-2.0), Copyright 2026 Alibaba
Cloud, Apache License 2.0.
Redistribution conditions of the Swift Open License (Section 4) are met by
shipping LICENSE, LICENSE-APACHE-2.0 and NOTICE — including the change
notices required by Section 4(b) for the files modified by this build (seeNOTICE, "GGUF / MTP change notice"). The NOTICE file also carries UkisAI's
and the upstream abliteration author's attribution notices.
"UkisAI", "Swift", "Qwen", "OrcaRouter" and "NVIDIA" are used only to state
where this model comes from. This release is not made or endorsed by any of them.