license: other
license_name: swift-open-license-1.0
license_link: https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b/blob/main/LICENSE
library_name: transformers
pipeline_tag: image-text-to-text
base_model: ajgazin/Swift-1.5-Qwen3.8-27B-Uncensored-MTP
base_model_relation: quantized
tags:
- qwen3_8
- abliterated
- uncensored
- autoround
- w4a16
- int4
- compressed-tensors
- mtp
Swift 1.5 Qwen3.8-27B Uncensored MTP W4A16 (AutoRound, BF16 MTP head)
A community 4-bit weight-only quantization of
ajgazin/Swift-1.5-Qwen3.8-27B-Uncensored-MTP,
an abliterated version of UkisAI's
Swift 1.5 Qwen3.8-27B. The parent projects
a single refusal direction out of Swift 1.5's residual-writing weights and edits the MTP head
consistently, so self-speculative decoding still works. It is not an official UkisAI release.
- W4A16: int4 weights (group size 128, symmetric), 16-bit activations, made with Intel
AutoRound 0.15.1 and exported as compressed-tensors. vLLM picks the int4 kernels up
automatically (Machete on Hopper, Marlin elsewhere). - 19.47 GB on disk versus about 56 GB for the BF16 parent.
- BF16 MTP head: the edited multi-token-prediction module is kept at full precision and
listed in the quantization ignore list, so MTP decoding works as it does on the parent.
Status: not yet evaluated. This checkpoint has not been benchmarked, load-tested in vLLM,
or checked for refusal behaviour. It uses the same recipe, layout and toolchain as
causal/Swift-1.5-Qwen3.8-27b-W4A16-AutoRound
and the Swift 1.0 version
causal/Swift-Qwen3.8-27b-W4A16-AutoRound-MTP-BF16,
which serves in vLLM 0.27.1.
Evaluation
None measured for this repository yet. The parent card reports these results for the BF16
parent, measured with Heretic's built-in evaluation; they are
copied here for reference.
| Model | Refusals | KL divergence |
|---|---|---|
| ajgazin/Swift-1.5-Qwen3.8-27B-Uncensored-MTP (BF16, against Swift 1.5) | 23/100 | 0.0884 |
| ukisai/Swift-1.5-Qwen3.8-27b (BF16) | 98/100 | 0 |
- Refusals: 100 prompts from
mlabonne/harmful_behaviors, greedy, up to 100 tokens, with
thinking closed immediately. - KL divergence: first-token distributions on 100 prompts from
mlabonne/harmless_alpaca. - Calibration here used general web text (
NeelNanda/pile-10k), not refusal data. 4-bit
rounding can shift refusal behaviour in either direction, and that has not been measured. - The parent card lists general benchmarks, reasoning-mode refusals, reasoning length and MTP
acceptance as not evaluated; the same holds for this quantization.
Quantization details
| Setting | Value |
|---|---|
| Method | AutoRound 0.15.1, W4A16, group size 128, symmetric |
| Quantized | 400 Linear layers: GatedDeltaNet in_proj_qkv / in_proj_z / out_proj, all MLP projections, full-attention q/k/v/o |
| Kept in BF16 | vision tower (110 Linear), GatedDeltaNet in_proj_a / in_proj_b (96), lm_head, embeddings, MTP head (8 Linear) |
| Calibration | NeelNanda/pile-10k, 128 samples x 2,048 tokens, 200 iterations, batch 4, seed 42 |
| Export | compressed-tensors (pack-quantized) |
| Toolchain | auto-round 0.15.1, transformers 5.17.0, vLLM 0.29.0 image (CUDA 13.0) |
| Cost | 46 min on one NVIDIA L40S, peak 17.4 GB VRAM |
The abliteration edits self_attn.o_proj, linear_attn.out_proj, mlp.down_proj andembed_tokens. The first three are quantized here along with the rest of the text model, except
for the two edited MTP tensors, which stay BF16 with the rest of the MTP head; embed_tokens
stays BF16. The layer selection follows dbirks/Qwen3.8-27B-W4A16-AutoRound,
except that the MTP head stays BF16.
How to use
vLLM
vllm serve causal/Swift-1.5-Qwen3.8-27B-Uncensored-MTP-W4A16-AutoRound \
--max-model-len 262144 \
--reasoning-parser qwen3 \
--enable-auto-tool-choice \
--tool-call-parser qwen3_coder \
--port 8000
The Swift 1.0 version, which has the same size and layout, loads in 18.5 GiB of GPU memory and
fits one full 262K-token request on a single 48 GB L40S. Lower --max-model-len on smaller GPUs.
Optional MTP decoding
--speculative-config '{"method":"mtp","num_speculative_tokens":3}'
Sampling
As for Swift and Qwen: temperature 1.0, top_p 0.95, top_k 20, min_p 0.
Intended use
The model answers requests the original declines. You are responsible for how you use it and
for complying with applicable law and the license.
License
A derivative of Swift 1.5 Qwen3.8-27B under the Swift Open License v1.0.
Qwen3.8-27B and orcarouter/Qwen3.8-27B-Uncensored are licensed under
Apache 2.0. See NOTICE.
Use is free for individuals and organizations with gross annual revenue, including affiliates,
of up to US$1,000,000. Above that threshold, commercial use requires a Swift Enterprise License
from UkisAI.
Changes from the parent: the model weights were quantized to int4 as described above
(model-*.safetensors, model_extra_tensors.safetensors, model.safetensors.index.json);config.json gained a quantization_config and quantization_config.json was added; the other
config, tokenizer and processor files were re-saved by transformers 5.17.0 during export; this
README replaces the parent's. The parent's abliteration/ folder and abliteration.json are not
included; see the parent repository for the refusal direction and scripts. LICENSE,
LICENSE-APACHE-2.0 and NOTICE are the parent's.
Credits
- Qwen for Qwen3.8-27B.
- UkisAI for Swift 1.5 Qwen3.8-27B.
- OrcaRouter for Qwen3.8-27B-Uncensored and its refusal direction.
- ajgazin for transferring the edit to Swift 1.5 with the MTP head.
- Arditi et al., Refusal in Language Models Is Mediated by a Single Direction (2024).
Quantization for this repository ran on a Modal L40S.