license: other
license_name: swift-open-license-1.0
license_link: https://ukisai.com/contact
base_model: ajgazin/Swift-1.5-Qwen3.8-27B-Uncensored-MTP
base_model_relation: quantized
library_name: transformers
language:
- en
- zh
pipeline_tag: image-text-to-text
tags: - awq
- compressed-tensors
- int4
- uncensored
- abliterated
- mtp
- speculative-decoding
- intel
- arc
- intel-arc
- xpu
- vllm
- qwen3.8
- swift
Swift-1.5-Qwen3.8-27B-Uncensored-AWQ-Int4-asym-G128-MTP-Int4
AWQ W4A16 (asymmetric (per-group zero points), group size 128) quantization of ajgazin/Swift-1.5-Qwen3.8-27B-Uncensored-MTP, with
UkisAI's own AWQ recipe for Swift 1.5 — the recipe.yaml published with
ukisai/Swift-1.5-Qwen3.8-27b-W4A16-AWQ — and the MTP draft head kept, so native
speculative decoding works. Built for Intel Arc / vLLM XPU; acompressed-tensors checkpoint, so stock vLLM loads it too.
Why this recipe
Swift 1.5's selling point is short reasoning. Quantization can quietly undo that,
and the recipe decides how much. Measured on GPQA-Diamond under Swift 1.5's
published protocol (EvalScope, temperature 1.0, top-p 0.95, top-k 20, xhigh),
each pair compared question by question (95% interval of the per-question
reasoning-length ratio):
| comparison | reasoning length | score Δ |
|---|---|---|
| GPTQ int4, gptqmodel, c4 calibration (our previous build) vs its BF16 source | 1.35x (1.24–1.48) | −0.5 ± 1.8 pp |
| abliteration (ajgazin BF16) vs Swift 1.5 BF16 | 1.00x (0.92–1.09) | −1.0 ± 1.9 pp |
| ukisai's AWQ (this recipe) on Swift 1.5 vs Swift 1.5 BF16 | 1.02x (0.95–1.10) | −0.8 ± 1.5 pp |
The scores sit within noise throughout; the reasoning length does not. All of the
extra reasoning came from the GPTQ recipe, none from the abliteration, and this
AWQ recipe adds none. Hence this checkpoint: the AWQ recipe on the abliterated model.
This checkpoint:
Not yet measured on this checkpoint. The figures above are for ukisai's AWQ of the original Swift 1.5 and for the abliteration in BF16. The recipe is the same, and the abliteration added no reasoning length, so the expectation is about 1.0x Swift 1.5 BF16. It is an expectation, not a measurement.
The MTP draft head
ukisai's recipe ignores mtp.*, but transformers never instantiates the draft
head at all, so llm-compressor writes a checkpoint without one. ukisai re-attached
the 15 BF16 tensors afterwards — and left them out of ignore, so the
checkpoint's single Linear config group covers them: vLLM builds the drafter as
W4A16 and cannot load BF16 weights into it. This build re-attaches them too, and
then either lists them in ignore or packs them. MTP draft head: int4. The 8 drafter linears are rounded to
int4 (round-to-nearest, asymmetric, group 128) and packed in the
body's own format by requant-mtp-int4.py; the 15 norms stay BF16. Its note:mtp.* linears RTN int4 g128 asym (compressed-tensors pack-quantized, library-checked) from BF16 by quantize/requant-mtp-int4.py (805c132); rel err 0.1071. On two Arc Pro B70s, the same int4 drafter on ukisai's AWQ accepted 3.309
tokens per step against 3.316 for our GPTQ checkpoint's int4 drafter (30 runs each, same
prompts): unchanged. The target verifies every draft token, so the drafter can only cost
speed, never change the output.
greglechin/Swift-1.5-Qwen3.8-27B-Uncensored-AWQ-Int4-asym-G128-MTP-BF16 is the same body with the BF16 drafter, for stacks that reject this one.
Provenance
Qwen/Qwen3.8-27B (base)
└─ ukisai/Swift-1.5-Qwen3.8-27b (Swift 1.5, reasoning-efficient finetune)
└─ ajgazin/Swift-1.5-Qwen3.8-27B-Uncensored-MTP
(uncensored, orcarouter rank-1 refusal direction projected out, BF16)
└─ this repo (AWQ W4A16 asym g128, int4 MTP drafter, 19 GB)
Quantization
| Tool | llm-compressor 0.13.0, compressed-tensors 0.18.0, transformers 5.14.1, torch 2.13.0+cu130 |
| Recipe | ukisai's, unchanged; as run: recipe-as-run.yaml in this repo |
| AWQ | smoothing input_layernorm → q/k/v (16 full-attention layers), post_attention_layernorm → gate/up, up → down; duo_scaling: both, n_grid: 20; GDN projections and o_proj are plain RTN |
| Weights | int4, group 128, asymmetric (per-group zero points); 400 packed modules |
| Unquantized | lm_head, the vision tower (333 tensors), GDN in_proj_a/in_proj_b |
| Calibration | HuggingFaceH4/ultrachat_200k train_sft @ 8049631c405a: 256 conversations (shuffle(seed=42).select(range(256))), rendered with the model's chat template, ≤2048 tokens, 308,054 tokens; sha256 0012068e9416cfc6… |
| Time | 35.6 min of oneshot on NVIDIA RTX PRO 6000 Blackwell Workstation Edition (36.6 min end to end) |
| Output | 2 shards, 18.9 GB, 2423 tensors |
ukisai's manifest records a tokenized sha256 for their calibration set; their
preprocessing is not published and none of the variants tried reproduces it, so
the calibration set here is equivalent in kind, not provably identical.
Serving (vLLM)
vllm serve /model \
--kv-cache-dtype fp8 \
--speculative-config '{"method":"mtp","num_speculative_tokens":5}' \
--chat-template /model/chat_template.jinja
compressed-tensors is detected from the config. On Intel Arc, Qwen3.8 MTP needs a
patched vLLM: the build used here is github.com/greglechin/vllm-xpu. Asymmetric int4 runs about 8% slower on prefill than symmetric there; decode was within noise of our symmetric GPTQ build (35.7 against 36.7 steps/s).
Limitations — please read
- Calibration is English chat (ultrachat_200k). Code-heavy or non-English use may
quantize slightly worse than general chat. - This is an uncensored model. It will attempt requests an aligned model declines. You
are responsible for how you use it. - Uncensoring is inherited, not verified here. All refusal-removal properties come
from the upstream ablation; this repo only changes numeric precision, and no refusal
testing was run on it. - The upstream publishes a KL divergence and states what it did not measure. It reports KL 0.0884 against
ukisai/Swift-1.5-Qwen3.8-27bon first-token distributions over 100 harmless prompts, and 23/100 refusals, both under Heretic's built-in evaluation with the reference rows measured the same way. Its refusal direction is recovered from orcarouter's published weights rather than fitted fresh, and it verifies that recovery by reproducing orcarouter's 131 edited tensors to 99.75% bit-identical. What it explicitly does not evaluate: general benchmarks, refusal behaviour in thinking mode, whether Swift 1.5's shorter reasoning traces survive, and MTP acceptance — which on a speculative-decoding deployment is the number that decides decode throughput. Measure acceptance on your own traffic before promoting this over a checkpoint you have already gated.
Licence
Not Apache-2.0. These weights inherit the Swift Open License v1.0 from
ukisai/Swift-1.5-Qwen3.8-27b. Personal, research, educational, evaluation and commercial
use are free for individuals and organizations with annual recurring revenue,
including affiliates, of up to US$1,000,000. Above that threshold, commercial use
requires a separate Swift Enterprise License — contact
UkisAI.
Credits
- Qwen — base model
- UkisAI — Swift 1.5, and the AWQ recipe used here
- ajgazin — directional ablation
- llm-compressor — quantization