← back to catalog · registered 2026-10-01 14:58

causal/Swift-1.5-Qwen3.8-27B-Uncensored-MTP-W4A16-AutoRound

causal 27B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/causal%2FSwift-1.5-Qwen3.8-27B-Uncensored-MTP-W4A16-AutoRound"
Response includes
  • classification m-uncensored
  • files 23
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-10-01

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
other
Tags
transformers safetensors qwen3_5 image-text-to-text qwen3_8 abliterated uncensored autoround w4a16 int4 compressed-tensors mtp

Related

Total size
18.1 GB
Files
23
Quantizations
1
Registered
2026-10-01 14:58
Last updated on HF
2026-10-01 14:34

Files by quantization

Auxiliary files 23 files 18.1 GB
model-00004-of-00007.safetensors 3.00 GB 452c24bb download
model-00001-of-00007.safetensors 2.99 GB ee4a654e download
model-00002-of-00007.safetensors 2.98 GB 0a85101c download
model-00003-of-00007.safetensors 2.98 GB f004eede download
model-00006-of-00007.safetensors 2.37 GB 720b113c download
model-00007-of-00007.safetensors 2.37 GB 677b4d4a download
model_extra_tensors.safetensors 810 MB 2fce17ae download
model-00005-of-00007.safetensors 667 MB 0994bb8d download
tokenizer.json 19.1 MB 06b95093 download
model.safetensors.index.json 194 KB e3a119c1 download
config.json 15.3 KB 8ce72ff0 download
LICENSE 13.0 KB 209a5720 download
LICENSE-APACHE-2.0 11.3 KB f938136e download
quantization_config.json 11.1 KB 2294a867 download
chat_template.jinja 8.74 KB c0c686f9 download
README.md 6.20 KB f1cc0278 download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.19 KB 43c4343e download
tokenizer_config.json 1.14 KB 1d134cd2 download
NOTICE 1.11 KB c4ad1a71 download
preprocessor_config.json 390 B 2ea84a43 download
generation_config.json 205 B a382b041 download
.gitignore 47.0 B b5c9bd04 download

README current version from Hugging Face


license: other
license_name: swift-open-license-1.0
license_link: https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b/blob/main/LICENSE
library_name: transformers
pipeline_tag: image-text-to-text
base_model: ajgazin/Swift-1.5-Qwen3.8-27B-Uncensored-MTP
base_model_relation: quantized
tags:

  • qwen3_8
  • abliterated
  • uncensored
  • autoround
  • w4a16
  • int4
  • compressed-tensors
  • mtp

Swift 1.5 Qwen3.8-27B Uncensored MTP W4A16 (AutoRound, BF16 MTP head)

A community 4-bit weight-only quantization of
ajgazin/Swift-1.5-Qwen3.8-27B-Uncensored-MTP,
an abliterated version of UkisAI's
Swift 1.5 Qwen3.8-27B. The parent projects
a single refusal direction out of Swift 1.5's residual-writing weights and edits the MTP head
consistently, so self-speculative decoding still works. It is not an official UkisAI release.

  • W4A16: int4 weights (group size 128, symmetric), 16-bit activations, made with Intel
    AutoRound 0.15.1 and exported as compressed-tensors. vLLM picks the int4 kernels up
    automatically (Machete on Hopper, Marlin elsewhere).
  • 19.47 GB on disk versus about 56 GB for the BF16 parent.
  • BF16 MTP head: the edited multi-token-prediction module is kept at full precision and
    listed in the quantization ignore list, so MTP decoding works as it does on the parent.

Status: not yet evaluated. This checkpoint has not been benchmarked, load-tested in vLLM,
or checked for refusal behaviour. It uses the same recipe, layout and toolchain as
causal/Swift-1.5-Qwen3.8-27b-W4A16-AutoRound
and the Swift 1.0 version
causal/Swift-Qwen3.8-27b-W4A16-AutoRound-MTP-BF16,
which serves in vLLM 0.27.1.

Evaluation

None measured for this repository yet. The parent card reports these results for the BF16
parent, measured with Heretic's built-in evaluation; they are
copied here for reference.

Model Refusals KL divergence
ajgazin/Swift-1.5-Qwen3.8-27B-Uncensored-MTP (BF16, against Swift 1.5) 23/100 0.0884
ukisai/Swift-1.5-Qwen3.8-27b (BF16) 98/100 0
  • Refusals: 100 prompts from mlabonne/harmful_behaviors, greedy, up to 100 tokens, with
    thinking closed immediately.
  • KL divergence: first-token distributions on 100 prompts from mlabonne/harmless_alpaca.
  • Calibration here used general web text (NeelNanda/pile-10k), not refusal data. 4-bit
    rounding can shift refusal behaviour in either direction, and that has not been measured.
  • The parent card lists general benchmarks, reasoning-mode refusals, reasoning length and MTP
    acceptance as not evaluated; the same holds for this quantization.

Quantization details

Setting Value
Method AutoRound 0.15.1, W4A16, group size 128, symmetric
Quantized 400 Linear layers: GatedDeltaNet in_proj_qkv / in_proj_z / out_proj, all MLP projections, full-attention q/k/v/o
Kept in BF16 vision tower (110 Linear), GatedDeltaNet in_proj_a / in_proj_b (96), lm_head, embeddings, MTP head (8 Linear)
Calibration NeelNanda/pile-10k, 128 samples x 2,048 tokens, 200 iterations, batch 4, seed 42
Export compressed-tensors (pack-quantized)
Toolchain auto-round 0.15.1, transformers 5.17.0, vLLM 0.29.0 image (CUDA 13.0)
Cost 46 min on one NVIDIA L40S, peak 17.4 GB VRAM

The abliteration edits self_attn.o_proj, linear_attn.out_proj, mlp.down_proj and
embed_tokens. The first three are quantized here along with the rest of the text model, except
for the two edited MTP tensors, which stay BF16 with the rest of the MTP head; embed_tokens
stays BF16. The layer selection follows dbirks/Qwen3.8-27B-W4A16-AutoRound,
except that the MTP head stays BF16.

How to use

vLLM

vllm serve causal/Swift-1.5-Qwen3.8-27B-Uncensored-MTP-W4A16-AutoRound \
  --max-model-len 262144 \
  --reasoning-parser qwen3 \
  --enable-auto-tool-choice \
  --tool-call-parser qwen3_coder \
  --port 8000

The Swift 1.0 version, which has the same size and layout, loads in 18.5 GiB of GPU memory and
fits one full 262K-token request on a single 48 GB L40S. Lower --max-model-len on smaller GPUs.

Optional MTP decoding

--speculative-config '{"method":"mtp","num_speculative_tokens":3}'

Sampling

As for Swift and Qwen: temperature 1.0, top_p 0.95, top_k 20, min_p 0.

Intended use

The model answers requests the original declines. You are responsible for how you use it and
for complying with applicable law and the license.

License

A derivative of Swift 1.5 Qwen3.8-27B under the Swift Open License v1.0.
Qwen3.8-27B and orcarouter/Qwen3.8-27B-Uncensored are licensed under
Apache 2.0. See NOTICE.

Use is free for individuals and organizations with gross annual revenue, including affiliates,
of up to US$1,000,000. Above that threshold, commercial use requires a Swift Enterprise License
from UkisAI.

Changes from the parent: the model weights were quantized to int4 as described above
(model-*.safetensors, model_extra_tensors.safetensors, model.safetensors.index.json);
config.json gained a quantization_config and quantization_config.json was added; the other
config, tokenizer and processor files were re-saved by transformers 5.17.0 during export; this
README replaces the parent's. The parent's abliteration/ folder and abliteration.json are not
included; see the parent repository for the refusal direction and scripts. LICENSE,
LICENSE-APACHE-2.0 and NOTICE are the parent's.

Credits

Quantization for this repository ran on a Modal L40S.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.