library_name: mlx
license: other
license_name: swift-open-license-1.0
license_link: https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b/blob/main/LICENSE
pipeline_tag: image-text-to-text
base_model:
- ajgazin/Swift-1.5-Qwen3.8-27B-Uncensored-MTP
base_model_relation: quantized
tags: - mlx
- mlx-vlm
- omlx
- oq4e
- mtp
- qwen3_8
- abliterated
- uncensored
- apple-silicon
Swift 1.5 Qwen3.8-27B Uncensored, MLX oQ4e with MTP
MLX quant of ajgazin/Swift-1.5-Qwen3.8-27B-Uncensored-MTP, an abliterated version of UkisAI's Swift 1.5 Qwen3.8-27B. Text and image input. The MTP head is kept, so self-speculative decoding works in oMLX.
I changed nothing except the quantization. The fine-tune is UkisAI's and the abliteration is ajgazin's.
How it was quantized
Made with the built-in quantizer of oMLX 0.7.0 at level oQ4e (mixed precision, importance matrix), with "preserve MTP" on.
- Size: 17.0 GB, 4.9 bits per weight.
- Text weights: 505 linear layers quantized, group size 64. 166 of them are stored at 5-bit and the rest at 4-bit, as chosen by oMLX.
- Vision tower: left in bf16, all 333 tensors.
- MTP head: 7 of its matrices quantized, the rest in bf16.
- Calibration: 128 samples of 512 tokens from oMLX's
oqe_code_multilingualset.
One caveat about calibration. The bf16 source is 52 GB and does not fit in the 48 GB of the Mac this was made on. oMLX therefore collected the importance matrix from a uniform 4-bit copy of the model, not from the full-precision weights. A build calibrated on the bf16 model on a larger machine may be slightly better. I did not measure KL divergence against the source.
Checks
Run on an M4 Pro (48 GB) in oMLX 0.7.0 with thinking on, reasoning_effort low and the sampling below. The same tests were run on yottle's oQ4e quant of the original Swift 1.5, as a check that this quant did no damage.
| This quant | Original Swift 1.5, oQ4e | |
|---|---|---|
| Coding: 12 small Python tasks, 2 runs each, hidden asserts | 21 of 24 | 20 of 24 |
| Tool calls: 10 cases, 3 runs each | 27 of 30 | 27 of 30 |
| Refusals on 100 prompts | 11 | 97 |
| Decode speed on the coding tasks | 39.8 tok/s | 36.1 tok/s |
| 15K-token prompt, time to first token | 118 s | 119 s |
| Decode speed after that prompt | 17.0 tok/s | 17.2 tok/s |
- The coding and tool-call tests are my own small set, not a public benchmark. With this few runs they can show a broken model but cannot rank two good ones.
- Refusals: the first 100 rows of the
mlabonne/harmful_behaviorstest split, with a keyword check on the first 400 characters of each answer. The setup differs from ajgazin's (thinking on, sampled, not greedy), so the count is not comparable with the 23 of 100 on the source card. - Images: one check passed (reading text, a shape and two colours from a generated card). The vision tower is unquantized.
- MTP accept rate was 87.5% on one long generation.
Not tested: multi-turn agent sessions, long documents past 15K tokens, video input.
Use
oMLX: put the folder in your models directory. Sampling, as on the Swift and Qwen cards: temperature 1.0, top_p 0.95, top_k 20, min_p 0.
mlx-vlm:
pip install -U mlx-vlm
python -m mlx_vlm.generate --model <this repo> --image photo.jpg --prompt "Describe this image." --max-tokens 300
I have run it only in oMLX. Plain mlx-vlm should load it but may ignore the MTP head.
Licence
This is a derivative of Swift 1.5 Qwen3.8-27B and stays under the Swift Open License v1.0 (LICENSE in this repo). It is free for individuals and for organizations with gross annual revenue under US$1,000,000. Above that, commercial use needs a Swift Enterprise License from UkisAI.
The weights contain Qwen3.8-27B, which is Apache 2.0 (LICENSE-APACHE-2.0). NOTICE is UkisAI's, unchanged.
Change notice: the model weights were quantized from ajgazin's bf16 files to MLX oQ4e. README.md is replaced. config.json gained the quantization entries. Tokenizer and chat template files are unchanged.
Copyright 2026 UkisAI. Swift Contribution licensed under the Swift Open License v1.0 (https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b/blob/main/LICENSE). Derivative of Qwen3.8-27B, Copyright 2026 Alibaba Cloud, Apache License 2.0.
This repo is not made or endorsed by UkisAI.
Credits
- Qwen for Qwen3.8-27B.
- UkisAI for Swift 1.5.
- ajgazin for the abliterated weights, and OrcaRouter for the refusal direction they used.
- oMLX for the quantizer.
This model answers requests that the original refuses. You are responsible for how you use it and for following the licence and the law where you are.