license: apache-2.0
base_model: orcarouter/Qwen3.8-27B-Uncensored
base_model_relation: quantized
library_name: mlx
pipeline_tag: text-generation
language:
- zh
- en
tags: - mlx
- apple-silicon
- qwen
- qwen3.8
- qwen3_5
- 3-bit
- mixed-precision
- imatrix
- quantized
- abliterated
- uncensored
Qwen3.8-27B-Uncensored · MLX 3-bit (mixed precision)
An MLX build of orcarouter/Qwen3.8-27B-Uncensored for Apple Silicon, quantized to fit in under 10 GB.
It is a 3-bit-class quantization: instead of one bit width for every layer, an importance-weighted plan gives each linear layer 2, 3 or 4 bits under a fixed size budget. The result averages 2.9 bits per weight including scales. That is a little smaller than a uniform 3-bit model (3.25 bits per weight at group size 128, 3.5 at 64).
| Size on disk | 9.82 GB (9,820,551,838 bytes, 2 shards) |
| Effective bits per weight | 2.915 (incl. quantization scales and biases) |
| Peak memory, short prompt | ~10.5 GB |
| Perplexity | 7.39 vs 6.76 for the BF16 source (+9%, see Evaluation) |
| Modality | Text only (vision tower and MTP head are not included) |
| Requires | mlx-lm >= 0.32 |
⚠️ Disclaimer
The source model has had its safety alignment removed by abliteration, and this quantization keeps that behaviour. It will comply with harmful, unethical or illegal requests that the original Qwen3.8-27B would refuse.
- It is meant for research: interpretability, refusal-mechanism and safety studies, red-teaming and robustness evaluation.
- You are responsible for how you use it and for everything it generates. Don't deploy it to end users without your own safety and moderation layers.
- Use must follow the Apache 2.0 license and the laws that apply to you. The uploader accepts no liability for misuse.
See the source model card for the abliteration method and its refusal benchmarks.
Usage
pip install -U mlx-lm
Command line:
mlx_lm.generate --model nyaaorick/Qwen3.8-27B-Uncensored-MLX-3bit \
--prompt "用三句话介绍一下杭州。" --max-tokens 1024
Python:
from mlx_lm import load, generate
model, tokenizer = load("nyaaorick/Qwen3.8-27B-Uncensored-MLX-3bit")
messages = [{"role": "user", "content": "Explain why the sky is blue."}]
prompt = tokenizer.apply_chat_template(messages, add_generation_prompt=True)
print(generate(model, tokenizer, prompt=prompt, max_tokens=1024, verbose=True))
OpenAI-compatible server:
mlx_lm.server --model nyaaorick/Qwen3.8-27B-Uncensored-MLX-3bit --port 8080
The model thinks before answering by default. To turn thinking off, pass enable_thinking=False to apply_chat_template. Sampling defaults come from generation_config.json (temperature 1.0, top-p 0.95, top-k 20).
Hardware: weights take ~9.2 GiB of unified memory. A 16 GB Mac can run it with short contexts and little else open; 24 GB or more is recommended for long contexts.
Quantization
Bit allocation
| Tier | Group size | Effective bits | Share of linear-layer params | Layers |
|---|---|---|---|---|
| 4-bit | 128 | 4.25 | 11.0% | 76 |
| 3-bit | 64 | 3.5 | 19.2% | 157 |
| 2-bit | 64 | 2.5 | 69.8% | 264 |
embed_tokensis 3-bit, group size 64.lm_headand every linear layer in the first two and last two blocks are fixed at 4-bit.- Norms, convolutions and other small tensors stay in BF16.
Per-module breakdown (number of layers at 2 / 3 / 4 bits)
| Module | 2-bit | 3-bit | 4-bit |
|---|---|---|---|
linear_attn.in_proj_qkv |
20 | 25 | 3 |
linear_attn.in_proj_z |
15 | 30 | 3 |
linear_attn.in_proj_a |
0 | 8 | 40 |
linear_attn.in_proj_b |
0 | 41 | 7 |
linear_attn.out_proj |
45 | 0 | 3 |
self_attn.q_proj |
0 | 15 | 1 |
self_attn.k_proj |
2 | 13 | 1 |
self_attn.v_proj |
5 | 7 | 4 |
self_attn.o_proj |
15 | 0 | 1 |
mlp.gate_proj |
50 | 10 | 4 |
mlp.up_proj |
52 | 8 | 4 |
mlp.down_proj |
60 | 0 | 4 |
lm_head |
0 | 0 | 1 |
The exact per-layer assignment is in quantization/plan.json, and in the quantization section of config.json.
Method
- Activation statistics. One forward pass over ~100k calibration tokens records the mean squared input activation of every linear layer (an importance matrix).
- Error per tier. Each layer's weights are quantized at each tier. The reconstruction error is weighted by those activation statistics, so input channels that matter more count more.
- Greedy allocation under a byte budget. Every layer starts at 2-bit. The script repeatedly applies the upgrade (2→3 or 3→4) that removes the most weighted error per extra byte, until the 9.8 GB budget is spent. The budget counts actual bytes, including scales, the embedding table and the BF16 tensors, so the final size is guaranteed.
- Apply.
mlx_lm'squantize_modelis run with the per-layer plan.
Calibration data: about 70% Chinese and 30% English. The Chinese text is from Wikipedia (wikimedia/wikipedia, 20231101.zh) and the English from the mlx-lm calibration_v5 text, cut into ~1000-token chunks and shuffled.
Hardware used: one NVIDIA A100 80 GB with MLX's CUDA backend (mlx 0.32.3, mlx-lm 0.32.0). Planning took ~8 minutes (peak 72 GB of GPU memory) and quantizing about 1 minute.
Evaluation
Perplexity with mlx_lm.perplexity: 128 samples × 512 tokens from allenai/tulu-3-sft-mixture, seed 123, same settings for both models.
| Model | Perplexity |
|---|---|
BF16 source (orcarouter/Qwen3.8-27B-Uncensored) |
6.763 ± 0.088 |
| This model | 7.389 ± 0.087 |
No downstream benchmarks have been run on this build. With about 70% of the weights at 2-bit, expect a noticeable drop compared with 4-bit builds on hard reasoning, maths and long-form tasks.
Reproduce
The scripts in quantization/ produce this repo from the BF16 source:
hf download orcarouter/Qwen3.8-27B-Uncensored --local-dir ./src
python quantization/build_calib.py # needs ./zh.parquet (zh Wikipedia shard) and ./en.txt
python quantization/imatrix_plan.py # writes ./plan.json
python quantization/apply_plan.py # writes ./out
Limitations
- Text only.
mlx_lmdrops the vision tower and the MTP speculative-decoding head of the source model. - Safety guardrails removed (see the disclaimer). It also inherits the biases and limitations of Qwen3.8-27B.
- The quantization was calibrated mostly on Chinese text, and the Chinese Wikipedia text is largely Traditional characters.
License and credits
Apache 2.0, inherited from Qwen/Qwen3.8-27B. The full text is in LICENSE.
- Base model: Qwen/Qwen3.8-27B by the Qwen team.
- Abliterated BF16 weights: orcarouter/Qwen3.8-27B-Uncensored.
- This repo only adds the MLX conversion and mixed-precision quantization. It is not affiliated with Qwen, Alibaba or OrcaRouter.