license: apache-2.0
base_model: orcarouter/Qwen3.8-27B-Uncensored
base_model_relation: quantized
library_name: mlx
pipeline_tag: image-text-to-text
tags:
- mlx
- mtplx
- qwen
- qwen3
- abliterated
- uncensored
- vision
- quantized
- 3-bit
- speculative-decoding
- mtp
Qwen3.8-27B-Uncensored-MTPLX-3bit
MTPLX 3-bit conversion of orcarouter/Qwen3.8-27B-Uncensored — the abliterated (refusal-removed) BF16 build of Qwen3.8-27B — for speculative decoding on Apple Silicon via mtplx.
Sibling: orcarouter/Qwen3.8-27B-Uncensored-MTPLX-4bit (best performer in this family).
Converted straight from the bf16 safetensors weights (not from any GGUF quant), so there is no dequantize-requantize drift in the chain.
Abliterated / uncensored. This model has had refusal behavior removed. It will comply with requests the base Qwen model would refuse. Use responsibly and in accordance with your local laws and the model's license. No additional content filtering is applied beyond what is in the checkpoint.
Why this release exists
OrcaRouter's uncensored Qwen3.8 build preserves the full BF16 base — 27B dense, hybrid Gated-DeltaNet / full-attention, 64 layers, 262K context, vision + MTP — with refusal ablated for red-teaming and open-ended use. MTPLX 3-bit keeps that checkpoint intact and adds a bf16 MTP sidecar: at ~13.6 GB on disk it fits 24 GB unified-memory Macs while delivering strong speculative speedups (2.14x at D2 verified).
This repo is the 3-bit affine point. For the highest speedups and near-lossless draft acceptance (~98%), use the 4-bit sibling.
Available MTPLX conversions
| Repo | Bits | Size (on disk) | Use case |
|---|---|---|---|
| this repo (3-bit) | 3 | ~13.6 GB (12.6 GiB) | tight unified memory |
| orcarouter/Qwen3.8-27B-Uncensored-MTPLX-4bit | 4 | ~16.9 GB | best performer, recommended |
Source bf16 remains at orcarouter/Qwen3.8-27B-Uncensored.
Which one to pick:
- Highest throughput + acceptance → 4-bit MTPLX (2.56x at D2, 98.4%/98.3% acceptance)
- Tightest fit → this 3-bit MTPLX (2.14x at D2, 93.1%/86.9% — still excellent)
- 36 GB+ Mac with headroom → either; prefer 4-bit
Conversion notes
- Source:
orcarouter/Qwen3.8-27B-Uncensored(bf16, 18 shards + vision,orcarouter--Qwen3.8-27B-Uncensoredlocal cache), forge-local viamtplx==2.9.1 - Recipe:
body_bits=3,body_group_size=64,body_mode=affine,mtp_policy=keep_bf16, quantized trunk + bf16 MTP sidecar - MTP contract:
base_hidden_variant=post_norm,hidden_variant=post_norm,concat_order=embedding_hidden,mtp_position_mode=local,mtp_quant_group_size=64,mtp_quant_mode=affine - Output: 3 safetensors shards +
model-vision.safetensors(879 MB, 333 tensors, bf16) +mtp.safetensors(810 MB, bf16 sidecar) +model.safetensors.index.json - Forged at: 2026-08-24T12:13:04+03:00 on Apple M5 Pro (18 CPU / 20 GPU, 24 GB unified memory) — tuned for 24 GB Macs · macOS 27.0 arm64,
mtplx_runtime.jsonships in-repo as provenance
Vision
The vision encoder is inside these weights (stored as model-vision.safetensors, auto-loaded) — no separate mmproj to fetch:
mtplx run --model orcarouter/Qwen3.8-27B-Uncensored-MTPLX-3bit \
--prompt "Describe this image." --image photo.jpg --depth 2
Text-only needs nothing extra.
MTPLX usage
# install
pip install -U mtplx
# single-turn generation (verified optimum is --depth 2 for this 3-bit)
mtplx run --model orcarouter/Qwen3.8-27B-Uncensored-MTPLX-3bit \
--prompt "Write a short story about a robot learning to paint." \
--depth 2 --max-tokens 512
# multimodal
mtplx run --model orcarouter/Qwen3.8-27B-Uncensored-MTPLX-3bit \
--prompt "Describe this image." --image photo.jpg --depth 2
# OpenAI-compatible local server
mtplx serve --model orcarouter/Qwen3.8-27B-Uncensored-MTPLX-3bit \
--depth 2 --port 8080
# then: curl http://localhost:8080/v1/chat/completions ...
# custom sampling (verified defaults: temp 0.6, top_p 0.95, top_k 20)
mtplx run --model orcarouter/Qwen3.8-27B-Uncensored-MTPLX-3bit \
--prompt "Explain quantum entanglement simply." \
--depth 2 --temperature 0.6 --top-p 0.95 --top-k 20
Verification results
Verified locally with mtplx forge verify (MLX backend, 512-token budget, quality gate enabled). Depth 0 is plain AR.
| Depth | tok/s | vs AR | Acceptance by position | Quality |
|---|---|---|---|---|
| 0 (AR) | 20.7 | 1.00x | — | pass |
| 1 | 33.7 | 1.63x | 95.7% | pass |
| 2 | 44.3 | 2.14x | 93.1% / 86.9% | pass |
| 3 | 40.4 | 1.95x | 94.0% / 88.0% / 79.7% | pass |
- Recommended depth: 2 (
mtp_depth_winsat D2; D3 passes quality but loses throughput) - Verdict:
mtp_depth_wins;quality_rejected=[],failure_reasons=[] - Hardware: Apple M5 Pro (18 CPU / 20 GPU, 24 GB) · macOS 27.0 arm64 (Apple Silicon, MLX), mtplx 2.9.1, artifact
sha256:a8cdfcaf... - All depths passed the quality gate. D2 is the fastest verified depth; D3's raw acceptance is slightly higher but its extra draft cost outweighs the gain for this 3-bit trunk.
Training details (source model)
- Base: Qwen/Qwen3.8-27B — dense 27B, hybrid Gated-DeltaNet / full-attention, 64 layers, 262K context
- Post-training: Abliteration (refusal-direction removal) by OrcaRouter — BF16, no quantization in the source; vision + MTP preserved
- License: Apache 2.0 (inherited)
- Context & modalities: 262K text, vision-language, function calling, reasoning
Source chain
Qwen/Qwen3.8-27B (base)
→ orcarouter/Qwen3.8-27B-Uncensored (BF16 abliterated)
→ this repo (MTPLX 3-bit conversion)
Related in this collection:
- orcarouter/Qwen3.8-27B-Uncensored-MTPLX-4bit — 4-bit MTPLX, best performer