license: apache-2.0
base_model: orcarouter/Qwen3.8-27B-Uncensored
base_model_relation: quantized
library_name: mlx
pipeline_tag: image-text-to-text
tags:
- mlx
- mtplx
- qwen
- qwen3
- abliterated
- uncensored
- vision
- quantized
- 4-bit
- speculative-decoding
- mtp
Qwen3.8-27B-Uncensored-MTPLX-4bit
MTPLX 4-bit conversion of orcarouter/Qwen3.8-27B-Uncensored — the abliterated (refusal-removed) BF16 build of Qwen3.8-27B — for speculative decoding on Apple Silicon via mtplx. Best performer in this MTPLX collection.
Sibling: orcarouter/Qwen3.8-27B-Uncensored-MTPLX-3bit.
Converted straight from the bf16 safetensors weights (not from any GGUF quant), so there is no dequantize-requantize drift in the chain.
Abliterated / uncensored. This model has had refusal behavior removed. It will comply with requests the base Qwen model would refuse. Use responsibly and in accordance with your local laws and the model's license. No additional content filtering is applied beyond what is in the checkpoint.
Why this release exists
OrcaRouter's uncensored Qwen3.8 preserves the full BF16 base — 27B dense, 64 layers, 262K context, vision + MTP — with refusal ablated. MTPLX 4-bit keeps that checkpoint byte-for-byte, quantizes the trunk to 4-bit affine (group 64), and leaves the MTP head in bf16. At ~16.9 GB on disk it is still comfortable on 36 GB unified-memory Macs, and it delivers the highest verified speedup in the five-model set: 2.56x at depth 2 with near-lossless draft acceptance (98.4% / 98.3%). If you can afford the extra ~3 GB over the 3-bit, this is the one to run.
Available MTPLX conversions
| Repo | Bits | Size (on disk) | tok/s (D2) | vs AR | Use case |
|---|---|---|---|---|---|
| orcarouter/Qwen3.8-27B-Uncensored-MTPLX-3bit | 3 | ~13.6 GB | 44.3 | 2.14x | tight unified memory |
| this repo (4-bit) | 4 | ~16.9 GB (15.7 GiB) | 45.1 | 2.56x | best performer, recommended |
Source bf16 remains at orcarouter/Qwen3.8-27B-Uncensored.
Conversion notes
- Source:
orcarouter/Qwen3.8-27B-Uncensored(bf16, 18 shards + vision,orcarouter--Qwen3.8-27B-Uncensoredlocal cache), forge-local viamtplx==2.9.1 - Recipe:
body_bits=4,body_group_size=64,body_mode=affine,mtp_policy=keep_bf16, quantized trunk + bf16 MTP sidecar - MTP contract:
base_hidden_variant=post_norm,hidden_variant=post_norm,concat_order=embedding_hidden,mtp_position_mode=local,mtp_quant_group_size=64,mtp_quant_mode=affine - Output: 3 safetensors shards +
model-vision.safetensors(879 MB, 333 tensors, bf16) +mtp.safetensors(810 MB, bf16 sidecar) +model.safetensors.index.json - Forged at: 2026-08-24T12:00:11+03:00 on Apple M5 Pro (18 CPU / 20 GPU, 24 GB unified memory) — tuned for 24 GB Macs · macOS 27.0 arm64,
mtplx_runtime.jsonships in-repo as provenance
Vision
The vision encoder is inside these weights (stored as model-vision.safetensors, auto-loaded) — no separate mmproj to fetch:
mtplx run --model orcarouter/Qwen3.8-27B-Uncensored-MTPLX-4bit \
--prompt "Describe this image." --image photo.jpg --depth 2
Text-only needs nothing extra.
MTPLX usage
# install
pip install -U mtplx
# single-turn generation (verified optimum is --depth 2)
mtplx run --model orcarouter/Qwen3.8-27B-Uncensored-MTPLX-4bit \
--prompt "Write a short story about a robot learning to paint." \
--depth 2 --max-tokens 512
# multimodal
mtplx run --model orcarouter/Qwen3.8-27B-Uncensored-MTPLX-4bit \
--prompt "Describe this image." --image photo.jpg --depth 2
# OpenAI-compatible local server
mtplx serve --model orcarouter/Qwen3.8-27B-Uncensored-MTPLX-4bit \
--depth 2 --port 8080
# then: curl http://localhost:8080/v1/chat/completions ...
# explicit sampling (verified defaults)
mtplx run --model orcarouter/Qwen3.8-27B-Uncensored-MTPLX-4bit \
--prompt "Explain quantum entanglement simply." \
--depth 2 --temperature 0.6 --top-p 0.95 --top-k 20
Verification results
Verified locally with mtplx forge verify (MLX backend, 512-token budget, quality gate enabled). Depth 0 is plain AR.
| Depth | tok/s | vs AR | Acceptance by position | Quality |
|---|---|---|---|---|
| 0 (AR) | 17.6 | 1.00x | — | pass |
| 1 | 30.3 | 1.72x | 98.4% | pass |
| 2 | 45.1 | 2.56x | 98.4% / 98.4% | pass |
| 3 | 41.9 | 2.38x | 96.4% / 86.1% / 75.9% | pass |
- Recommended depth: 2 (
mtp_depth_winsat D2; D3 passes quality and is still very fast, but D2 is the throughput peak) - Verdict:
mtp_depth_wins;quality_rejected=[],failure_reasons=[] - Hardware: Apple M5 Pro (18 CPU / 20 GPU, 24 GB) · macOS 27.0 arm64 (Apple Silicon, MLX), mtplx 2.9.1, artifact
sha256:f6b4c762... - All depths passed the quality gate. Single-position acceptance at D1/D2 is effectively lossless (98%+), so the speedup comes with no visible quality trade-off at the verified temperature (0.6). D3 also passes but its third-position acceptance (75.9%) costs more than it saves vs D2.
This is the best-verified point among the five MTPLX models in this batch (highest tok/s and highest acceptance at the recommended depth).
Training details (source model)
- Base: Qwen/Qwen3.8-27B — dense 27B, hybrid Gated-DeltaNet / full-attention, 64 layers, 262K context
- Post-training: Abliteration (refusal-direction removal) by OrcaRouter — BF16, no quantization in the source; vision + MTP preserved
- License: Apache 2.0 (inherited)
- Context & modalities: 262K text, vision-language, function calling, reasoning
Source chain
Qwen/Qwen3.8-27B (base)
→ orcarouter/Qwen3.8-27B-Uncensored (BF16 abliterated)
→ this repo (MTPLX 4-bit conversion)
Related in this collection:
- orcarouter/Qwen3.8-27B-Uncensored-MTPLX-3bit — 3-bit MTPLX sibling