license: apache-2.0
library_name: mlx
pipeline_tag: image-text-to-text
base_model:
- Qwen/Qwen3.8-27B
- orcarouter/Qwen3.8-27B-Uncensored
tags: - mlx
- qwen3_5
- qwen3.8
- bfloat16
- bf16
- vision-language
- reasoning
- uncensored
- abliterated
- mtp
- apple-silicon
Qwen3.8-27B-Uncensored MLX BF16
Full-precision BF16 MLX conversion oforcarouter/Qwen3.8-27B-Uncensored,
including native vision-language support and a separately packaged native MTP drafter.
This is a community conversion for Apple silicon. It is not an official Qwen or OrcaRouter release.
No quantization, fine-tuning, pruning, or additional weight modification was applied during this
conversion. The upstream model is an abliterated (refusal-direction removed) derivative ofQwen/Qwen3.8-27B; consult its model card for evaluation details and limitations.
What is included
| Component | Format | Tensors | Notes |
|---|---|---|---|
| Main model | MLX BF16 | 1,184 | Text/reasoning model plus all 333 vision tensors |
mtp/ drafter |
MLX BF16 | 15 | Native Qwen3.5/Qwen3.8 MTP head for speculative decoding |
| Total | MLX BF16 | 1,199 | Matches the complete upstream tensor inventory |
The model has a 262,144-token configured context window. Real usable context depends on memory,
KV-cache settings, prompt modality, and the MLX-VLM version.
Requirements
- Apple silicon Mac
- Current
mlx-vlmwith Qwen3.5/Qwen3.8 and MTP support - Approximately 55 GB for weights, plus working memory and KV cache
One installation option:
uv tool install mlx-vlm --with jinja2
Text generation
mlx_vlm.generate \
--model onchainengineer/Qwen3.8-27B-Uncensored-MLX-BF16 \
--prompt "Explain why the sky is blue." \
--thinking-mode enabled \
--max-tokens 512
For deterministic output, add --temperature 0. For maximum response quality, leave weights and
KV cache unquantized; long-context workloads may optionally trade fidelity for memory with MLX-VLM's
KV-cache controls.
Native MTP speculative decoding
hf download onchainengineer/Qwen3.8-27B-Uncensored-MLX-BF16 \
--local-dir ./Qwen3.8-27B-Uncensored-MLX-BF16
mlx_vlm.generate \
--model ./Qwen3.8-27B-Uncensored-MLX-BF16 \
--draft-model ./Qwen3.8-27B-Uncensored-MLX-BF16/mtp \
--draft-kind mtp \
--prompt "Write a clear technical explanation of speculative decoding." \
--thinking-mode enabled \
--max-tokens 512
Point --draft-model at the snapshot's mtp subdirectory. MTP can improve longer
generations, but very short outputs may be slower because drafter setup dominates.
Vision-language generation
mlx_vlm.generate \
--model onchainengineer/Qwen3.8-27B-Uncensored-MLX-BF16 \
--image /absolute/path/to/image.png \
--prompt "Describe this image precisely." \
--max-tokens 256
Conversion provenance
- Source:
orcarouter/Qwen3.8-27B-Uncensored - Source revision:
9878936be9458522b5aeed0e13476bb8426f57f0 - Source inventory: 18 safetensors shards, 1,199 tensors, all BF16
- Converter:
mlx-vlm 0.6.15 - MLX:
0.32.1 - Transformers:
5.15.1 - Conversion target:
bfloat16, without quantization
Main conversion:
mlx_vlm.convert \
--hf-path /path/to/orcarouter-Qwen3.8-27B-Uncensored \
--mlx-path /path/to/Qwen3.8-27B-Uncensored-MLX-BF16 \
--dtype bfloat16
MTP extraction:
python -m mlx_vlm.speculative.drafters.qwen3_5_mtp.split \
--model /path/to/orcarouter-Qwen3.8-27B-Uncensored \
--output /path/to/Qwen3.8-27B-Uncensored-MLX-BF16/mtp
Validation
The published artifact was checked locally on a 256 GB M3 Ultra Mac Studio:
- all 1,184 main tensors are BF16; all 15 MTP tensors are BF16
- no model quantization metadata is present
- deterministic text generation loaded and returned the expected answer
- reasoning-style generation loaded and produced a coherent mathematical explanation
- vision inference identified the subject and dominant colors in a test image
- native MTP loaded successfully and achieved 100% drafted-token acceptance on the deterministic smoke test
- measured peak unified memory was approximately 55.0 GB for text, 55.8 GB for vision, and 56.4 GB with MTP
These are functional smoke tests, not a claim of comprehensive benchmark parity. Performance varies by
hardware, prompt, software version, thermal state, and generation settings.
Safety and intended use
This checkpoint intentionally reduces refusal behavior. That does not make every generated answer
accurate, safe, legal, private, or appropriate. Users are responsible for evaluating outputs and for
complying with applicable law and the Apache-2.0 license. Do not use it to facilitate harm, unauthorized
access, privacy violations, fraud, or other unlawful activity. Apply appropriate safeguards for any
deployment exposed to other users.
License and attribution
Apache License 2.0. See LICENSE. This conversion retains attribution to Qwen and OrcaRouter; review
the upstream model cards for provenance, training/modification details, evaluations, and known limitations.