← back to catalog · registered 2026-08-22 13:56

onchainengineer/Qwen3.8-27B-Uncensored-MLX-BF16

onchainengineer Qwen 27B multimodal
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/onchainengineer%2FQwen3.8-27B-Uncensored-MLX-BF16"
Response includes
  • classification m1
  • files 24
  • hub_downloads_all_time 753
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
753
209 last 30d - stable
Likes
0
Model age
7w ago
created 2026-08-21
Downloads over time
Now820→from54↑1,419%
1630960389754 on Aug 19820 on Oct 11AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
mlx safetensors qwen3_5 qwen3.8 bfloat16 bf16 vision-language reasoning uncensored abliterated mtp apple-silicon

Related

Total size
51.0 GB
Files
24
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-23 05:00

Files by quantization

Auxiliary files 24 files 51.0 GB
model-00009-of-00011.safetensors 4.98 GB 0fcb919f download
model-00005-of-00011.safetensors 4.98 GB ba66348e download
model-00004-of-00011.safetensors 4.96 GB f7a941b3 download
model-00006-of-00011.safetensors 4.96 GB 20ff7a0d download
model-00003-of-00011.safetensors 4.96 GB d658aa34 download
model-00007-of-00011.safetensors 4.96 GB f71531fd download
model-00008-of-00011.safetensors 4.96 GB 4f803bf7 download
model-00002-of-00011.safetensors 4.96 GB b9533ddc download
model-00001-of-00011.safetensors 4.87 GB 0451f0af download
model-00010-of-00011.safetensors 4.03 GB 5f411081 download
model-00011-of-00011.safetensors 2.37 GB 1219c9ea download
tokenizer.json 19.1 MB 06b95093 download
vocab.json 6.41 MB 0aa0ce06 download
model.safetensors.index.json 113 KB 3425801b download
LICENSE 11.3 KB f938136e download
chat_template.jinja 8.74 KB c0c686f9 download
README.md 5.04 KB 89a4c83d download
config.json 4.62 KB 5abfcb09 download
tokenizer_config.json 1.14 KB 1d134cd2 download
processor_config.json 991 B 8f29fe38 download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
generation_config.json 202 B 023756cf download
.gitattributes 156 B 2070e5cc download

README current version from Hugging Face


license: apache-2.0
library_name: mlx
pipeline_tag: image-text-to-text
base_model:

  • Qwen/Qwen3.8-27B
  • orcarouter/Qwen3.8-27B-Uncensored
    tags:
  • mlx
  • qwen3_5
  • qwen3.8
  • bfloat16
  • bf16
  • vision-language
  • reasoning
  • uncensored
  • abliterated
  • mtp
  • apple-silicon

Qwen3.8-27B-Uncensored MLX BF16

Full-precision BF16 MLX conversion of
orcarouter/Qwen3.8-27B-Uncensored,
including native vision-language support and a separately packaged native MTP drafter.

This is a community conversion for Apple silicon. It is not an official Qwen or OrcaRouter release.
No quantization, fine-tuning, pruning, or additional weight modification was applied during this
conversion. The upstream model is an abliterated (refusal-direction removed) derivative of
Qwen/Qwen3.8-27B; consult its model card for evaluation details and limitations.

What is included

Component Format Tensors Notes
Main model MLX BF16 1,184 Text/reasoning model plus all 333 vision tensors
mtp/ drafter MLX BF16 15 Native Qwen3.5/Qwen3.8 MTP head for speculative decoding
Total MLX BF16 1,199 Matches the complete upstream tensor inventory

The model has a 262,144-token configured context window. Real usable context depends on memory,
KV-cache settings, prompt modality, and the MLX-VLM version.

Requirements

  • Apple silicon Mac
  • Current mlx-vlm with Qwen3.5/Qwen3.8 and MTP support
  • Approximately 55 GB for weights, plus working memory and KV cache

One installation option:

uv tool install mlx-vlm --with jinja2

Text generation

mlx_vlm.generate \
  --model onchainengineer/Qwen3.8-27B-Uncensored-MLX-BF16 \
  --prompt "Explain why the sky is blue." \
  --thinking-mode enabled \
  --max-tokens 512

For deterministic output, add --temperature 0. For maximum response quality, leave weights and
KV cache unquantized; long-context workloads may optionally trade fidelity for memory with MLX-VLM's
KV-cache controls.

Native MTP speculative decoding

hf download onchainengineer/Qwen3.8-27B-Uncensored-MLX-BF16 \
  --local-dir ./Qwen3.8-27B-Uncensored-MLX-BF16

mlx_vlm.generate \
  --model ./Qwen3.8-27B-Uncensored-MLX-BF16 \
  --draft-model ./Qwen3.8-27B-Uncensored-MLX-BF16/mtp \
  --draft-kind mtp \
  --prompt "Write a clear technical explanation of speculative decoding." \
  --thinking-mode enabled \
  --max-tokens 512

Point --draft-model at the snapshot's mtp subdirectory. MTP can improve longer
generations, but very short outputs may be slower because drafter setup dominates.

Vision-language generation

mlx_vlm.generate \
  --model onchainengineer/Qwen3.8-27B-Uncensored-MLX-BF16 \
  --image /absolute/path/to/image.png \
  --prompt "Describe this image precisely." \
  --max-tokens 256

Conversion provenance

  • Source: orcarouter/Qwen3.8-27B-Uncensored
  • Source revision: 9878936be9458522b5aeed0e13476bb8426f57f0
  • Source inventory: 18 safetensors shards, 1,199 tensors, all BF16
  • Converter: mlx-vlm 0.6.15
  • MLX: 0.32.1
  • Transformers: 5.15.1
  • Conversion target: bfloat16, without quantization

Main conversion:

mlx_vlm.convert \
  --hf-path /path/to/orcarouter-Qwen3.8-27B-Uncensored \
  --mlx-path /path/to/Qwen3.8-27B-Uncensored-MLX-BF16 \
  --dtype bfloat16

MTP extraction:

python -m mlx_vlm.speculative.drafters.qwen3_5_mtp.split \
  --model /path/to/orcarouter-Qwen3.8-27B-Uncensored \
  --output /path/to/Qwen3.8-27B-Uncensored-MLX-BF16/mtp

Validation

The published artifact was checked locally on a 256 GB M3 Ultra Mac Studio:

  • all 1,184 main tensors are BF16; all 15 MTP tensors are BF16
  • no model quantization metadata is present
  • deterministic text generation loaded and returned the expected answer
  • reasoning-style generation loaded and produced a coherent mathematical explanation
  • vision inference identified the subject and dominant colors in a test image
  • native MTP loaded successfully and achieved 100% drafted-token acceptance on the deterministic smoke test
  • measured peak unified memory was approximately 55.0 GB for text, 55.8 GB for vision, and 56.4 GB with MTP

These are functional smoke tests, not a claim of comprehensive benchmark parity. Performance varies by
hardware, prompt, software version, thermal state, and generation settings.

Safety and intended use

This checkpoint intentionally reduces refusal behavior. That does not make every generated answer
accurate, safe, legal, private, or appropriate. Users are responsible for evaluating outputs and for
complying with applicable law and the Apache-2.0 license. Do not use it to facilitate harm, unauthorized
access, privacy violations, fraud, or other unlawful activity. Apply appropriate safeguards for any
deployment exposed to other users.

License and attribution

Apache License 2.0. See LICENSE. This conversion retains attribution to Qwen and OrcaRouter; review
the upstream model cards for provenance, training/modification details, evaluations, and known limitations.

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-23Add StrongREJECT generation cap accounting0deb0927.6 KB
    Loading...
  2. 2026-08-23Add comprehensive end-to-end benchmark resultsfaa23057.5 KB
    Loading...
  3. 2026-08-21Add full BF16 MLX conversion with native vision and MTP1ea993e5 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration