← back to catalog · registered 2026-08-24 13:02

barozp/Qwen3.8-27B-Uncensored-MTPLX-4bit

barozp Qwen 27B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/barozp%2FQwen3.8-27B-Uncensored-MTPLX-4bit"
Response includes
  • classification m1
  • files 17
  • hub_downloads_all_time 188
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
188
137 last 30d - active
Likes
5
Model age
6w ago
created 2026-08-24
Downloads over time
Now261→from12↑2,075%
09519128612 on Aug 26261 on Oct 11AugSepOct
Aug 26 → Oct 11 · 47 snapshots · spans 46 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
mlx safetensors qwen3_5 mtplx qwen qwen3 abliterated uncensored vision quantized 4-bit speculative-decoding

Related

Total size
15.7 GB
Files
17
Quantizations
1
Registered
2026-08-24 13:02
Last updated on HF
2026-08-25 22:53

Files by quantization

Auxiliary files 17 files 15.8 GB
model-00002-of-00003.safetensors 4.99 GB ******** download
model-00001-of-00003.safetensors 4.96 GB ******** download
model-00003-of-00003.safetensors 4.14 GB ******** download
model-vision.safetensors 879 MB ******** download
mtp.safetensors 810 MB ******** download
tokenizer.json 19.1 MB ******** download
model.safetensors.index.json 202 KB c6b399a4 download
chat_template.jinja 8.74 KB c0c686f9 download
README.md 6.23 KB 6d13312d download
config.json 4.43 KB c81a1664 download
mtplx_runtime.json 3.94 KB 09114e09 download
.gitattributes 1.53 KB 52373fe2 download
tokenizer_config.json 1.13 KB 8d610519 download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
processor_config.json 367 B ced889f7 download
generation_config.json 202 B 023756cf download

README current version from Hugging Face


license: apache-2.0
base_model: orcarouter/Qwen3.8-27B-Uncensored
base_model_relation: quantized
library_name: mlx
pipeline_tag: image-text-to-text
tags:

  • mlx
  • mtplx
  • qwen
  • qwen3
  • abliterated
  • uncensored
  • vision
  • quantized
  • 4-bit
  • speculative-decoding
  • mtp

Qwen3.8-27B-Uncensored-MTPLX-4bit

MTPLX 4-bit conversion of orcarouter/Qwen3.8-27B-Uncensored — the abliterated (refusal-removed) BF16 build of Qwen3.8-27B — for speculative decoding on Apple Silicon via mtplx. Best performer in this MTPLX collection.

Sibling: orcarouter/Qwen3.8-27B-Uncensored-MTPLX-3bit.

Converted straight from the bf16 safetensors weights (not from any GGUF quant), so there is no dequantize-requantize drift in the chain.

Abliterated / uncensored. This model has had refusal behavior removed. It will comply with requests the base Qwen model would refuse. Use responsibly and in accordance with your local laws and the model's license. No additional content filtering is applied beyond what is in the checkpoint.

Why this release exists

OrcaRouter's uncensored Qwen3.8 preserves the full BF16 base — 27B dense, 64 layers, 262K context, vision + MTP — with refusal ablated. MTPLX 4-bit keeps that checkpoint byte-for-byte, quantizes the trunk to 4-bit affine (group 64), and leaves the MTP head in bf16. At ~16.9 GB on disk it is still comfortable on 36 GB unified-memory Macs, and it delivers the highest verified speedup in the five-model set: 2.56x at depth 2 with near-lossless draft acceptance (98.4% / 98.3%). If you can afford the extra ~3 GB over the 3-bit, this is the one to run.

Available MTPLX conversions

Repo Bits Size (on disk) tok/s (D2) vs AR Use case
orcarouter/Qwen3.8-27B-Uncensored-MTPLX-3bit 3 ~13.6 GB 44.3 2.14x tight unified memory
this repo (4-bit) 4 ~16.9 GB (15.7 GiB) 45.1 2.56x best performer, recommended

Source bf16 remains at orcarouter/Qwen3.8-27B-Uncensored.

Conversion notes

  • Source: orcarouter/Qwen3.8-27B-Uncensored (bf16, 18 shards + vision, orcarouter--Qwen3.8-27B-Uncensored local cache), forge-local via mtplx==2.9.1
  • Recipe: body_bits=4, body_group_size=64, body_mode=affine, mtp_policy=keep_bf16, quantized trunk + bf16 MTP sidecar
  • MTP contract: base_hidden_variant=post_norm, hidden_variant=post_norm, concat_order=embedding_hidden, mtp_position_mode=local, mtp_quant_group_size=64, mtp_quant_mode=affine
  • Output: 3 safetensors shards + model-vision.safetensors (879 MB, 333 tensors, bf16) + mtp.safetensors (810 MB, bf16 sidecar) + model.safetensors.index.json
  • Forged at: 2026-08-24T12:00:11+03:00 on Apple M5 Pro (18 CPU / 20 GPU, 24 GB unified memory) — tuned for 24 GB Macs · macOS 27.0 arm64, mtplx_runtime.json ships in-repo as provenance

Vision

The vision encoder is inside these weights (stored as model-vision.safetensors, auto-loaded) — no separate mmproj to fetch:

mtplx run --model orcarouter/Qwen3.8-27B-Uncensored-MTPLX-4bit \
  --prompt "Describe this image." --image photo.jpg --depth 2

Text-only needs nothing extra.

MTPLX usage

# install
pip install -U mtplx

# single-turn generation (verified optimum is --depth 2)
mtplx run --model orcarouter/Qwen3.8-27B-Uncensored-MTPLX-4bit \
  --prompt "Write a short story about a robot learning to paint." \
  --depth 2 --max-tokens 512

# multimodal
mtplx run --model orcarouter/Qwen3.8-27B-Uncensored-MTPLX-4bit \
  --prompt "Describe this image." --image photo.jpg --depth 2

# OpenAI-compatible local server
mtplx serve --model orcarouter/Qwen3.8-27B-Uncensored-MTPLX-4bit \
  --depth 2 --port 8080
# then: curl http://localhost:8080/v1/chat/completions ...

# explicit sampling (verified defaults)
mtplx run --model orcarouter/Qwen3.8-27B-Uncensored-MTPLX-4bit \
  --prompt "Explain quantum entanglement simply." \
  --depth 2 --temperature 0.6 --top-p 0.95 --top-k 20

Verification results

Verified locally with mtplx forge verify (MLX backend, 512-token budget, quality gate enabled). Depth 0 is plain AR.

Depth tok/s vs AR Acceptance by position Quality
0 (AR) 17.6 1.00x — pass
1 30.3 1.72x 98.4% pass
2 45.1 2.56x 98.4% / 98.4% pass
3 41.9 2.38x 96.4% / 86.1% / 75.9% pass
  • Recommended depth: 2 (mtp_depth_wins at D2; D3 passes quality and is still very fast, but D2 is the throughput peak)
  • Verdict: mtp_depth_wins; quality_rejected=[], failure_reasons=[]
  • Hardware: Apple M5 Pro (18 CPU / 20 GPU, 24 GB) · macOS 27.0 arm64 (Apple Silicon, MLX), mtplx 2.9.1, artifact sha256:f6b4c762...
  • All depths passed the quality gate. Single-position acceptance at D1/D2 is effectively lossless (98%+), so the speedup comes with no visible quality trade-off at the verified temperature (0.6). D3 also passes but its third-position acceptance (75.9%) costs more than it saves vs D2.

This is the best-verified point among the five MTPLX models in this batch (highest tok/s and highest acceptance at the recommended depth).

Training details (source model)

  • Base: Qwen/Qwen3.8-27B — dense 27B, hybrid Gated-DeltaNet / full-attention, 64 layers, 262K context
  • Post-training: Abliteration (refusal-direction removal) by OrcaRouter — BF16, no quantization in the source; vision + MTP preserved
  • License: Apache 2.0 (inherited)
  • Context & modalities: 262K text, vision-language, function calling, reasoning

Source chain

Qwen/Qwen3.8-27B (base)
→ orcarouter/Qwen3.8-27B-Uncensored (BF16 abliterated)
→ this repo (MTPLX 4-bit conversion)

Related in this collection:

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration