license: apache-2.0
base_model: orcarouter/Qwen3.8-27B-Uncensored
base_model_relation: quantized
pipeline_tag: image-text-to-text
library_name: mlx
language:
- en
- zh
- ja
tags: - mlx
- optiq
- mixed-precision
- quantized
- 4-bit
- 8-bit
- uncensored
- abliterated
- qwen3.8
- vision
- mtp
- image-text-to-text
Qwen3.8-27B-Uncensored-OptiQ-4bit (MLX)
A sensitivity-aware mixed 4/8-bit OptiQ MLX quant oforcarouter/Qwen3.8-27B-Uncensored,
the abliterated (refusal-removed) BF16 build of Qwen's Qwen/Qwen3.8-27B.
Unlike builds that reuse the base Qwen sensitivity ranking, this quant takes its
per-layer KL-divergence sensitivity directly from the Orca uncensored weights
against the BF16 source (not a uniform-4-bit reference), then allocates 8-bit to
the sensitive layers and 4-bit to the rest. The native vision tower and MTP head
are preserved.
- ~19.2 GB on disk, 5.13 bits/weight (target 5.0)
- 497 quantized tensors: 220 @ 8-bit, 277 @ 4-bit (group size 64)
- Vision tower kept at BF16 in a sidecar (
optiq/optiq_vision.safetensors) - MTP head preserved as a sidecar (
optiq/mtp.safetensors) for speculative decoding - 262,144-token context, hybrid attention (48 linear + 16 full attention layers)
Benchmark — Capability Score 88.86
Full 6-benchmark OptiQ suite (reasoning mode), same Mac (M3 Ultra 96 GB), greedy decode.
| Benchmark | Score |
|---|---|
| MMLU (generative, reasoning) | 90.1% |
| GSM8K (1000) | 97.0% |
| IFEval (strict) | 86.3% |
| BFCL-V3 (simple) | 90.0% |
| HumanEval (pass@1) | 95.7% |
| HashHop | 74.0% |
| Capability Score | 88.86 |
For context, the reference mlx-community/Qwen3.8-27B-OptiQ-4bit reports 87.98.
KL divergence vs the BF16 source (64 prompts × 256 tokens): mean 0.101, median 0.008, p95 0.300.
Quantization recipe
optiq convert orcarouter/Qwen3.8-27B-Uncensored \
--method optiq \
--target-bpw 5.0 \
--candidate-bits 4,8 \
--group-size 64 \
--reference bf16 \
--calibration-mix optiq \
--n-calibration 40
mlx-optiq0.5.19- Reference: BF16 (the full BF16 source fits in 96 GB RAM, so sensitivities are measured against the true base rather than a 4-bit proxy)
- Calibration: OptiQ 6-domain mix, 40 samples
Optional: mixed-precision KV cache
A sensitivity-measured per-layer KV config is included in the source build notes and
reproduces 5.00 average KV bits (12 layers @ 4-bit + 4 layers @ 8-bit across the 16
full-attention layers; linear-attention layers are skipped):
optiq kv-cache ./Qwen3.8-27B-Uncensored-OptiQ-4bit --target-bits 5.0 --candidate-bits 4,8
optiq serve --model ./Qwen3.8-27B-Uncensored-OptiQ-4bit --kv-config ./kv_config.json
Usage
Text (stock mlx-lm)
pip install mlx-lm
from mlx_lm import load, generate
model, tokenizer = load("j-llm/Qwen3.8-27B-Uncensored-OptiQ-4bit")
print(generate(model, tokenizer, "What is 12*13?", max_tokens=64))
Image + text (stock mlx-vlm)
pip install mlx-vlm
from mlx_vlm import load, generate
from mlx_vlm.prompt_utils import apply_chat_template
from mlx_vlm.utils import load_config
model, processor = load("j-llm/Qwen3.8-27B-Uncensored-OptiQ-4bit")
config = load_config("j-llm/Qwen3.8-27B-Uncensored-OptiQ-4bit")
prompt = apply_chat_template(processor, config, "Describe this image.", num_images=1)
print(generate(model, processor, prompt, image=["test.png"], max_tokens=256).text)
The vision tower lives in the optiq/ sidecar and is restored automatically by both
OptiQ and current mlx-vlm. preprocessor_config.json is included so library and
app scanners (e.g. MLXBar) detect the model as image+text.
OptiQ serving (OpenAI + Anthropic API, MTP, KV)
pip install mlx-optiq
optiq serve --model j-llm/Qwen3.8-27B-Uncensored-OptiQ-4bit --mtp
Serving note:
optiq serve's vision path currently requiresmlx-lm==0.31.x
(_serve_single(self, request));mlx-lm0.32 changed that internal signature and
breaks the vision patch. Pinmlx-lm<0.32in the serving environment.
Model details
| Field | Value |
|---|---|
| Base | orcarouter/Qwen3.8-27B-Uncensored @ 8cb32d72080f6a47bf34dc5adf8067daa6c63a31 |
| Architecture | Qwen3_5ForConditionalGeneration — 64 layers, hidden 5120, hybrid Gated DeltaNet (48 linear + 16 full attention, interval 4) |
| Type | Abliterated (refusal-direction removed), then OptiQ mixed-precision quantized |
| Quantization | OptiQ mixed 4/8-bit, group size 64, affine; achieved 5.13 BPW |
| Vision | BF16 sidecar, 333 tensors |
| MTP | Preserved sidecar, 29 tensors (projections 4-bit, norms + mtp.fc BF16) |
| Context | 262,144 tokens |
| Size | ~19.2 GB |
Safety
This is an abliterated / uncensored model: safety refusals have been removed from
the base weights. It is shared for research and red-teaming. Run it behind access
controls and do not expose it publicly without moderation. The upstream orcarouter
release carries the same warning.
Provenance
- Source:
orcarouter/Qwen3.8-27B-Uncensored(BF16, 55.6 GB), revision8cb32d72… - Quantized with
mlx-optiq0.5.19 on Apple M3 Ultra (96 GB), BF16 reference, 40-sample calibration.