← back to catalog · registered 2026-08-25 07:02

mlx-community/Qwen3.8-27B-Uncensored-OptiQ-4bit

mlx-community Qwen 27B second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/mlx-community%2FQwen3.8-27B-Uncensored-OptiQ-4bit"
Response includes
  • classification m-uncensored
  • files 12
  • hub_downloads_all_time 16,592
  • author_summary 207 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
17K
Likes
33
Model age
6w ago
created 2026-08-25
Downloads over time
Now22.8K→from354↑6,331%
08.3K16.7K25K354 on Aug 2622.8K on Oct 11AugSepOct
Aug 26 → Oct 11 · 47 snapshots · spans 46 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en
Tags
mlx safetensors qwen3_5 optiq quantized 4bit mixed-precision qwen3.8 reasoning long-context mtp multi-token-prediction

Related

Total size
18.1 GB
Files
12
Quantizations
1
Registered
2026-08-25 07:02
Last updated on HF
2026-09-16 13:47

Files by quantization

Auxiliary files 12 files 18.1 GB
model-00003-of-00004.safetensors 4.99 GB 1f606c81 download
model-00002-of-00004.safetensors 4.97 GB 684a3d30 download
model-00001-of-00004.safetensors 4.96 GB 56740026 download
model-00004-of-00004.safetensors 3.17 GB 7d2c9ac2 download
tokenizer.json 19.1 MB 06b95093 download
model.safetensors.index.json 185 KB e1bc204b download
config.json 106 KB e20165e0 download
chat_template.jinja 8.74 KB c0c686f9 download
README.md 2.60 KB c750d757 download
.gitattributes 1.53 KB 52373fe2 download
tokenizer_config.json 1.13 KB 8d610519 download
generation_config.json 202 B 023756cf download

README current version from Hugging Face


license: apache-2.0
base_model: orcarouter/Qwen3.8-27B-Uncensored
base_model_relation: quantized
library_name: mlx
pipeline_tag: text-generation
language:

  • en
    tags:
  • mlx
  • optiq
  • quantized
  • 4bit
  • mixed-precision
  • qwen3.8
  • reasoning
  • long-context
  • mtp
  • multi-token-prediction
  • function-calling
  • tool-use
  • agentic
  • uncensored
  • apple-silicon

mlx-community/Qwen3.8-27B-Uncensored-OptiQ-4bit

Built with mlx-optiq, the MLX-native toolkit to quantize, fine-tune, and serve LLMs locally on Apple Silicon, no PyTorch and no cloud. All OptiQ quants · Docs

OptiQ mixed-precision quant of orcarouter/Qwen3.8-27B-Uncensored, a Qwen3.8-family reasoning model with a bundled MTP speculation head. 19 GB on disk.

What it is

Property Value
Base orcarouter/Qwen3.8-27B-Uncensored (Qwen3.8, 27B)
Method OptiQ mixed-precision, per-layer 4/8-bit
Bit allocation Reused from the Qwen3.8-27B OptiQ recipe: the architecture is identical, so the per-layer sensitivity ranking transfers directly and no per-model sweep is needed
Layer split 237 components at 4-bit, 261 at 8-bit
Group size 64
On disk 19 GB
MTP Speculation head preserved in optiq/mtp.safetensors for faster decode via optiq serve --draft-model

Following the naming llama.cpp uses for its mixed quants, the "4bit" label denotes the family, not the weighted average.

Run it

Qwen3.8 and the MTP sidecar register through OptiQ, so import optiq once before loading:

pip install "mlx-optiq>=0.4.27"
import optiq  # registers the arch + MTP sidecar
from mlx_lm import load, generate

model, tok = load("mlx-community/Qwen3.8-27B-Uncensored-OptiQ-4bit")
prompt = tok.apply_chat_template(
    [{"role": "user", "content": "Explain mixed-precision quantization in two sentences."}],
    tokenize=False, add_generation_prompt=True,
)
print(generate(model, tok, prompt=prompt, max_tokens=400))

For an OpenAI- and Anthropic-compatible endpoint with mixed-precision KV cache:

optiq serve --model mlx-community/Qwen3.8-27B-Uncensored-OptiQ-4bit

This is a reasoning model, so give it a generous token budget.

Links

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-25OptiQ mixed-precision 4-bit quant5dda6e12.6 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration