language: [en]
license: apache-2.0
pipeline_tag: text-generation
base_model: Qwen/Qwen3.8-27B
base_model_relation: quantized
quantized_by: mlasli
library_name: mlx
tags:
- qwen3.8
- qwen3
- abliterated
- heretic
- uncensored
- decensored
- text-generation
- roleplay
- mlx
- quantization
- apple-silicon
Qwen3.8-27B Heretic-Abliterated (8-bit MLX)
MLX quantization of
mlasli/Qwen3.8-27B-Heretic-Uncensored-BF16
— Qwen/Qwen3.8-27B with its refusal direction removed using
Heretic (single-direction abliteration with an
Optuna-based parameter search). The language backbone is abliterated; all other
capabilities are preserved.
What this is for: the same abliterated 27B weights, quantized to 28.6 GB in MLX format
(group-wise 8-bit, group size 64) for fast Apple-Silicon inference via mlx-lm.
Text-only. The MLX conversion drops the vision tower (only
language_model.*weights are
shipped), so this repo handles text only. For image input use the BF16 safetensors repo
or the GGUF repos with a separate mmproj.
Note on the parameter count badge: this is a 27B-parameter model. Hugging
Face's automatic scanner may show a lower count (e.g. "6B" / "8B") for MLX
quantizations because it counts the packed 32-bit integer elements (which hold
multiple low-bit weights) as individual parameters rather than unpacking them.
The true total inmodel.safetensors.index.jsonis ~27B.
Note on the parameter count badge: this is a 27B-parameter model. Hugging
Face's automatic scanner may show a lower count (e.g. "6B" / "8B") for MLX
quantizations because it counts the packed 32-bit integer elements (which hold
multiple low-bit weights) as individual parameters rather than unpacking them.
The true total inmodel.safetensors.index.jsonis ~27B.
What is Heretic?
Heretic is an abliteration method that removes a
model's safety-aligned refusal direction in one shot (unlike earlier multi-direction
approaches), trading a minimal amount of capability for a large drop in refusals. Its
Optuna search picks the ablation parameters on the Pareto front of
(compliance, first-token KL divergence).
Evaluation
Source: Independent eval (merged model)
- Compliance: 94.0% (harmful-behaviors, Zou et al. refusal detector, 50 prompts)
- Zou 29-substring refusal rate: 6.0%
- First-token KL divergence vs base: 0.0467
A stricter combined keyword detector reported 18.0% refusal, but manual review of the
flagged completions confirmed these are largely false positives (the model answers
directly and uses words like "illegal"/"harmful"/"violent" inside compliant responses).
The Zou number above is the more reliable refusal estimate.
Usage
pip install -U mlx-lm
# one-shot generation
mlx_lm.generate \
--model mlasli/Qwen3.8-27B-Heretic-Abliterated-MLX-8bit \
--prompt "Hi!" --max-tokens 256
# OpenAI-compatible server
mlx_lm.server \
--model mlasli/Qwen3.8-27B-Heretic-Abliterated-MLX-8bit --port 8080
Quantizations
MLX quants (group-wise 6/8-bit, group size 64 — not GGUF Q6_K/Q8_0):
| Quant | Size | Repository |
|---|---|---|
| 8-bit | 28.6 GB | Qwen3.8-27B-Heretic-Abliterated-MLX-8bit |
| 6-bit | 21.9 GB | Qwen3.8-27B-Heretic-Abliterated-MLX-6bit |
GGUF quants (llama.cpp / Ollama) are in separate repos:
| Quant | Repository |
|---|---|
| Q8_0 | Qwen3.8-27B-Heretic-Uncensored-Q8_0-GGUF |
| Q6_K | Qwen3.8-27B-Heretic-Uncensored-Q6_K-GGUF |
| Q4_K_M | Qwen3.8-27B-Heretic-Uncensored-Q4_K_M-GGUF |
Abliteration removes safety alignment. Use responsibly and in accordance with your local
laws and the upstream Apache-2.0 license.
Changelog
v1.0.0 — initial MLX release (2026-08-16)
- Initial MLX quantization of
mlasli/Qwen3.8-27B-Heretic-Uncensored-BF16. - Group-wise 8-bit (group size 64) affine quantization via
mlx_lm.convert.