← back to catalog · registered 2026-08-22 13:56

mlasli/Qwen3.8-27B-Heretic-Abliterated-MLX-8bit

mlasli Qwen 27B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/mlasli%2FQwen3.8-27B-Heretic-Abliterated-MLX-8bit"
Response includes
  • classification m3
  • files 14
  • hub_downloads_all_time 3,224
  • author_summary 23 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · lifetime
3K
1K last 30d - stable
Likes
2
Model age
7w ago
created 2026-08-16
Downloads over time
Now3.7K→from813↑351%
6701.8K2.9K4K813 on Aug 193.7K on Oct 11AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en
Tags
mlx safetensors qwen3_5 qwen3.8 qwen3 abliterated heretic uncensored decensored text-generation roleplay quantization

Related

Total size
26.6 GB
Files
14
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-09-01 02:36

Files by quantization

Auxiliary files 14 files 26.6 GB
model-00005-of-00006.safetensors 4.99 GB aa346f1e download
model-00002-of-00006.safetensors 4.99 GB 2918ba7b download
model-00003-of-00006.safetensors 4.97 GB 0509f4e4 download
model-00001-of-00006.safetensors 4.94 GB e0597a0f download
model-00004-of-00006.safetensors 4.93 GB 4bbf2e76 download
model-00006-of-00006.safetensors 1.80 GB 14c43bc1 download
tokenizer.json 19.1 MB 06b95093 download
model.safetensors.index.json 185 KB 0adfa998 download
chat_template.jinja 8.74 KB c0c686f9 download
README.md 4.30 KB 0ac2b7df download
config.json 4.01 KB b554a687 download
tokenizer_config.json 1.17 KB ffb89ca0 download
generation_config.json 214 B 8b9f95da download
.gitattributes 101 B 0fd17d2f download

README current version from Hugging Face


language: [en]
license: apache-2.0
pipeline_tag: text-generation
base_model: Qwen/Qwen3.8-27B
base_model_relation: quantized
quantized_by: mlasli
library_name: mlx
tags:

  • qwen3.8
  • qwen3
  • abliterated
  • heretic
  • uncensored
  • decensored
  • text-generation
  • roleplay
  • mlx
  • quantization
  • apple-silicon

Qwen3.8-27B Heretic-Abliterated (8-bit MLX)

MLX quantization of
mlasli/Qwen3.8-27B-Heretic-Uncensored-BF16
— Qwen/Qwen3.8-27B with its refusal direction removed using
Heretic (single-direction abliteration with an
Optuna-based parameter search). The language backbone is abliterated; all other
capabilities are preserved.

What this is for: the same abliterated 27B weights, quantized to 28.6 GB in MLX format
(group-wise 8-bit, group size 64) for fast Apple-Silicon inference via mlx-lm.

Text-only. The MLX conversion drops the vision tower (only language_model.* weights are
shipped), so this repo handles text only. For image input use the BF16 safetensors repo
or the GGUF repos with a separate mmproj.

Note on the parameter count badge: this is a 27B-parameter model. Hugging
Face's automatic scanner may show a lower count (e.g. "6B" / "8B") for MLX
quantizations because it counts the packed 32-bit integer elements (which hold
multiple low-bit weights) as individual parameters rather than unpacking them.
The true total in model.safetensors.index.json is ~27B.

Note on the parameter count badge: this is a 27B-parameter model. Hugging
Face's automatic scanner may show a lower count (e.g. "6B" / "8B") for MLX
quantizations because it counts the packed 32-bit integer elements (which hold
multiple low-bit weights) as individual parameters rather than unpacking them.
The true total in model.safetensors.index.json is ~27B.

What is Heretic?

Heretic is an abliteration method that removes a
model's safety-aligned refusal direction in one shot (unlike earlier multi-direction
approaches), trading a minimal amount of capability for a large drop in refusals. Its
Optuna search picks the ablation parameters on the Pareto front of
(compliance, first-token KL divergence).

Evaluation

Source: Independent eval (merged model)

  • Compliance: 94.0% (harmful-behaviors, Zou et al. refusal detector, 50 prompts)
  • Zou 29-substring refusal rate: 6.0%
  • First-token KL divergence vs base: 0.0467

A stricter combined keyword detector reported 18.0% refusal, but manual review of the
flagged completions confirmed these are largely false positives (the model answers
directly and uses words like "illegal"/"harmful"/"violent" inside compliant responses).
The Zou number above is the more reliable refusal estimate.

Usage

pip install -U mlx-lm

# one-shot generation
mlx_lm.generate \
  --model mlasli/Qwen3.8-27B-Heretic-Abliterated-MLX-8bit \
  --prompt "Hi!" --max-tokens 256

# OpenAI-compatible server
mlx_lm.server \
  --model mlasli/Qwen3.8-27B-Heretic-Abliterated-MLX-8bit --port 8080

Quantizations

MLX quants (group-wise 6/8-bit, group size 64 — not GGUF Q6_K/Q8_0):

Quant Size Repository
8-bit 28.6 GB Qwen3.8-27B-Heretic-Abliterated-MLX-8bit
6-bit 21.9 GB Qwen3.8-27B-Heretic-Abliterated-MLX-6bit

GGUF quants (llama.cpp / Ollama) are in separate repos:

Quant Repository
Q8_0 Qwen3.8-27B-Heretic-Uncensored-Q8_0-GGUF
Q6_K Qwen3.8-27B-Heretic-Uncensored-Q6_K-GGUF
Q4_K_M Qwen3.8-27B-Heretic-Uncensored-Q4_K_M-GGUF

Abliteration removes safety alignment. Use responsibly and in accordance with your local
laws and the upstream Apache-2.0 license.

Changelog

v1.0.0 — initial MLX release (2026-08-16)

  • Initial MLX quantization of mlasli/Qwen3.8-27B-Heretic-Uncensored-BF16.
  • Group-wise 8-bit (group size 64) affine quantization via mlx_lm.convert.

README history 4 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-09-01release: v2.1.0 re-ablation MLX quant9c728205.7 KB
    Loading...
  2. 2026-08-17docs: clarify 27B parameter count vs HF badge080ccd54.3 KB
    Loading...
  3. 2026-08-17docs: clarify 27B parameter count vs HF badge13edd7b3.9 KB
    Loading...
  4. 2026-08-16feat: initial MLX quantization (Heretic-Abliterated)80eb7a63.6 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration