← back to catalog · registered 2026-08-22 13:56

jialinyyzz/Qwen3.8-27B-abliterated-MLX-8bit

jialinyyzz Qwen 27B multimodal
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/jialinyyzz%2FQwen3.8-27B-abliterated-MLX-8bit"
Response includes
  • classification m1
  • files 19
  • hub_downloads_all_time 543
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
543
265 last 30d - stable
Likes
0
Model age
7w ago
created 2026-08-16
Downloads over time
Now650→from187↑248%
164341519696187 on Aug 19650 on Oct 11AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
mlx safetensors qwen3_5 abliterated uncensored heretic vision-language image-text-to-text conversational en base_model:Qwen/Qwen3.8-27B base_model:quantized:Qwen/Qwen3.8-27B

Related

Total size
27.5 GB
Files
19
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-16 21:26

Files by quantization

Auxiliary files 19 files 27.5 GB
model-00003-of-00006.safetensors 4.99 GB 9bd1d441 download
model-00002-of-00006.safetensors 4.99 GB c03deb59 download
model-00004-of-00006.safetensors 4.97 GB fdb11b4f download
model-00001-of-00006.safetensors 4.95 GB 71746e3d download
model-00005-of-00006.safetensors 4.93 GB 3b570493 download
model-00006-of-00006.safetensors 2.65 GB b8b9e238 download
tokenizer.json 19.1 MB 06b95093 download
vocab.json 6.41 MB 0aa0ce06 download
model.safetensors.index.json 213 KB 2201d720 download
chat_template.jinja 8.74 KB c0c686f9 download
README.md 5.10 KB 12ac698c download
config.json 4.91 KB d2da8a98 download
.gitattributes 1.53 KB 52373fe2 download
tokenizer_config.json 1.14 KB 1d134cd2 download
processor_config.json 991 B 8f29fe38 download
abliteration_info.json 625 B 923ff07e download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
generation_config.json 214 B 0bc3addd download

README current version from Hugging Face


license: apache-2.0
base_model: Qwen/Qwen3.8-27B
language: en
library_name: mlx
pipeline_tag: image-text-to-text
tags:

  • mlx
  • qwen3_5
  • abliterated
  • uncensored
  • heretic
  • vision-language

Qwen3.8-27B abliterated — MLX 8-bit

An abliterated (decensored) conversion of Qwen/Qwen3.8-27B
to Apple MLX (8-bit), with the vision tower preserved. Refusal behavior has been suppressed via
directional ablation; the model's built-in safety guardrails are largely removed.

⚠️ Experimental / research use only. This model will attempt to answer requests that the
original model would refuse. It ships without the base model's safety behavior. You are
responsible for how you use it and for complying with the base model's Apache-2.0 license and
applicable law.

What this is

  • Base: Qwen/Qwen3.8-27B (hybrid Gated-DeltaNet + attention VLM; 64 text layers + 27-layer vision tower)
  • Method: Heretic MPOA (projected abliteration), selected from a
    300-trial Optuna search (variant wide-177)
  • Format: MLX affine 8-bit (group size 64), 8.627 bits/weight, vision tower kept
  • One file, two uses: runs under mlx-vlm (image/video understanding) and under mlx-dspark
    (text + DSpark speculative decoding). Vision tensors (333 keys) are on disk; the text/DSpark path
    ignores them at load.

Abliteration recipe (wide-177)

parameter value
refusal direction single direction from layer ≈28 (direction_index 28.46)
attn.o_proj gentle: 0.83→0.64, layers 6–63 (attention pathway barely touched)
mlp.down_proj strong & uniform: 1.57→1.52, all 64 layers
KL divergence (orig ‖ abliterated) 0.056

Refusals are removed mainly through the MLP pathway across every layer while the attention
pathway is left nearly intact
— which is what preserves reasoning while suppressing refusals.

Refusal reduction

Heretic keyword-refusal scorer on mlabonne/harmful_behaviors test[:100]:

refusals / 100
base model 98
this model (w177) 11

This is an automated keyword metric with known false positives (a compliant answer containing
"illegal"/"harmful" is scored as a refusal) and false negatives (a refusal phrased without those
words is missed). Treat it as an indicator, not ground truth.

Capability evaluation

Original vs. this abliteration (lm-eval-harness loglikelihood; GSM8K generative CoT):

benchmark original w177 Δ
ARC-Challenge 0.570 0.567 −0.003
HellaSwag 0.750 0.750 0.000
Winogrande 0.753 0.770 +0.017
OpenBookQA 0.477 0.470 −0.007
GSM8K (CoT) 0.76 0.78 +0.02

No measurable capability loss on these five benchmarks (all deltas within ~±0.03 sampling noise).
This is not a claim of "lossless": not evaluated — code generation, agentic/tool-use
(SWE-bench etc.), GPQA, long-context, multilingual, and actual safety behavior; these may differ.
KL 0.056 means the harmless-prompt output distribution is close to the original, but only the
harmless first-token distribution and the benchmarks above were checked.

Performance (Apple M5 Max, 128 GB)

Median of 3 trials, mlx-dspark benchmark, across chat/code/math prompts:

config tok/s speedup
baseline 17.8 —
DSpark (--caps auto) 26.9 1.51×

DSpark helps most on structured content — 1.84× on math, 1.65× on code, ~1.0× on chat. Use
--mode dspark; cap=auto is optimal (larger fixed caps are slower on this hybrid because rejected
drafts must rebuild the Gated-DeltaNet recurrent state). DSpark auto-resolves its drafter
(RadixArk/Qwen3.8-27B-DSpark) from the model basename — no --drafter flag needed. Single lucky
prompts can hit ~2×, but the sustained median is ~1.5×. For maximum tok/s overall, the 4-bit variant
is faster in absolute terms (see that repo).

Usage

DSpark speculative decoding needs a second weight — the ~1.36B drafter. mlx-dspark
auto-downloads it from RadixArk/Qwen3.8-27B-DSpark (no manual assembly), so the simple command
just works:

pip install mlx-dspark
mlx-dspark generate --model ./Qwen3.8-27B-abliterated-MLX-8bit --mode dspark \
  --prompt "..." --max-new-tokens 512

Fully self-contained (no upstream dependency) — point at the bundled drafter mirror
Qwen3.8-27B-DSpark-drafter:

mlx-dspark generate --model ./Qwen3.8-27B-abliterated-MLX-8bit \
  --drafter ./Qwen3.8-27B-DSpark-drafter --mode dspark --prompt "..."

Image / video (vision tower preserved, no drafter needed):

pip install mlx-vlm
python -m mlx_vlm generate --model ./Qwen3.8-27B-abliterated-MLX-8bit \
  --image photo.jpg --prompt "Describe this image."

Provenance & license

  • Derived from Qwen/Qwen3.8-27B under Apache-2.0; this derivative inherits Apache-2.0.
  • Abliteration performed with Heretic (AGPL-3.0 tool; does not affect the model license).
  • No additional training; weights edited by directional ablation only.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-16Upload folder using huggingface_hubdddccfe5.1 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration