← back to catalog · registered 2026-08-22 13:56

wfiedler/Qwen3.8-27B-ABLITERATED-mlx-q3km

wfiedler Qwen 27B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/wfiedler%2FQwen3.8-27B-ABLITERATED-mlx-q3km"
Response includes
  • classification m1
  • files 15
  • hub_downloads_all_time 1,283
  • author_summary 3 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
1K
256 last 30d - stable
Likes
1
Model age
8w ago
created 2026-08-15
Downloads over time
Now1.4K→from689↑102%
6549231.2K1.5K689 on Aug 191.4K on Oct 11AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
mlx safetensors qwen3_5 qwen3.8 qwen 27b dense vision-language multimodal abliterated quantized tool-calling

Related

Total size
12.5 GB
Files
15
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-15 07:21

Files by quantization

Auxiliary files 15 files 12.5 GB
model-00002-of-00003.safetensors 5.00 GB 50f0e350 download
model-00001-of-00003.safetensors 4.99 GB 8680d521 download
model-00003-of-00003.safetensors 2.47 GB 7d41816d download
tokenizer.json 19.1 MB 06b95093 download
vocab.json 6.41 MB 0aa0ce06 download
model.safetensors.index.json 213 KB 496197cf download
config.json 126 KB c3016a3e download
chat_template.jinja 10.1 KB 764dd790 download
README.md 4.54 KB 2308c01d download
.gitattributes 1.53 KB 52373fe2 download
tokenizer_config.json 1.14 KB 1d134cd2 download
processor_config.json 991 B 8f29fe38 download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
generation_config.json 202 B 023756cf download

README current version from Hugging Face


license: apache-2.0
base_model: Blackfrost-AI/Qwen3.8-27B-ABLITERATED-BF16
tags:

  • qwen3.8
  • qwen
  • 27b
  • dense
  • vision-language
  • multimodal
  • abliterated
  • quantized
  • tool-calling
  • long-context
  • mlx
    pipeline_tag: image-text-to-text
    library_name: mlx

Qwen3.8-27B-ABLITERATED-mlx-q3km

MLX conversion of Blackfrost-AI/Qwen3.8-27B-ABLITERATED-BF16,
quantized with the mixed_3_4 recipe as an equivalent to GGUF Q3_K_M.

  • 3.910 bits per weight, 13.4 GB across 3 safetensors shards
  • Architecture qwen3_5 (Qwen3_5ForConditionalGeneration), multimodal, hybrid
    linear/full attention every 4th layer
  • Requires mlx-vlm >= 0.6.13 (earlier versions have no qwen3_5 module)

Why mixed_3_4 and not --q-bits 3

MLX's mixed-bit predicates are a deliberate port of llama.cpp's _K_M scheme,
not an approximation. Both give more bits to v_proj and down_proj in the
first eighth of layers, the last eighth, and every third layer between, plus a
high-bit lm_head. mixed_3_4 is 3-bit base with 4-bit on those sensitive
tensors, which is exactly Q3_K_M.

The measured result confirms it: 3.910 bpw here, against 3.91 bpw for the
official GGUF Q3_K_M (13.30 GB for ~27.2B params). A plain --q-bits 3 would
land near 3.5 bpw and would not match.

Note that mlx-vlm never quantizes the vision encoder at any level, so the
vision tower stays bf16.

Conversion

mlx_vlm.convert \
  --hf-path Blackfrost-AI/Qwen3.8-27B-ABLITERATED-BF16 \
  --mlx-path ./Qwen3.8-27B-ABLITERATED-mlx-q3km \
  --quantize --quant-predicate mixed_3_4

Usage

pip install -U mlx-vlm

# text
mlx_vlm.generate --model wfiedler/Qwen3.8-27B-ABLITERATED-mlx-q3km \
  --prompt "Explain unified memory in two sentences." --max-tokens 256

# vision
mlx_vlm.generate --model wfiedler/Qwen3.8-27B-ABLITERATED-mlx-q3km \
  --prompt "Read this receipt and list every line item with its price." \
  --image receipt.png --max-tokens 400

Quality vs the 5-bit conversion

Benchmarked against the q5 build (5.678 bpw, 18 GB) of the same base model, on
images with exact ground truth, 12 tasks x 3 trials at temperature 0.7.

task tests q5 this model
receipt OCR, 14 fields 14/14 x3 14/14 x3
chart bar values 6/6 x3 6/6 x3
table 5x5 numeric cells 8/8 x3 8/8 x3
small text small-font rare tokens 6/6 x3 6/6 x3
shapes (10) count 3 groups 6/6 x3 6/6 x3
shapes (20) count 4 scattered groups 3/3 exact 0/3 exact
photos species ID 3/3 3/3
signs real-world OCR 2/3 + 1 loop 3/3
logic / recall text reasoning 3/3, 3/3 3/3, 3/3

On reading images the two are indistinguishable. Every OCR-shaped task ties
perfectly: all receipt fields, all chart values, all table cells, every rare
token in the small print.

The one real gap is counting many objects. Given 20 scattered shapes in four
colour/type groups, q5 answered correctly three times out of three; this model
answered 17, 19 and 21, never correct. The difference is visible in the output:
q5 enumerates shape by shape before totalling, while the 3-bit build emits a
terse number. The low-bit quant lost the habit of decomposing before answering.

Neither model does reliable mental arithmetic without a scratchpad; both compute
a four-item subtotal wrongly at temperature 0.7.

q5 this model
on disk 18 GB 12 GB
bits per weight 5.678 3.910
weights resident 20.68 GB 14.11 GB
peak during image generation 23.24 GB 16.67 GB
generation 12.9 tok/s 18.9 tok/s

Measured on an Apple Silicon Mac with 64 GB unified memory, each model in its
own process. (GenerationResult.peak_memory is a process-global high-water mark
that does not reset between models, so comparing two in one process makes the
second inherit the first's peak.)

Use this build for OCR, document and chart work — 46% faster and 6.6 GB
lighter for no measurable loss. Use the 5-bit build when the task involves
counting or enumerating things in a scene.

Provenance and caveats

This is an abliterated model: the base has had refusal behaviour removed by its
authors. It will comply with requests a stock instruct model would decline.
Evaluate it before using it anywhere user-facing.

Quantization is lossy. The benchmark above covers a narrow slice of behaviour
and says nothing about long-context, tool-calling, or code quality at 3 bits.
All credit for the model itself goes to Blackfrost-AI and the Qwen team.

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-15Fix shard count in model card180b9e34.5 KB
    Loading...
  2. 2026-08-15Upload folder using huggingface_hub16598094.5 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration