← back to catalog · registered 2026-08-22 13:56

underlotus/Qwen3.6-27B-uncensored-heretic-v2-Native-MTP-Preserved-oQ4-mtp

underlotus Qwen 27B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/underlotus%2FQwen3.6-27B-uncensored-heretic-v2-Native-MTP-Preserved-oQ4-mtp"
Response includes
  • classification m3
  • files 13
  • benchmarks 11 entries
  • hub_downloads_all_time 583
  • author_summary 3 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · lifetime
583
130 last 30d - stable
Likes
0
Model age
2mo ago
created 2026-08-04
Downloads over time
Now626→from161↑289%
138316494673161 on Aug 5626 on Oct 11626 on Oct 9AugSepOct
Aug 5 → Oct 11 · 50 snapshots · spans 67 days

Benchmarks

Benchmark Score Source
Entertainment 1.7 UGI
Hazardous 2.4 UGI
Natural Intelligence 26.28 UGI
Political lean -26.8% UGI
Sensitive-Info 18.31 UGI
SocPol 1.6 UGI
UGI 43.88 UGI
Willingness (10) 9.5 UGI
W10-Adherence 10 UGI
W10-Direct 9 UGI
Writing 38.41 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh
Tags
mlx safetensors qwen3_5 oq 4bit mixed-precision uncensored mtp qwen apple-silicon image-text-to-text conversational

Related

Total size
15.9 GB
Files
13
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-15 04:19

Files by quantization

Auxiliary files 13 files 15.9 GB
model-00002-of-00004.safetensors 4.70 GB 95c6773e download
model-00003-of-00004.safetensors 4.69 GB da13b81d download
model-00001-of-00004.safetensors 4.66 GB 1216b864 download
model-00004-of-00004.safetensors 1.80 GB 9f7d27ad download
tokenizer.json 19.1 MB 6f32ce20 download
model.safetensors.index.json 207 KB 6ba6ae46 download
config.json 37.7 KB fc23a3aa download
chat_template.jinja 11.7 KB 177eacb1 download
README.md 2.62 KB a9ae54e8 download
.gitattributes 1.53 KB 52373fe2 download
tokenizer_config.json 1.13 KB c487bad4 download
preprocessor_config.json 390 B 2ea84a43 download
generation_config.json 226 B 2dd033e0 download

README current version from Hugging Face


license: apache-2.0
language:

  • en
  • zh
    base_model: llmfan46/Qwen3.6-27B-uncensored-heretic-v2-Native-MTP-Preserved
    pipeline_tag: image-text-to-text
    tags:
  • mlx
  • oq
  • 4bit
  • mixed-precision
  • uncensored
  • mtp
  • qwen
  • apple-silicon

Qwen3.6-27B uncensored heretic v2 — oQ4 (MTP preserved)

Mixed-precision quant of llmfan46/Qwen3.6-27B-uncensored-heretic-v2-Native-MTP-Preserved, produced with oQ (oMLX v0.5.4). MTP head preserved. Standard MLX safetensors — compatible with oMLX, mlx-lm, LM Studio, and any MLX-capable app.

What is oQ?

Unlike uniform 4-bit quantization, oQ is a data-driven mixed-precision quantizer that calibrates per-layer sensitivity and allocates bits where they matter most. Critical layers (embeddings, LM head, the most sensitive transformer layers) are automatically promoted to 8-bit, while less sensitive layers stay at 4-bit. Typical result: ~4.6 bits-per-weight.

Benchmarked on Qwen3.5-35B-A3B (oMLX project):

Benchmark mlx-lm 4-bit oQ4
MMLU (300) 79.7% 83.3%
TruthfulQA (300) 87.7% 88.0%
HumanEval (full) 87.2% 85.4%
MBPP (300) 71.7% 74.3%

Performance (oMLX on M4 10-core)

Context PP tok/s TG tok/s Peak Mem
1k 62.1 12.4 16.7 GB
4k 60.8 11.5 18.2 GB
Batch TG tok/s Speedup
1× 12.4 1.00×
2× 11.8 0.95×
4× 42.2 3.40×

Full benchmark →

💡 MTP preserved — compared to the non-MTP quant of the abliterated variant (6.5 tok/s), this model is nearly 2× faster on token generation.

Why this quant

The original BF16 weights require ~55 GB. This oQ4 quant runs in ~16–18 GB on Apple Silicon while keeping the full MTP stack intact.

Quick start

# oMLX
omlx serve --model underlotus/Qwen3.6-27B-uncensored-heretic-v2-Native-MTP-Preserved-oQ4-mtp
# mlx-lm
from mlx_lm import load, generate

model, tokenizer = load("underlotus/Qwen3.6-27B-uncensored-heretic-v2-Native-MTP-Preserved-oQ4-mtp")
response = generate(model, tokenizer, prompt="Hello!", max_tokens=256)
print(response)

Original model

  • Base: Qwen/Qwen3.6-27B
  • Uncensored: Heretic v2 MPOA pipeline, 94% fewer refusals, KL divergence 0.0021
  • MTP: Multi-token prediction head preserved

License

Apache 2.0, inherited from base model.

README history 5 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-15docs: fix MTP layer count (1 head, not 15 layers)54965d42.6 KB
    Loading...
  2. 2026-08-04Upload README.md with huggingface_hubfbd3c192.6 KB
    Loading...
  3. 2026-08-04Upload README.md with huggingface_hub657923e2.1 KB
    Loading...
  4. 2026-08-04Upload README.md with huggingface_hub5d1cf9a1.2 KB
    Loading...
  5. 2026-08-04Upload Qwen3.6-27B-uncensored-heretic-v2-Native-MTP-Preserved-oQ4-mtp via oMLX52801d9351 B
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration