← back to catalog · registered 2026-08-22 13:56

froggeric/Qwen3.6-27B-Uncensored-Heretic-v2-MLX-8bit

froggeric Qwen 27B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/froggeric%2FQwen3.6-27B-Uncensored-Heretic-v2-MLX-8bit"
Response includes
  • classification m3
  • files 18
  • benchmarks 11 entries
  • hub_downloads_all_time 3,696
  • author_summary 8 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · lifetime
4K
171 last 30d - cooling
Likes
2
Model age
5mo ago
created 2026-04-27
Downloads over time
Now3.8K→from772↑386%
6231.8K2.9K4.1K772 on Apr 293.8K on Oct 11AprMayJunJulAugSepOct
Apr 29 → Oct 11 · 63 snapshots · spans 165 days

Benchmarks

Benchmark Score Source
Entertainment 1.5 UGI
Hazardous 2.9 UGI
Natural Intelligence 29.17 UGI
Political lean -24.5% UGI
Sensitive-Info 17.24 UGI
SocPol 1 UGI
UGI 43.16 UGI
Willingness (10) 9.5 UGI
W10-Adherence 10 UGI
W10-Direct 9 UGI
Writing 39.42 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh multilingual
Tags
mlx safetensors qwen3_5 mlx-lm mlx-vlm qwen3.6 conversational vision multimodal uncensored abliterated heretic

Related

Total size
27.5 GB
Files
18
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-05-06 08:30

Files by quantization

Auxiliary files 18 files 27.5 GB
model-00003-of-00006.safetensors 4.99 GB 0dfccc36 download
model-00002-of-00006.safetensors 4.99 GB 83675696 download
model-00004-of-00006.safetensors 4.97 GB 09d9313d download
model-00001-of-00006.safetensors 4.95 GB 70b3d4ac download
model-00005-of-00006.safetensors 4.93 GB 8f3c48e3 download
model-00006-of-00006.safetensors 2.65 GB 265391ad download
tokenizer.json 19.1 MB 6f32ce20 download
model.safetensors.index.json 213 KB 2201d720 download
tokenizer_config.json 12.6 KB 5af7d7e4 download
chat_template.jinja 11.0 KB 7e823502 download
README.md 10.8 KB d2082f76 download
chat_template.README.md 6.03 KB f2d4808a download
config.json 4.52 KB b723c81f download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.16 KB 33818c7f download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
generation_config.json 226 B 16d319af download

README current version from Hugging Face


language:

  • en
  • zh
  • multilingual
    license: apache-2.0
    library_name: mlx
    tags:
  • mlx
  • mlx-lm
  • mlx-vlm
  • qwen3.6
  • qwen3_5
  • conversational
  • vision
  • multimodal
  • uncensored
  • abliterated
  • heretic
    base_model:
  • llmfan46/Qwen3.6-27B-uncensored-heretic-v2
    pipeline_tag: image-text-to-text
    quantization: 8-bit

Qwen3.6-27B Uncensored Heretic v2

MLX 8-bit · Apple Silicon native

Text · Vision · Video · Thinking · Tool Calling

MLX 6-bit
MLX 4-bit
LM Studio
License


2026-04-30 — Re-converted from updated source. The upstream model (llmfan46) was re-done using MPOA (Magnitude-Preserving Orthogonal Ablation) replacing the earlier ARA method. Key improvements: refusals dropped from 13/100 to 6/100, KL divergence from 0.0035 to 0.0021, and reported issues with EOS spam and generation interruptions are fixed. If you downloaded before April 30, re-download for the better version.


Why this model?

Two things set this apart from other Qwen 3.6 conversions:

1. Architecture-aware uncensoring. Qwen 3.6 uses a hybrid attention design — linear (DeltaNet-style) and traditional softmax blocks, mixed 3:1. Most abliteration tools treat them the same. llmfan46 applied separate parameters for each attention type using the Heretic tool with the MPOA (Magnitude-Preserving Orthogonal Ablation) method, yielding one of the lowest KL divergences of any uncensored Qwen variant — dramatically fewer refusals with negligible capability loss.

2. A fixed chat template. The official Qwen 3.6 template is broken on every C++ runtime (LM Studio, llama.cpp, MLX). Tool calls crash, the developer role throws errors, and empty thinking blocks waste your context window. This model ships with a rewritten template that fixes all five issues and adds a thinking toggle (<|think_on|> / <|think_off|>) you can drop into any message.


Quick start

Text

from mlx_lm import load, generate

model, tokenizer = load("froggeric/Qwen3.6-27B-Uncensored-Heretic-v2-MLX-8bit")
response = generate(model, tokenizer, prompt="Hello", max_tokens=256, temp=0.7)
print(response)

Vision

from mlx_vlm import load, generate
from mlx_vlm.prompt_utils import apply_chat_template

model, processor = load("froggeric/Qwen3.6-27B-Uncensored-Heretic-v2-MLX-8bit")
image = ["path/to/image.jpg"]
prompt = "Describe this image."
formatted = apply_chat_template(processor, model.config, prompt, num_images=len(image))
result = generate(model, processor, formatted, image, max_tokens=256, temp=0.7)
print(result.text)

CLI

# Text
mlx_lm.generate \
  --model froggeric/Qwen3.6-27B-Uncensored-Heretic-v2-MLX-8bit \
  --prompt "Hello"

# Vision
mlx_vlm.generate \
  --model froggeric/Qwen3.6-27B-Uncensored-Heretic-v2-MLX-8bit \
  --image image.jpg --prompt "Describe this image"

Requirements: mlx-lm >= 0.31.2, mlx-vlm >= 0.4.4


System prompt

The first line of your system prompt must be:

You are Qwen, created by Alibaba Cloud. You are a helpful assistant.

The model underperforms without it. You can append anything after that line.


Thinking toggle

Drop <|think_on|> or <|think_off|> anywhere in your system or user prompt. The template intercepts the tag, strips it from context so the model never sees it, and flips the mode.

Fast answer, no reasoning:

System: You are a coding assistant. <|think_off|>
User: What's 2+2?

Deep reasoning:

System: You are a coding assistant. <|think_on|>
User: Implement a red-black tree in Rust.

Chat template fixes

The official Qwen 3.6 Jinja template has five bugs that break real usage. This model ships with a rewritten template that fixes all of them:

Bug Impact Fix
` items` filter in tool calls Crashes on every C++ runtime (LM Studio, llama.cpp, MLX)
` safe` filter Python-only, does not exist in C++ Jinja
developer role Modern APIs send it; official template throws an error Maps to system
Empty thinking blocks Wraps every past turn in tags, even with nothing inside — wastes context tokens Only emitted when reasoning_content is non-empty
</thinking> hallucination Model sometimes generates the wrong closing tag; parser fails Detects which tag was used and splits on that

Works in LM Studio, llama.cpp (--jinja), vLLM, MLX, oMLX, and any engine that supports HuggingFace Jinja templates.


The uncensoring

This model uses Heretic v1.2.0 with the MPOA (Magnitude-Preserving Orthogonal Ablation) method.

How it works

Heretic identifies the "refusal direction" in the model's residual stream by comparing activations on harmless vs. harmful prompts, then orthogonalizes specific weight matrices against that direction so the model can no longer express refusal behavior.

MPOA preserves the norm of the original weight matrices during abliteration, maintaining the model's activation distributions and thus its capabilities — unlike simple orthogonal projection which can distort the activation landscape.

What llmfan46 did differently

Standard Heretic treats all attention blocks identically. Qwen 3.6's hybrid architecture mixes linear attention (DeltaNet-style) and traditional softmax attention in a 3:1 ratio. llmfan46 applied separate abliteration parameters for each attention type, allowing more precise removal of refusal behavior with less collateral damage to model capabilities.

This approach was submitted as a pull request to Heretic but was not merged — not because it doesn't work, but because the extra parameters increase optimization time. For this specific architecture, it produces superior results.


How it compares

Community results

r/LocalLLaMA users have been A/B-testing various uncensored Qwen 3.6 variants — Heretic, HauhauCS Aggressive, abliterix, and simple orthogonal projection. The pattern is consistent: Heretic produces the best balance of refusal removal and output quality.

Community discussion →

Why

Most abliteration methods treat all layers identically. Qwen 3.6's hybrid attention (3:1 linear-to-softmax ratio) means a single parameter set either under-abliterate the DeltaNet blocks or over-abliterate the softmax blocks. Architecture-aware abliteration — separate parameters per attention type — is the key differentiator.

A note on SSM conv1d "repair"

Some uncensored variants apply a pre-processing step that rescales SSM conv1d weights before abliteration, claiming to fix "outlier" tensors in the DeltaNet linear attention layers. This technique (originating as "Sig-ScaleSync") was benchmarked with 284 data points across perplexity, needle-in-a-haystack, and repetition tests at multiple context lengths (4K–128K). Result: perplexity degraded at every length with no improvement in NIAH or repetition. The unrepaired original weights perform best.

Abliterating a degraded baseline can yield a lower measured KL divergence — but that measures distance from a worse starting point, not better preservation of the original model's capabilities.


Sampling

From the official Qwen authors. Reserve 128K+ context for thinking mode.

Mode temp top_p top_k min_p repeat_penalty presence_penalty
Thinking (coding) 0.6 0.95 20 0 1.0 off
Thinking (general) 1.0 0.95 20 0 1.0 1.5
Non-thinking 0.7 0.8 20 0 1.0 1.5

GGUF runtimes use presence_penalty (0 = off). MLX / LM Studio use repeat_penalty (1.0 = off).


This conversion

Source llmfan46/Qwen3.6-27B-uncensored-heretic-v2 (BF16 safetensors, MPOA abliteration)
Quantization 8-bit (8.6 bits/weight, ~28 GB across 6 shards)
Chat template Fixed Jinja template with tool calling, developer role, thinking toggle, and hallucination handling
Minimum RAM ~32 GB (28 GB weights + overhead)
Architecture details
Spec Value
Architecture Dense — 27.8B params, all active per token
Layers 64 (3x linear attention + 1x full attention, 16 repetitions)
Attention 24 Q heads, 4 KV heads (GQA), head_dim 256
Linear attention 16 QK heads, 48 V heads, head_dim 128
FFN intermediate_size 17408
Context 262K native, 1M+ with YaRN
RoPE theta 10M, partial_rotary_factor 0.25, mrope_interleaved
Vocab 248K tokens
Multimodal Text, image, video
Multi-token prediction Supported (1 draft layer)
model_type qwen3_5

Credits

Role Author
Original model Alibaba Cloud (Qwen team)
Refusal direction research Arditi et al.
MPOA method Jim Lai
Heretic tool Philipp Weidmann
Architecture-aware abliteration + uncensored variant llmfan46
Fixed chat template + MLX conversion froggeric

Links

License

Apache-2.0, inherited from Qwen3.6.

README history 5 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-05-01Upload README.md with huggingface_hubd48630510.8 KB
    Loading...
  2. 2026-05-01Add files using upload-large-folder toold7dd0589.2 KB
    Loading...
  3. 2026-04-30Upload README.md with huggingface_hub9584d438.7 KB
    Loading...
  4. 2026-04-27Upload README.md with huggingface_hubdebe5e18.4 KB
    Loading...
  5. 2026-04-27Add files using upload-large-folder toolb4451ed5.9 KB
    Loading...

Discussions 1 thread

  1. 2026-04-306bit Variants?open2 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration