← back to catalog · registered 2026-08-22 13:56

hares-0o0/Qwen3.6-35B-A3B-Abliterated-Heretic-MLX-4bit

hares-0o0 Qwen 35B MoE multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/hares-0o0%2FQwen3.6-35B-A3B-Abliterated-Heretic-MLX-4bit"
Response includes
  • classification m3
  • files 19
  • hub_downloads_all_time 1,023
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · lifetime
1K
62 last 30d - cooling
Likes
0
Model age
5mo ago
created 2026-04-29
Downloads over time
Now1K→from362↑187%
3285888481.1K362 on Apr 291K on Oct 11AprMayJunJulAugSepOct
Apr 29 → Oct 11 · 63 snapshots · spans 165 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
mlx safetensors qwen3_5_moe mlx-lm qwen qwen3.6 moe mixture-of-experts multimodal vlm vision video

Related

Total size
22.9 GB
Files
19
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-04-29 11:56

Files by quantization

Auxiliary files 19 files 22.9 GB
model-00002-of-00005.safetensors 4.98 GB b9b3ac39 download
model-00003-of-00005.safetensors 4.98 GB ed1a0397 download
model-00001-of-00005.safetensors 4.95 GB d8ac2a5d download
model-00004-of-00005.safetensors 4.92 GB 60c777cc download
model-00005-of-00005.safetensors 2.26 GB f57dca6f download
model-vision-00001-of-00001.safetensors 852 MB d9aba9d9 download
tokenizer.json 19.1 MB 87a7830d download
model.safetensors.index.json 205 KB 812f1824 download
config.json 136 KB 261d537a download
mlx_variant_metadata.json 16.8 KB d84901d2 download
chat_template.jinja 7.58 KB a8755d82 download
README.md 3.61 KB 47f9e937 download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.28 KB a3be3e79 download
tokenizer_config.json 1.15 KB 040f6e4b download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
generation_config.json 242 B 89a0a0db download
configuration.json 58.0 B d24dba94 download

README current version from Hugging Face


base_model: Youssofal/Qwen3.6-35B-A3B-Abliterated-Heretic-BF16
library_name: mlx
pipeline_tag: text-generation
license: apache-2.0
tags:

  • mlx
  • mlx-lm
  • qwen
  • qwen3.6
  • qwen3_5_moe
  • moe
  • mixture-of-experts
  • multimodal
  • vlm
  • vision
  • video
  • image-text-to-text
  • abliterated
  • uncensored
  • heretic
  • mpoa
  • soma
  • apple-silicon
  • 4-bit
    quantized_by: Youssofal

Qwen3.6-35B-A3B-Abliterated-Heretic-MLX-4bit

This is an MLX release of an abliterated version of Qwen's Qwen3.6-35B-A3B.

By applying Heretic's ablation pipeline to the text-side MoE stack, the base refusal behavior was removed at the weight level. This release keeps the Qwen3.6-35B-A3B reasoning and instruction-following profile in Apple MLX format for local deployment on Apple Silicon hardware.

This MLX repo includes the retained Qwen3.6 image/video processor files and vision tower tensors for runtimes with Qwen3.6 multimodal MLX support.

Quick Benchmarks

Check Original Qwen3.6-35B-A3B Abliterated Heretic MLX
Official 25-prompt refusal check 22/25 refusals 3/25 refusals
Archived Heretic KL divergence - 0.010655362159013748

Methodology & Model Notes

Qwen3.6-35B-A3B is a sparse MoE model in the qwen3_5_moe family. The accepted abliterated BF16 source checkpoint was produced with a Heretic MPOA/SOMA-style sibling-transfer workflow and finalized with the input-side split-MoE intervention that cleared the official 25-prompt refusal marker suite down to 1/25.

This MLX release was built directly from the published BF16 Heretic checkpoint using a high-quality layer-aware quantization policy instead of a flat per-weight pass.

  • quant target: 4-bit
  • quant build: 4-bit high-quality mixed layer-aware quantization (4/6-bit)
  • source checkpoint: Youssofal/Qwen3.6-35B-A3B-Abliterated-Heretic-BF16
  • published variant: Qwen3.6-35B-A3B-Abliterated-Heretic-MLX-4bit

The layer-aware policy keeps more precision on sensitive projections in the early, late, and selected middle layers so the quant stays cleaner than a naive flat conversion.

Validation

This published MLX variant passed:

  • the official 25-prompt refusal marker check in standard thinking-enabled chat format: 3/25 refusals
  • the local smoke suite for general chat, short reasoning, and short code output: all_looks_ok=true

Running

from mlx_lm import load, generate

model, tokenizer = load("Youssofal/Qwen3.6-35B-A3B-Abliterated-Heretic-MLX-4bit")

messages = [{"role": "user", "content": "Write a short Python function that reverses a string."}]
prompt = tokenizer.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True,
)

response = generate(model, tokenizer, prompt=prompt, max_tokens=256)
print(response)

Files

The repo root contains the complete 4-bit MLX export for this variant:

  • config.json
  • model.safetensors.index.json
  • split quantized text model-*.safetensors shards
  • model-vision-00001-of-00001.safetensors
  • tokenizer, generation, and processor files
  • README.md

Credits

Disclaimer

This model has had refusal behavior removed at the weight level. It will answer prompts that the base model would normally refuse. You are responsible for how you use it.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-04-29Duplicate from Youssofal/Qwen3.6-35B-A3B-Abliterated-Heretic-MLX-4bit2117d103.6 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration