← back to catalog · registered 2026-08-22 13:56

Youssofal/Qwen3.6-35B-A3B-Abliterated-Heretic-MLX-6bit

Youssofal Qwen 35B MoE multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/Youssofal%2FQwen3.6-35B-A3B-Abliterated-Heretic-MLX-6bit"
Response includes
  • classification m3
  • files 20
  • hub_downloads_all_time 10,029
  • author_summary 17 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · lifetime
10K
336 last 30d - cooling
Likes
4
Model age
5mo ago
created 2026-04-16
Downloads over time
Now10.1K→from1.7K↑508%
1.2K4.5K7.7K11K1.7K on Apr 1510.1K on Oct 11AprMayJunJulAugSepOct
Apr 15 → Oct 11 · 67 snapshots · spans 179 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
mlx safetensors qwen3_5_moe mlx-lm qwen qwen3.6 moe mixture-of-experts multimodal vlm vision video

Related

Total size
28.0 GB
Files
20
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-04-26 01:04

Files by quantization

Auxiliary files 20 files 28.1 GB
model-00003-of-00006.safetensors 4.96 GB 820beca7 download
model-00001-of-00006.safetensors 4.88 GB 8bc0d169 download
model-00002-of-00006.safetensors 4.83 GB 06aecb2e download
model-00004-of-00006.safetensors 4.83 GB b2fdad6c download
model-00005-of-00006.safetensors 4.82 GB 591049df download
model-00006-of-00006.safetensors 2.91 GB aeb7fe4a download
model-vision-00001-of-00001.safetensors 852 MB d9aba9d9 download
tokenizer.json 19.1 MB 87a7830d download
model.safetensors.index.json 205 KB e8ce326c download
config.json 136 KB 7bccefe3 download
mlx_variant_metadata.json 16.7 KB 5c1f48ce download
chat_template.jinja 7.58 KB a8755d82 download
README.md 3.59 KB bdde8cb5 download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.28 KB a3be3e79 download
tokenizer_config.json 1.15 KB 040f6e4b download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
generation_config.json 242 B 89a0a0db download
configuration.json 58.0 B d24dba94 download

README current version from Hugging Face


base_model: Youssofal/Qwen3.6-35B-A3B-Abliterated-Heretic-BF16
library_name: mlx
pipeline_tag: text-generation
license: apache-2.0
tags:

  • mlx
  • mlx-lm
  • qwen
  • qwen3.6
  • qwen3_5_moe
  • moe
  • mixture-of-experts
  • multimodal
  • vlm
  • vision
  • video
  • image-text-to-text
  • abliterated
  • uncensored
  • heretic
  • mpoa
  • soma
  • apple-silicon
  • 6-bit
    quantized_by: Youssofal

Qwen3.6-35B-A3B-Abliterated-Heretic-MLX-6bit

This is an MLX release of an abliterated version of Qwen's Qwen3.6-35B-A3B.

By applying Heretic's ablation pipeline to the text-side MoE stack, the base refusal behavior was removed at the weight level. This release keeps the Qwen3.6-35B-A3B reasoning and instruction-following profile in Apple MLX format for local deployment on Apple Silicon hardware.

This MLX repo includes the retained Qwen3.6 image/video processor files and vision tower tensors for runtimes with Qwen3.6 multimodal MLX support.

Quick Benchmarks

Check Original Qwen3.6-35B-A3B Abliterated Heretic MLX
Official 25-prompt refusal check 22/25 refusals 1/25 refusals
Archived Heretic KL divergence - 0.010655362159013748

Methodology & Model Notes

Qwen3.6-35B-A3B is a sparse MoE model in the qwen3_5_moe family. The accepted abliterated BF16 source checkpoint was produced with a Heretic MPOA/SOMA-style sibling-transfer workflow and finalized with the input-side split-MoE intervention that cleared the official 25-prompt refusal marker suite down to 1/25.

This MLX release was built directly from the published BF16 Heretic checkpoint using a high-quality layer-aware quantization policy instead of a flat per-weight pass.

  • quant target: 6-bit
  • quant build: 6-bit tuned layer-aware quantization
  • source checkpoint: Youssofal/Qwen3.6-35B-A3B-Abliterated-Heretic-BF16
  • published variant: Qwen3.6-35B-A3B-Abliterated-Heretic-MLX-6bit

The layer-aware policy keeps more precision on sensitive projections in the early, late, and selected middle layers so the quant stays cleaner than a naive flat conversion.

Validation

This published MLX variant passed:

  • the official 25-prompt refusal marker check in standard thinking-enabled chat format: 1/25 refusals
  • the local smoke suite for general chat, short reasoning, and short code output: all_looks_ok=true

Running

from mlx_lm import load, generate

model, tokenizer = load("Youssofal/Qwen3.6-35B-A3B-Abliterated-Heretic-MLX-6bit")

messages = [{"role": "user", "content": "Write a short Python function that reverses a string."}]
prompt = tokenizer.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True,
)

response = generate(model, tokenizer, prompt=prompt, max_tokens=256)
print(response)

Files

The repo root contains the complete 6-bit MLX export for this variant:

  • config.json
  • model.safetensors.index.json
  • split quantized text model-*.safetensors shards
  • model-vision-00001-of-00001.safetensors
  • tokenizer, generation, and processor files
  • README.md

Credits

Disclaimer

This model has had refusal behavior removed at the weight level. It will answer prompts that the base model would normally refuse. You are responsible for how you use it.

README history 6 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-04-26Fix MLX vision tower packagingc4772773.6 KB
    Loading...
  2. 2026-04-17Replace MLX 6-bit export with corrected rebuild8fcc11f3.3 KB
    Loading...
  3. 2026-04-16Update original refusal score to 22/251e7e5a13.4 KB
    Loading...
  4. 2026-04-16Fix original refusal score tablec19bc1c3.4 KB
    Loading...
  5. 2026-04-16Update uploads-in-progress bannerf6fc7c93.5 KB
    Loading...
  6. 2026-04-16Add uploads-in-progress notice089e1483.6 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration