← back to catalog · registered 2026-08-22 13:56

marafx2025/Qwen3.6-35B-A3B-uncensored-heretic-APEX-GGUF

marafx2025 Qwen 35B GGUF MoE multimodal second-order 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/marafx2025%2FQwen3.6-35B-A3B-uncensored-heretic-APEX-GGUF"
Response includes
  • classification m3
  • files 10
  • benchmarks 11 entries
  • hub_downloads_all_time 3,144
  • author_summary 3 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · lifetime
3K
334 last 30d - stable
Likes
2
Model age
5mo ago
created 2026-04-27
Downloads over time
Now3.3K→from461↑611%
3201.4K2.5K3.6K461 on Apr 293.3K on Oct 11AprMayJunJulAugSepOct
Apr 29 → Oct 11 · 63 snapshots · spans 165 days

Benchmarks

Benchmark Score Source
Entertainment 1.9 UGI
Hazardous 2.4 UGI
Natural Intelligence 28.42 UGI
Political lean -12.5% UGI
Sensitive-Info 18.88 UGI
SocPol 1.5 UGI
UGI 44.25 UGI
Willingness (10) 9.5 UGI
W10-Adherence 9 UGI
W10-Direct 10 UGI
Writing 37.6 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
gguf quantized apex moe mixture-of-experts qwen3 vlm vision uncensored heretic base_model:llmfan46/Qwen3.6-35B-A3B-uncensored-heretic base_model:quantized:llmfan46/Qwen3.6-35B-A3B-uncensored-heretic

Related

Total size
136 GB
Files
10
Quantizations
2
Registered
2026-08-22 13:56
Last updated on HF
2026-04-27 20:29

Files by quantization

mmproj 1 file 861 MB
mmproj.gguf 861 MB 1c625f05 download
Auxiliary files 9 files 136 GB
Qwen3.6-35B-A3B-uncensored-heretic-APEX-I-Balanced.gguf 23.9 GB 7e4d28f5 download
Qwen3.6-35B-A3B-uncensored-heretic-APEX-Balanced.gguf 23.9 GB b1d1fd7e download
Qwen3.6-35B-A3B-uncensored-heretic-APEX-I-Quality.gguf 21.3 GB 667bd985 download
Qwen3.6-35B-A3B-uncensored-heretic-APEX-Quality.gguf 21.3 GB 07323bea download
Qwen3.6-35B-A3B-uncensored-heretic-APEX-I-Compact.gguf 16.1 GB aa80e8a5 download
Qwen3.6-35B-A3B-uncensored-heretic-APEX-Compact.gguf 16.1 GB 9453e789 download
Qwen3.6-35B-A3B-uncensored-heretic-APEX-I-Mini.gguf 13.3 GB 956dd8cd download
README.md 4.52 KB cad2c424 download
.gitattributes 2.15 KB 085de57d download

README current version from Hugging Face


license: apache-2.0
base_model: llmfan46/Qwen3.6-35B-A3B-uncensored-heretic
tags:

  • gguf
  • quantized
  • apex
  • moe
  • mixture-of-experts
  • qwen3
  • vlm
  • vision
  • uncensored
  • heretic

⚡ Each donation = another big MoE quantized

I host 25+ free APEX MoE quantizations as independent research. My only local hardware is an NVIDIA DGX Spark (122 GB unified memory) — enough for ~30-50B-class MoEs, but bigger ones (200B+) require rented compute on H100/H200/Blackwell, typically $20-100 per quant.
If APEX quants are useful to you, your support directly funds those bigger runs.

🎉 Patreon (Monthly)  |  ☕ Buy Me a Coffee  |  ⭐ GitHub Sponsors

💚 Big thanks to Hugging Face for generously donating additional storage — much appreciated.

Qwen3.6 35B-A3B Uncensored Heretic APEX GGUF

APEX (Adaptive Precision for EXpert Models) quantizations of llmfan46/Qwen3.6-35B-A3B-uncensored-heretic.

Brought to you by the LocalAI team | APEX Project | Technical Report

Available Files

File Profile Size Best For
Qwen3.6-35B-A3B-uncensored-heretic-APEX-I-Balanced.gguf I-Balanced 24 GB Best overall quality/size ratio
Qwen3.6-35B-A3B-uncensored-heretic-APEX-Balanced.gguf Balanced 24 GB General purpose
Qwen3.6-35B-A3B-uncensored-heretic-APEX-I-Quality.gguf I-Quality 22 GB Highest quality with imatrix
Qwen3.6-35B-A3B-uncensored-heretic-APEX-Quality.gguf Quality 22 GB Highest quality standard
Qwen3.6-35B-A3B-uncensored-heretic-APEX-I-Compact.gguf I-Compact 17 GB Consumer GPUs, best quality/size
Qwen3.6-35B-A3B-uncensored-heretic-APEX-Compact.gguf Compact 17 GB Consumer GPUs
Qwen3.6-35B-A3B-uncensored-heretic-APEX-I-Mini.gguf I-Mini 14 GB Smallest viable, fastest inference
mmproj.gguf Vision projector ~1 GB Required for image understanding

What is APEX?

APEX is a quantization strategy for Mixture-of-Experts (MoE) models. It classifies tensors by role (routed expert, shared expert, attention) and applies a layer-wise precision gradient — edge layers get higher precision, middle layers get more aggressive compression. I-variants use diverse imatrix calibration (chat, code, reasoning, tool-calling, agentic traces, Wikipedia).

The key insight: in MoE models, expert FFN tensors make up the bulk of model weight but only ~8/256 experts activate per token. APEX compresses middle-layer experts more aggressively while preserving edge layers (first/last 5) and keeping attention, SSM/Mamba, and shared expert tensors at higher precision.

See the APEX project for full details, technical report, and scripts.

Architecture

  • Model: Qwen3.6 35B-A3B Uncensored Heretic (uncensored fine-tune)
  • Base: Qwen 3.6 35B-A3B
  • Layers: 40
  • Experts: 256 routed + shared (8 active per token)
  • Total Parameters: ~35B
  • Active Parameters: ~3B per token
  • Attention: Hybrid (full attention every 4th layer, linear/Mamba otherwise)
  • Vision: Built-in vision encoder (mmproj included)
  • APEX Config: 5+5 symmetric edge gradient across 40 layers
  • Calibration: v1.3 diverse dataset (chat, code, reasoning, multilingual, tool-calling, Wikipedia)

Run with LocalAI

local-ai run mudler/Qwen3.6-35B-A3B-uncensored-heretic-APEX-GGUF@Qwen3.6-35B-A3B-uncensored-heretic-APEX-I-Balanced.gguf

Credits

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-04-27Duplicate from mudler/Qwen3.6-35B-A3B-uncensored-heretic-APEX-GGUFbc74baf4.5 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration