← back to catalog · registered 2026-08-22 13:56

darkmaniac7/Qwen3.6-35B-A3B-abliterated-MNN

darkmaniac7 Qwen 35B MoE multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/darkmaniac7%2FQwen3.6-35B-A3B-abliterated-MNN"
Response includes
  • classification m1
  • files 13
  • hub_downloads_all_time 755
  • author_summary 21 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
755
185 last 30d - stable
Likes
1
Model age
5mo ago
created 2026-04-17
Downloads over time
Now901→from124↑627%
85383681979124 on Apr 15901 on Oct 11AprMayJunJulAugSepOct
Apr 15 → Oct 11 · 65 snapshots · spans 179 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en
Quantizations
BF16
Tags
chat qwen qwen3.6 moe mixture-of-experts multimodal abliterated uncensored heretic mnn tokforge text-generation

Related

Total size
970 MB
Files
13
Quantizations
2
Registered
2026-08-22 13:56
Last updated on HF
2026-10-05 08:42

Files by quantization

BF16 1 file 970 MB
embeddings_bf16.bin 970 MB fa0611f1 download
Auxiliary files 12 files 20.3 GB
llm.mnn.weight 19.9 GB a64a53b0 download
visual.mnn.weight 242 MB 99e86013 download
llm.mnn.json 113 MB b80a8605 download
llm.mnn 51.8 MB 56aea5e0 download
tokenizer.mtok 6.98 MB 9762fc50 download
tokenizer.txt 2.82 MB ee6ce011 download
visual.mnn 535 KB 0a1cf274 download
llm_config.json 8.44 KB 89618817 download
README.md 4.79 KB 00ff2012 download
.gitattributes 1.77 KB 91964e2c download
export_args.json 1.10 KB b37003d0 download
config.json 382 B c45038c8 download

README current version from Hugging Face


license: apache-2.0
language:

  • en
    pipeline_tag: text-generation
    tags:
  • chat
  • qwen
  • qwen3.6
  • moe
  • mixture-of-experts
  • multimodal
  • abliterated
  • uncensored
  • heretic
  • mnn
  • tokforge
    base_model:
  • Youssofal/Qwen3.6-35B-A3B-Abliterated-Heretic-BF16
  • Qwen/Qwen3.6-35B-A3B

TokForge

Runs on-device in the TokForge app.

Qwen3.6-35B-A3B-abliterated-MNN

MNN-format 4-bit quantization of the Heretic-abliterated Qwen3.6-35B-A3B multimodal MoE, packaged for the TokForge Android MNN-fork runtime.

What this is

  • Source (upstream abliteration): Youssofal/Qwen3.6-35B-A3B-Abliterated-Heretic-BF16
  • Upstream of that: Qwen/Qwen3.6-35B-A3B
  • Abliteration methodology: Heretic input-side split-MoE transfer (MPOA/SOMA-style). Upstream reports 1/25 refusal rate on the official 25-prompt harmful-behaviors check (vs. 22/25 for the base model) with KL-divergence 0.0107 vs. base.
  • Architecture: qwen3_5_moe (40 layers, 256 experts, 8 active per token, hybrid linear + full attention, DeepStack vision).

Bundle contents (parity with taobao-mnn/Qwen3.6-35B-A3B-MNN)

  • config.json — MNN LLM geometry (model_type qwen3_5_moe, jinja chat template, MRoPE, vision offsets)
  • llm_config.json — backend defaults (cpu, thread_num=4, precision=low, memory=low)
  • llm.mnn / llm.mnn.weight — quantized MNN graph + external weight blob (Q4, block 64, HQQ)
  • embeddings_bf16.bin — BF16 embedding table (separated from quantized weights via --seperate_embed)
  • tokenizer.mtok — MNN binary tokenizer
  • visual.mnn / visual.mnn.weight — vision transformer (DeepStack VLM)

Quantization scheme

Identical to the base taobao-mnn/Qwen3.6-35B-A3B-MNN bundle:

Flag Value
--quant_bit 4
--quant_block 64
--lm_quant_bit 4
--lm_quant_block 64
--embed_bit 16
--hqq enabled
--seperate_embed enabled

This parity means the abliterated variant should behave identically to the base Qwen3.6 bundle in terms of load time, memory footprint, and decode throughput on TokForge-supported devices.

VLM capability (added post-release, issue #217)

This bundle now ships with visual.mnn + visual.mnn.weight — DeepStack vision transformer for image input. llm_config.json has is_visual: true with the full vision config block (image_mean, image_norm, image_size: 420, vision_start, vision_end, image_pad, num_grid_per_side: 48, has_deepstack: true).

The vision tower in Qwen3.6-35B-A3B is architecturally identical across the base Qwen release and the Heretic-abliterated variant (abliteration targets only the LLM decoder MLP layers, verified by structural comparison of all 333 *.visual.* weight keys in both safetensors). The visual assets here are therefore drop-in compatible with taobao-mnn/Qwen3.6-35B-A3B-MNN and byte-identical to those in the base bundle.

Attention-stack fields (attention_type: mix, sliding_window: 4, layer_nums: 40) match the base taobao-mnn/Qwen3.6-35B-A3B-MNN bundle exactly.

Original conversion note

The first upload of this repo (pre-#217) shipped text-only: the ONNX visual export ran during the initial llmexport.py run, but the final MNNConvert step did not emit visual.mnn / visual.mnn.weight into the output dir. The exporter has since been patched with a --visual_only flag in the TokForge MNN fork to allow re-emitting vision assets without a full re-conversion. See the upstream issue thread for details.

Target runtime

TokForge Android — MNN fork with .mtok tokenizer, DeepStack VLM support, and the libMNN cherry-pick that lets the base Qwen3.6-35B-A3B load on 24 GB devices (RedMagic SM8850 verified at ~6.66 cold / ~8.32 warm tok/s).

Usage with upstream MNN llm_demo

git clone https://github.com/alibaba/MNN.git
cd MNN && mkdir build && cd build
cmake .. -DMNN_LOW_MEMORY=true -DMNN_CPU_WEIGHT_DEQUANT_GEMM=true \
         -DMNN_BUILD_LLM=true -DMNN_SUPPORT_TRANSFORMER_FUSE=true
make -j

./llm_demo /path/to/Qwen3.6-35B-A3B-abliterated-MNN/config.json prompt.txt

License & safety

  • Apache-2.0 (inherited from Qwen and the Youssofal abliteration).
  • This is a safety-reduced / uncensored variant. It refuses far less than the base model on the MPOA/SOMA refusal benchmark. Deploy with appropriate user-facing controls and local policy.
  • Export pipeline: alibaba/MNN llmexport (tq-merged branch, TokForge fork).

README history 4 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-10-05Card metadata: license, base_model, base_model_relationdde15b74.8 KB
    Loading...
  2. 2026-07-04Add TokForge app linksbb83b124.8 KB
    Loading...
  3. 2026-04-18Update README: VLM now supported (#217)00816cd4.5 KB
    Loading...
  4. 2026-04-17Initial upload: Qwen3.6-35B-A3B abliterated-heretic, MNN format (Q4/block64/H...3e672333.9 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration