← back to catalog · registered 2026-08-22 13:56

groxaxo/Qwen3.6-35B-A3B-Abliterated-Heretic-GGUF

groxaxo Qwen 35B GGUF MoE 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/groxaxo%2FQwen3.6-35B-A3B-Abliterated-Heretic-GGUF"
Response includes
  • classification m3
  • files 12
  • benchmarks 16 entries
  • hub_downloads_all_time 4,209
  • author_summary 27 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · lifetime
4K
2K last 30d - stable
Likes
1
Model age
5mo ago
created 2026-04-21
Downloads over time
Now5K→from936↑437%
7312.3K3.9K5.4K936 on Apr 225K on Oct 11AprMayJunJulAugSepOct
Apr 22 → Oct 11 · 64 snapshots · spans 172 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Arena-Battles 25948 LM-Arena
LM Arena Elo 1321.1938733711484 LM-Arena
Arena-Elo-Lower 1316.3871802173958 LM-Arena
Arena-Elo-Upper 1326.0005665249007 LM-Arena
Arena-Rank 61 LM-Arena
Entertainment 1.2 UGI
Hazardous 2.9 UGI
Natural Intelligence 15.9 UGI
Political lean -26.8% UGI
Sensitive-Info 17.11 UGI
SocPol 1.4 UGI
UGI 33.07 UGI
Willingness (10) 6.5 UGI
W10-Adherence 10 UGI
W10-Direct 3 UGI
Writing 30.24 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Quantizations
BF16 IQ4 Q2_K Q3_K Q4_K Q5_K Q6_K Q8_0
Tags
gguf quantized abliterated qwen3 moe base_model:Qwen/Qwen3-30B-A3B base_model:quantized:Qwen/Qwen3-30B-A3B license:apache-2.0 endpoints_compatible region:us conversational

Related

Total size
232 GB
Files
12
Quantizations
9
Registered
2026-08-22 13:56
Last updated on HF
2026-08-22 07:53

Files by quantization

BF16 2 files 65.5 GB
Qwen3.6-35B-A3B-Abliterated-Heretic-BF16.gguf 64.6 GB 22196784 download
mmproj-BF16.gguf 861 MB 356dfaa3 download
Q8_0 1 file 34.4 GB
Qwen3.6-35B-A3B-Abliterated-Heretic-Q8_0.gguf 34.4 GB 546dba71 download
Q6_K 1 file 26.6 GB
Qwen3.6-35B-A3B-Abliterated-Heretic-Q6_K.gguf 26.6 GB 154dc3b2 download
Q5_K 1 file 23.0 GB
Qwen3.6-35B-A3B-Abliterated-Heretic-Q5_K_M.gguf 23.0 GB f4d2d1de download
Q4_K 2 files 38.2 GB
Qwen3.6-35B-A3B-Abliterated-Heretic-Q4_K_M.gguf 19.7 GB ae2fb73a download
Qwen3.6-35B-A3B-Abliterated-Heretic-Q4_K_S.gguf 18.5 GB e07e729f download
IQ4 1 file 17.6 GB
Qwen3.6-35B-A3B-Abliterated-Heretic-IQ4_XS.gguf 17.6 GB ba313487 download
Q3_K 1 file 15.6 GB
Qwen3.6-35B-A3B-Abliterated-Heretic-Q3_K_M.gguf 15.6 GB 2e689a69 download
Q2_K 1 file 12.1 GB
Qwen3.6-35B-A3B-Abliterated-Heretic-Q2_K.gguf 12.1 GB f6a81bc9 download
Auxiliary files 2 files 6.63 KB
README.md 4.37 KB ded113e5 download
.gitattributes 2.27 KB e2269dff download

README current version from Hugging Face


license: apache-2.0
tags:

  • quantized
  • gguf
  • abliterated
  • qwen3
  • moe
    base_model:
  • Qwen/Qwen3-30B-A3B
  • Qwen/Qwen3.6-35B-A3B

Qwen3.6-35B-A3B-Abliterated-Heretic GGUF

Overview

Qwen3.6-35B-A3B-Abliterated-Heretic-GGUF is a GGUF release for llama.cpp-compatible runtimes and local inference, published by groxaxo.
It is intended for open-source evaluation, reproducible experimentation, and compatible local or
hosted inference workflows. The wording below is deliberately limited to what can be verified
from this repository's metadata and artifacts.

The repository name identifies a behavior-modified or reduced-filtering lineage. That label describes the source or conversion history; it is not a guarantee of unrestricted behavior in every prompt or runtime. Test outputs carefully before sharing or deploying them.

At a glance

Field Details
Format GGUF
Source / base Qwen/Qwen3-30B-A3B
Intended task image-text-to-text
License apache-2.0

What is included

  • *.gguf (10 files)
  • Additional configuration, tokenizer, processor, or shard files (10 visible artifacts total)

Quick start

llama.cpp

Download a .gguf file that fits your available memory, then run it with a current llama.cpp
build:

llama-cli \
  -m /path/to/model.gguf \
  -p "Write a concise technical summary."

For vision or any-to-any models, download the matching multimodal projection file when one is
provided and follow the source model's modality-specific instructions.

Compatibility and responsible use

  • Use a runtime that explicitly supports this format, architecture, and modality.
  • Keep configuration, tokenizer, processor, projection, and weight files from the same revision together.
  • Review the source model card and license before redistribution or deployment.
  • Hardware needs depend on parameter count, context length, cache precision, quantization, and concurrency.
  • Report reproducible issues with the runtime version, hardware, launch command, and a minimal example.

Quantization or conversion changes numerical behavior, memory use, and throughput relative to the source checkpoint; validate quality on your own workload.

Generated outputs may be inaccurate or unsuitable for a given use case. Users are responsible for
testing behavior, applying appropriate safeguards, and complying with applicable licenses and laws.

Quantized GGUF files for Qwen3.6-35B-A3B-Abliterated-Heretic.

Available Quantizations

File Quant Size Notes
Qwen3.6-35B-A3B-Abliterated-Heretic-Q8_0.gguf Q8_0 ~35 GB Best quality, largest
Qwen3.6-35B-A3B-Abliterated-Heretic-Q6_K.gguf Q6_K ~27 GB High quality
Qwen3.6-35B-A3B-Abliterated-Heretic-Q5_K_M.gguf Q5_K_M ~24 GB Good quality/size balance
Qwen3.6-35B-A3B-Abliterated-Heretic-Q4_K_M.gguf Q4_K_M ~20 GB Recommended
Qwen3.6-35B-A3B-Abliterated-Heretic-Q4_K_S.gguf Q4_K_S ~19 GB Smaller Q4 variant
Qwen3.6-35B-A3B-Abliterated-Heretic-Q3_K_M.gguf Q3_K_M ~16 GB Lower bitrate
Qwen3.6-35B-A3B-Abliterated-Heretic-IQ4_XS.gguf IQ4_XS ~18 GB Importance-matrix quant
Qwen3.6-35B-A3B-Abliterated-Heretic-Q2_K.gguf Q2_K ~13 GB Smallest, lowest quality
mmproj-BF16.gguf BF16 ~861 MB Multimodal projection

Source

Usage

Use with llama.cpp, LM Studio, Ollama, or any GGUF-compatible inference engine.

# Example with llama.cpp
./llama-server -m Qwen3.6-35B-A3B-Abliterated-Heretic-Q4_K_M.gguf --mmproj mmproj-BF16.gguf -ngl 99

Model Details

Qwen3.6-35B-A3B is a Mixture-of-Experts model with 256 experts (8 active), totaling ~35B parameters but only ~3B active per token. Features SSM (State Space Model) layers alongside attention.

This "Abliterated-Heretic" version has had alignment/refusal training removed via abliteration techniques.

README history 5 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-22Polish model card overview and usage notes841d0074.4 KB
    Loading...
  2. 2026-08-22Polish model card overview and usage notesddda3694.4 KB
    Loading...
  3. 2026-08-22Polish model card overview and usage notescb8b6a23.6 KB
    Loading...
  4. 2026-04-22Upload README.md with huggingface_hub932b2842 KB
    Loading...
  5. 2026-04-21Add README1fd10881.4 KB
    Loading...

Discussions 1 thread

  1. 2026-09-24Original checkpoint availability for offline preservationopen1 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration