← back to catalog · registered 2026-08-22 13:56

markmonger/Qwen3.6-40B-Claude-4.6-Opus-Deckard-Heretic-Uncensored-Thinking-iMatrix-MTP-GGUF

markmonger Qwen 40B GGUF multimodal second-order 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/markmonger%2FQwen3.6-40B-Claude-4.6-Opus-Deckard-Heretic-Uncensored-Thinking-iMatrix-MTP-GGUF"
Response includes
  • classification m3
  • files 3
  • hub_downloads_all_time 3,327
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · lifetime
3K
309 last 30d - cooling
Likes
0
Model age
3mo ago
created 2026-06-24

Training datasets

Corpora the author lists in the model card. Datasets tracked in our /datasets catalog carry a category badge linking to the workflow stage. Others open on Hugging Face.

Downloads over time
Now3.6K→from1.1K↑216%
1K2K2.9K3.8K1.1K on Jul 13.6K on Oct 11JulAugSepOct
Jul 1 → Oct 11 · 56 snapshots · spans 102 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
gguf qwen3_5 qwen3.6 mtp speculative-decoding fine-tune unsloth heretic uncensored abliterated multi-stage tuned 40B
Total size
40.7 GB
Files
3
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-10-09 03:53

Files by quantization

Auxiliary files 3 files 40.7 GB
q8_0-v1.gguf 40.7 GB b05e40d7 download
README.md 2.89 KB 6a75636e download
.gitattributes 1.53 KB 73b008d1 download

README current version from Hugging Face


license: apache-2.0
tags:

  • qwen3_5
  • qwen3.6
  • gguf
  • mtp
  • speculative-decoding
  • fine-tune
  • unsloth
  • heretic
  • uncensored
  • abliterated
  • multi-stage tuned
  • 40B
  • dense
  • vision
  • multimodal
  • mmproj
  • long-context
    base_model: DavidAU/Qwen3.6-40B-Claude-4.6-Opus-Deckard-Heretic-Uncensored-Thinking-NEO-CODE-Di-IMatrix-MAX-GGUF
    datasets:
    • TeichAI/claude-4.5-opus-high-reasoning-250x
    • DavidAU/PkDick-Deckard-5-Datasets
    • allenai/WildChat-1M
      pipeline_tag: text-generation

This repository represents an effort to graft and re-train an MTP layer for DavidAU's Qwen3.6-40B model, inspired by Piehsoft's untrained MTP grafts.

Currently only offering q8_0 quantizations, but this will eventually expand once MTP retraining is finalized.

Version 1 (complete)

A simple retraining attempt that used the following procedure:

  1. Ran the 40B model on 117 random prompts
  2. Record what it was thinking at every position by introspecting llama.cpp
    • Specifically, recorded which 64 tokens it thought were most likely next and how confident it was in each.
    • Note: Substituted h_t with embedding(input_token) to move forward with the data that was immediately accessible.
  3. Use the recorded data to retrain the MTP layer: align what the MTP's predictions with reality
    • This data did not include hidden states, but since the trained head is shown real-world verifier logprobs during training its bias toward the most common tokens is effectively re-calibrated.

Compared to Piehsoft's untrained MTP head, this methodology

  • Significantly improved acceptances rates in structured text / coding scenarios
  • Provided little to no impact on high-entropy scenarios such as creative writing / casual chat.

Ultimately proved that improving acceptance rates via retraining was possible. Further improvements seem to require:

  • Larger corpus (dataset)
  • Higher text entropy within corpus
  • Most importantly, data collection of the verifier hidden states

Version 2 (in-progress)

EDIT: Training using hidden states ended up generating significantly worse results than the methodology used in v1. Alternative paths are being considered, but this may potentially end as a stalemate.

Training is in-progress, which builds on-top of the v1. This version likely won't be a final version, but rather will be used to help steer direction. The main changes will be:

  • Collect data using a modified version of llama.cpp that exposes data from the 40B model's hidden states (h_t)
  • Use a dataset with high entropy: WildChat-1M
    • Not training on the full dataset (would take a long time), will use a subset of randomly selected samples.
  • Increase epochs from 3 to 5 (and tweak a few other training settings)

README history 10 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-10-09Update README.mdd9dda242.8 KB
    Loading...
  2. 2026-06-25Update README.mdc3d03fc2.9 KB
    Loading...
  3. 2026-06-24Update README.md7979f8d3.4 KB
    Loading...
  4. 2026-06-24Update README.mdf423f893.3 KB
    Loading...
  5. 2026-06-24Update README.mdd50e0a43.3 KB
    Loading...
  6. 2026-06-24Update README.mdb6625d73.3 KB
    Loading...
  7. 2026-06-24Update README.md726d7bd3.3 KB
    Loading...
  8. 2026-06-24Update README.md64122be3.2 KB
    Loading...
  9. 2026-06-24Update README.md767d97d3.1 KB
    Loading...
  10. 2026-06-24initial commit01da52428 B
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration