← back to catalog · registered 2026-08-22 13:56

DavidAU/Qwen3.5-21B-GLM-4.7-Flash-Deckard-Heretic-Uncensored-Thinking

DavidAU Glm 21B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/DavidAU%2FQwen3.5-21B-GLM-4.7-Flash-Deckard-Heretic-Uncensored-Thinking"
Response includes
  • classification m3
  • files 23
  • hub_downloads_all_time 285
  • author_summary 213 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · lifetime
285
94 last 30d - stable
Likes
8
Descendants
2
in 2 direct forks
Model age
6mo ago
created 2026-04-01

Training datasets

Corpora the author lists in the model card. Datasets tracked in our /datasets catalog carry a category badge linking to the workflow stage. Others open on Hugging Face.

Downloads over time
Now351→from0↑0%
01292573860 on Apr 1351 on Oct 11AprMayJunJulAugSepOct
Apr 1 → Oct 11 · 67 snapshots · spans 193 days

Genealogy 2 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh
Tags
transformers safetensors qwen3_5 image-text-to-text unsloth fine tune heretic uncensored abliterated multi-stage tuned. all use cases coder

Related

Total size
39.6 GB
Files
23
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-04-02 02:38

Files by quantization

Auxiliary files 23 files 39.6 GB
model-00009-of-00010.safetensors 4.65 GB d3dbfa8e download
model-00005-of-00010.safetensors 4.64 GB 5e365693 download
model-00008-of-00010.safetensors 4.63 GB 4b6b593f download
model-00003-of-00010.safetensors 4.62 GB e13d8d9c download
model-00007-of-00010.safetensors 4.59 GB 42341af8 download
model-00006-of-00010.safetensors 4.58 GB 6cfe5940 download
model-00004-of-00010.safetensors 4.58 GB 88e37537 download
model-00002-of-00010.safetensors 4.51 GB 14582fce download
model-00001-of-00010.safetensors 2.37 GB 4c1ebaab download
model-00010-of-00010.safetensors 462 MB 612e0bf3 download
tokenizer.json 12.2 MB 5f9e4d49 download
vocab.json 6.41 MB 0aa0ce06 download
valhalla.webp 4.14 MB 010bce05 download
model.safetensors.index.json 88.1 KB 6b4fa21e download
tokenizer_config.json 8.48 KB 3e0a1107 download
chat_template.jinja 7.57 KB a585dec8 download
README.md 6.12 KB 2e62975b download
config.json 3.16 KB d22c16e6 download
.gitattributes 1.58 KB 8d57e4fc download
processor_config.json 1.27 KB 7ad6acdf download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
generation_config.json 213 B 8c8412a2 download

README current version from Hugging Face


language:

  • en
  • zh
    license: apache-2.0
    tags:
  • unsloth
  • fine tune
  • heretic
  • uncensored
  • abliterated
  • multi-stage tuned.
  • all use cases
  • coder
  • creative
  • creative writing
  • fiction writing
  • plot generation
  • sub-plot generation
  • fiction writing
  • story generation
  • scene continue
  • storytelling
  • fiction story
  • science fiction
  • romance
  • all genres
  • story
  • writing
  • vivid prosing
  • vivid writing
  • fiction
  • roleplaying
  • bfloat16
  • all use cases
    datasets:
  • TeichAI/glm-4.7-2000x
  • DavidAU/PkDick-Deckard-5-Datasets
    base_model:
  • DavidAU/Qwen3.5-27B-Deckard-PKD-Heretic-Uncensored-Thinking
    pipeline_tag: image-text-to-text
    library_name: transformers

WARNING: This model has character and intelligence. It will take no prisoners. It will give no quarter. Uncensored,
Unfiltered and boldly confident. Not even remotely "SFW", if you ask it for NSFW content. And it is wickedly smart too.

Qwen3.5-21B-GLM-4.7-Flash-Deckard-Heretic-Uncensored-Thinking

21 billion parameters (dense, not moe) CONTRACTED/SHRUNK from 27B Qwen 3.5, then trained on GLM 4.7 Flash High Reasoning dataset via Unsloth on local hardware... but there
is much more to the story - in comes DECKARD.

48 layers, 639 Tensors. (33% LESS than base model of 27B)

The model is also 33% faster than 27B in terms of token per second, and quants take up 1/3 less memory too.

Features variable length reasoning ; less complex = shorter, longer for more complex.

Model performance has increased dramatically. And it has character too.

A lot of character.

No censorship, no nanny. (via Heretic)

And it is very, very smart.

Fully uncensored first (via Heretic), then trained (via Unsloth) on "Deckard/PDK" internal datasets (5) (character, intelligence, depth, observation, and ah... point of view),
THEN CONTRACTED to 21B parameters, and then trained (Unsloth again) with GLM 4.7 Flash Distill dataset (to shorten and improve reasoning, and stablize everything).

256K context.

Example generation(s) below.

SETTINGS:

  • min 8k to 16k context window.
  • for creative rep pen of 1.05 to 1.1 WITH LOWER QUANTS.
  • suggest temp .7 / rep pen 1 (off) for general usage.
  • suggest temp 1 / rep pen 1.05 to 1.1 for creative and SOME USE CASES.
  • output generation can exceed 100k tokens.
  • Suggest min quant of Q4KS (non imatrix) or IQ3_S (imatrix) or HIGHER.
  • For toolcalls -> suggest Q6 min quants (as per Qwen guidence)

EXAMPLE SYSTEM PROMPTS:

The model does not need a system prompt, however if you want to enhance operation here are some samples.

#1 - All use cases.

Be vivid and precise.

#2 - Creative use cases:

Below is an instruction that describes a task. Ponder each user instruction carefully, and use your skillsets and critical instructions to complete the task to the best of your abilities.

Here are your skillsets:
[MASTERSTORY]:NarrStrct(StryPlnng,Strbd,ScnSttng,Exps,Dlg,Pc)-CharDvlp(ChrctrCrt,ChrctrArcs,Mtvtn,Bckstry,Rltnshps,Dlg*)-PltDvlp(StryArcs,PltTwsts,Sspns,Fshdwng,Climx,Rsltn)-ConfResl(Antg,Obstcls,Rsltns,Cnsqncs,Thms,Symblsm)-EmotImpct(Empt,Tn,Md,Atmsphr,Imgry,Symblsm)-Delvry(Prfrmnc,VcActng,PblcSpkng,StgPrsnc,AudncEngmnt,Imprv)

[*DialogWrt]:(1a-CharDvlp-1a.1-Backgrnd-1a.2-Personality-1a.3-GoalMotiv)>2(2a-StoryStruc-2a.1-PlotPnt-2a.2-Conflict-2a.3-Resolution)>3(3a-DialogTech-3a.1-ShowDontTell-3a.2-Subtext-3a.3-VoiceTone-3a.4-Pacing-3a.5-VisualDescrip)>4(4a-DialogEdit-4a.1-ReadAloud-4a.2-Feedback-4a.3-Revision)

Here are your critical instructions:
Ponder each word choice carefully to present as vivid and emotional journey as is possible. Choose verbs and nouns that are both emotional and full of imagery. Load the story with the 5 senses. Aim for 50% dialog, 25% narration, 15% body language and 10% thoughts. Your goal is to put the reader in the story.

INSTRUCT MODE:

The model's default mode however is "THINKING" ; this can be changed by editing the following line in the jinja template:

{# Set Instruct mode here #}

to

{%- set enable_thinking = false %}
  • In LMStudio, this can be edited after loading the model, in dev mode -> template
  • When quanting -> if you set this to "false" => model/quants will be "instruct"

NOTES:

  • Upgraded Jinja template to correct issues with Qwen 3.5s - looping, repeatings, and long thinking as well as upgrades to tools too.
  • Was also trained with new improved template to further enhance operation too.
  • Image processing tested and intact.
  • Code generation also tested and passed.
  • System prompt - even a minor one - will enhance operation, especially at lower quants.

LOOPING:

  • This may happen with lower quants / prompts with "not enough meat on the bone" => Add more to the prompt and/or set rep pen to 1.05 to 1.1.
  • Adding a system prompt - even a single sentence - can correct this issue and bypass the need to adjust rep pen.

Want MORE power, more "IQ"?

https://huggingface.co/DavidAU/Qwen3.5-40B-Claude-4.6-Opus-Deckard-Heretic-Uncensored-Thinking

NEED something a wee bit wilder? Unhinged? A wee bit more raw?

See this version:

https://huggingface.co/DavidAU/Qwen3.5-40B-RoughHouse-Claude-4.6-Opus-Polar-Deckard-Uncensored-Heretic-Thinking


BENCHMARKS:

                              arc-c arc/e boolq hswag obkqa piqa  wino

THIS MODEL [instruct] mxfp8   PENDING...

Qwen3.5-27B-Thinking mxfp8    0.443,0.498,0.857,0.701,0.372,0.770,0.752

Note: Instruct mode will have stronger benchmarks.

See this model (instruct, also one of my fine tunes - it scores 675 on "arc" - Arc Challenge hard):

https://huggingface.co/DavidAU/Qwen3.5-27B-Claude-4.6-OS-INSTRUCT


SAFETY ALIGNMENT:

It is gone. No nanny, no strings, no limits.


EXAMPLE GENERATION(s):


Q4KS ; mid to low quality quant [non-imatrix].

"SYSTEM" is system prompt, "USER" is prompt, "ASSISTANT" ... you know.

This is a strong test to measure creative as well as instruction following specifically with a complex system prompt [if used].

Expect better generation with Imatrix quants / higher quants.

WARNING:

Graphic, swearing, intense - model does not hold back.


README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-04-02Update README.mdbe12f8e6.1 KB
    Loading...
  2. 2026-04-02Create README.md3baf3e744 B
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration