← back to catalog · registered 2026-08-22 13:56

DBMe/Qwen3.5-9B-ultra-uncensored-heretic-v2-exl3

DBMe Qwen 9B second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/DBMe%2FQwen3.5-9B-ultra-uncensored-heretic-v2-exl3"
Response includes
  • classification m3
  • files 4
  • hub_downloads_all_time 112
  • author_summary 7 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · lifetime
112
33 last 30d - stable
Likes
0
Model age
6mo ago
created 2026-04-13
Downloads over time
Now126→from13↑869%
7519413713 on Apr 15126 on Oct 11126 on Oct 9AprMayJunJulAugSepOct
Apr 15 → Oct 11 · 65 snapshots · spans 179 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
exllamav3 exl3 text-generation base_model:llmfan46/Qwen3.5-9B-ultra-uncensored-heretic base_model:quantized:llmfan46/Qwen3.5-9B-ultra-uncensored-heretic license:apache-2.0 region:us

Related

Total size
0 B
Files
4
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-04-14 10:11

Files by quantization

Auxiliary files 4 files 78.6 KB
metrics_graph.png 72.4 KB f23a6801 download
README.md 3.40 KB 4bd23d4f download
.gitattributes 1.48 KB a6344aac download
metrics.json 1.34 KB b8b20d71 download

README current version from Hugging Face


base_model: llmfan46/Qwen3.5-9B-ultra-uncensored-heretic-v2
base_model_relation: quantized
quantized_by: DBMe
library_name: exllamav3
pipeline_tag: text-generation
license: apache-2.0
tags:

  • exl3

DBMe/Qwen3.5-9B-ultra-uncensored-heretic-v2-exl3

EXL3 (ExLlamaV3) quantizations of llmfan46/Qwen3.5-9B-ultra-uncensored-heretic-v2. All credit for the original model goes to the original authors.

📊 Available Quantizations & VRAM

The model weights are stored in separate branches. Please switch to a branch to download.
Note: VRAM estimates include PyTorch context overhead (~0.8GB) and assume an unquantized FP16 KV cache.

Target BPW Head BPW Branch (Download Link) WikiText-2 PPL (2048 ctx)¹ 2K ctx 4K ctx 8K ctx 16K ctx 32K ctx
4.0 h6 4.0bpw_h6 7.2847 ~7.74 GB ~7.99 GB ~8.49 GB ~9.49 GB ~11.49 GB
5.0 h6 5.0bpw_h6 7.2596 ~8.55 GB ~8.8 GB ~9.3 GB ~10.3 GB ~12.3 GB
6.0 h6 6.0bpw_h6 7.3097 ~9.35 GB ~9.6 GB ~10.1 GB ~11.1 GB ~13.1 GB
8.0 h8 8.0bpw_h8 7.2881 ~11.2 GB ~11.45 GB ~11.95 GB ~12.95 GB ~14.95 GB

¹ Evaluated against WikiText-2 with ExLlamaV3 using a strided 2048-token context window (-c 2048) in llama.cpp parity mode (-g). Lower is better.
(Higher BPW = higher quality, lower BPW = fits in less VRAM).

📥 How to Download

It's recommended to use the huggingface-cli to download specific branches. (Do not use git clone as it will download all branches!)

Ensure you have the CLI installed:

pip install -U "huggingface_hub[cli]"

Download a specific branch (e.g., 4.0bpw_h6):

# Example: Downloading the 4.0bpw_h6 branch
huggingface-cli download DBMe/Qwen3.5-9B-ultra-uncensored-heretic-v2-exl3 --revision 4.0bpw_h6 --local-dir Qwen3.5-9B-ultra-uncensored-heretic-v2-exl3-4.0bpw_h6

💻 Supported Engines

These models are highly optimized for modern GPUs and can be run using:

  • TabbyAPI: A fast, OpenAI-compatible API server. (Set model_name: "Qwen3.5-9B-ultra-uncensored-heretic-v2-exl3-<BranchName>" in your config)
  • Text-Generation-WebUI: A local web interface. (Select the exllamav3 loader)
  • ExLlamaV3 (Native): Python library for custom integration.

📈 Perplexity Degradation Curve

(Lower is better)
Perplexity Graph

⚙️ Advanced: Quantization Environment & Settings

🔬 Quantization Settings

  • Codebook: mcg

  • Output Scales: always

  • Calibration Rows: 250

  • Calibration Cols: 2048

  • Calibration Dataset: ExLlamaV3 Default (Wiki/C4/Code)

  • High Quality (HQ) Mode: False

  • ExLlamaV3: 0.0.29 (Commit: cb1a436)

  • Hardware: NVIDIA L4

README history 6 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-04-14Update README.md2f03c033.4 KB
    Loading...
  2. 2026-04-14Hotpatch: Update main README with VRAM Matrix216ac6c3.4 KB
    Loading...
  3. 2026-04-14Update main README with VRAM Matrix36aacad3.1 KB
    Loading...
  4. 2026-04-14Update main README with VRAM Matrixd7d0c213 KB
    Loading...
  5. 2026-04-13Update main README with VRAM Matrix157ecbe2.8 KB
    Loading...
  6. 2026-04-13Update main README with VRAM Matrix475143b2.6 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration