← back to catalog · registered 2026-08-22 13:56

cjxzdzh/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP-GPTQ-INT4

cjxzdzh Qwen 24B second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/cjxzdzh%2FQwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP-GPTQ-INT4"
Response includes
  • classification m3
  • files 17
  • hub_downloads_all_time 3,915
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · lifetime
4K
231 last 30d - cooling
Likes
2
Model age
2mo ago
created 2026-08-02
Downloads over time
Now4K→from528↑655%
3551.7K3K4.3K528 on Aug 54K on Oct 11AugSepOct
Aug 5 → Oct 11 · 50 snapshots · spans 67 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en
Tags
gptqmodel safetensors qwen3_5 qwen3.6 gptq int4 quantized dynamic-quantization mtp multi-token-prediction uncensored reasoning

Related

Total size
18.2 GB
Files
17
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-02 12:22

Files by quantization

Auxiliary files 17 files 18.2 GB
model-00004-of-00005.safetensors 3.99 GB 31a6fac7 download
model-00002-of-00005.safetensors 3.98 GB edcce338 download
model-00003-of-00005.safetensors 3.97 GB 48858e22 download
model-00001-of-00005.safetensors 3.28 GB 05e96534 download
model-00005-of-00005.safetensors 3.00 GB bf57be63 download
tokenizer.json 19.1 MB 06b95093 download
model.safetensors.index.json 224 KB 4580b435 download
quant_log.csv 18.8 KB 09b8c9a0 download
chat_template.jinja 11.5 KB 82faea87 download
config.json 5.93 KB a40ecbaa download
README.md 4.04 KB ecdc8689 download
quantize_config.json 1.84 KB fbc0d54e download
.gitattributes 1.53 KB 52373fe2 download
tokenizer_config.json 1.29 KB bde3a7fd download
processor_config.json 1.16 KB 33818c7f download
generation_config.json 214 B dce86d49 download
generation_config.json.bak-default-1.0 214 B fd8faca8 download

README current version from Hugging Face


language:

  • en
    license: apache-2.0
    tags:
  • qwen3.6
  • gptq
  • int4
  • quantized
  • dynamic-quantization
  • mtp
  • multi-token-prediction
  • uncensored
  • reasoning
  • consumer-hardware
    pipeline_tag: text-generation
    base_model:
  • DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP
    library_name: gptqmodel
    inference_library_version: "7.3.2"
    quantization_config:
    quant_method: gptq
    bits: 4
    group_size: 128
    desc_act: true

Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP-GPTQ-INT4

GPTQ INT4 quantization of the original DavidAU/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP, optimized for dual-RTX-2080-Ti deployment.

Quantization Details

  • Method: GPTQ (via modelcloud/gptqmodel v7.3.2)
  • Base precision: INT4 (4-bit weights, group_size=128, symmetric, desc_act=true)
  • Dynamic per-layer override: Layers 68-79 projection layers quantized at INT8 instead of INT4 to preserve reasoning quality in the upper transformer layers
  • LM head: Not quantized (preserved at BF16)
  • Total size: ~19 GB (vs ~56 GB BF16 original)

Deployment Configuration

This quantization was specifically optimized and tested on a dual-RTX-2080-Ti (22GB each) setup with NVLink interconnect, running a custom vLLM fork designed for consumer GPUs.

  • Inference engine: weicj/vLLM-2080Ti-Definitive v0.1.14 (fork of vLLM with SM75-specific patches and INT4 GPTQ kernel optimizations)
  • GPU config: 2× NVIDIA RTX 2080 Ti (22GB VRAM each, compute capability 7.5), connected via NVLink (2 links per GPU, 25.781 GB/s each)
  • Tensor parallelism: TP=2
  • Context length: 262K tokens
  • MAX_NUM_SEQS: 2
  • GPU memory utilization: 0.93
  • MTP depth: 3

Benchmark Results

Measured on dual RTX 2080 Ti (NVLink) with the configuration above. All tests generate 1024 output tokens.

Prompt Length (tokens) TTFT (ms) ITL Avg (ms) ITL Std (ms) Prefill Time (ms) Prefill Speed (tok/s) Output Time (ms) Decode Speed (tok/s)
4,096 3,265 39.99 0.50 3,260 1,256 15,358 66.68
8,192 6,730 40.62 0.79 6,725 1,218 11,903 86.03
16,384 14,034 41.46 0.64 14,029 1,168 17,537 58.39
32,768 29,943 42.84 1.10 29,937 1,095 16,194 63.23
65,536 66,835 45.29 1.25 66,829 981 16,531 61.94
131,072 160,941 50.18 2.40 160,936 814 19,620 52.19
  • TTFT: Time To First Token
  • ITL: Inter-Token Latency (average between consecutive output tokens)
  • Prefill Speed: prompt tokens processed per second
  • Decode Speed: output tokens generated per second

Usage

Load with vLLM:

from vllm import LLM

llm = LLM(
model="cjxzdzh/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP-GPTQ-INT4",
quantize="gptq",
tensor_parallel_size=2,
max_model_len=262144,
gpu_memory_utilization=0.93,
max_num_seqs=2,
)

Load with GPTQModel directly:

from gptqmodel import GPTQModel, QuantizeConfig

model = GPTQModel.from_quantized(
"cjxzdzh/Qwen3.6-27B-Fable-Fusion-711-Uncensored-Heretic-NM-DAU-MTP-GPTQ-INT4",
device_map="auto",
)

Notes

  • The dynamic INT8 override on layers 68-79 projection layers is intentional and required for maintaining reasoning quality on this model family. This is a per-layer selective quantization strategy, not a uniform INT8 quantization.
  • The forked vLLM (weicj/vLLM-2080Ti-Definitive) includes critical patches for SM75 architecture and INT4 GPTQ kernel optimizations that are not present in upstream vLLM.
  • For benchmark results on the base model, refer to the original repository.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-02Upload GPTQ INT4 quantization - dual RTX 2080 Ti 22GB NVLink optimized8ee934b4 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration