← back to catalog · registered 2026-09-28 08:57

derikn/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-AWQ-INT4

derikn 27B multimodal second-order
curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/derikn%2FQwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU-AWQ-INT4"
Response includes
  • classification m3
  • files 21
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-09-28

Training datasets

Corpora the author lists in the model card. Datasets tracked in our /datasets catalog carry a category badge linking to the workflow stage. Others open on Hugging Face.

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
transformers safetensors qwen3_5 image-text-to-text qwen qwen3.8 heretic uncensored finetune awq int4 4-bit

Related

Total size
18.2 GB
Files
21
Quantizations
1
Registered
2026-09-28 08:57
Last updated on HF
2026-09-28 07:30

Files by quantization

Auxiliary files 21 files 18.2 GB
model-00003-of-00007.safetensors 3.00 GB 5365be16 download
model-00004-of-00007.safetensors 3.00 GB f54e1df7 download
model-00001-of-00007.safetensors 3.00 GB 3d33846a download
model-00002-of-00007.safetensors 2.97 GB 4770fb54 download
model-00006-of-00007.safetensors 2.37 GB ab536a99 download
model-00007-of-00007.safetensors 2.37 GB 7f82c172 download
model_extra_tensors.safetensors 810 MB 4eea6f7a download
model-00005-of-00007.safetensors 734 MB 4fe0d88f download
tokenizer.json 19.1 MB 87a7830d download
model.safetensors.index.json 187 KB 74aba99f download
LICENSE 11.3 KB f938136e download
config.json 10.3 KB 434b91cb download
chat_template.jinja 8.74 KB 5a39aa2b download
README.md 8.03 KB a4f8826a download
quantization_config.json 6.12 KB 647b4a88 download
NOTICE.txt 3.24 KB a1cbc29a download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.19 KB 43c4343e download
tokenizer_config.json 1.17 KB 4d9ac0cf download
preprocessor_config.json 443 B 8ed39680 download
generation_config.json 214 B c53835dc download

README current version from Hugging Face


base_model: DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU
base_model_relation: quantized
library_name: transformers
pipeline_tag: image-text-to-text
license: apache-2.0
datasets:

  • NeelNanda/pile-10k
    tags:
  • qwen
  • qwen3.8
  • heretic
  • uncensored
  • finetune
  • awq
  • int4
  • 4-bit
  • auto-round
  • vllm
  • xpu
  • intel-arc
  • image-text-to-text
  • vision
  • mtp
  • reasoning
  • tool-calling
    quantized_by: derikn

Qwen3.8-27B TURBO-Fable Cold-Fusion 735-882 Heretic Uncensored AWQ INT4 (W4A16, group 128)

4-bit weight-only quantization of
DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU,
built with Intel's AutoRound and tested on Intel Arc Pro B70 GPUs.

This is an unofficial community quantization by derikn. DavidAU did not author it and does not endorse
it. For the model itself, its capabilities and its evaluation results, read the
base model card.

Two things to know about the base model before you use this: it is a community fine-tune of
Qwen/Qwen3.8-27B, not the upstream model, and it is an uncensored build. Its outputs are not
filtered for safety. Quality and behaviour come from the fine-tune, not from this quantization.

What this is

Base model DavidAU/Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAU (Apache 2.0 as declared)
Upstream model Qwen/Qwen3.8-27B (Apache 2.0)
Parameters 27,781,427,952 (see the note on the parameter count below)
Quantization AWQ, weight-only, 4-bit (W4A16)
Group size 128, symmetric, gemm version
Method AutoRound 0.15.1, auto_awq export
Kept at 16-bit lm_head, the vision tower, the MTP head, the linear-attention projections, the visual merger
Native context 262,144 tokens, same as the base model
Vision the vision tower and the Qwen2VL image processor are intact, so image inputs work
MTP the head is retained at 16-bit and usable, see the MTP section below
Chat template included, unchanged from the base model

Only the linear weights are quantized. No fine-tuning and no extra training data went into this: it is
the base model's weights at 4 bits.

101 modules are held at 16-bit, listed in quantization_config.modules_to_not_convert: the 96
linear-attention in_proj_a and in_proj_b projections, the 2 visual merger layers, the vision blocks,
the MTP head, and lm_head.

Note on the parameter count

The Safetensors panel on this page reports a model size well below the real one. That is not the size of
the model. The 4-bit weights are packed into int32 tensors, eight weights per int32, and the panel's
total counts that packed storage instead of the parameters it encodes.

This checkpoint holds 27,781,427,952 parameters, the same count as the base model. The Hub API's own
per-dtype breakdown adds up to that figure while the panel's single total does not. I32 and BF16
describe storage, not the precision of the model: packed 4-bit weights, and the tensors held at
bfloat16.

How it was made

Quantized from the bf16 checkpoint with AutoRound 0.15.1 on a single Intel Arc Pro B70, with weights
streamed to the card so VRAM stays bounded. The run was driven from the Quantize tab of a private
dashboard, so this is the pipeline's invocation rather than a script you can run as-is:

python quantize.py \
  --model <base-model-dir> \
  --format auto_awq --device xpu:0 --low-gpu-mem \
  --n_samples 128 --seqlen 1024 --batch_size 8

Calibration used the quantizer's default set, NeelNanda/pile-10k, at 128 samples of 1024 tokens
each. The wrapper is not published, so that exact command will not run for you. These are the AutoRound
settings it passes, which are the whole recipe:

  • export format auto_awq: AWQ, 4-bit weight-only, group size 128, symmetric, gemm version
  • lm_head, the vision tower, the MTP head, the linear-attention projections and the visual merger excluded
  • calibration on NeelNanda/pile-10k, 128 samples of 1024 tokens, batch size 8
  • weights streamed to the device during tuning, to keep VRAM bounded

Running it

Tested on vLLM XPU (vllm/vllm-openai-xpu:nightly) with tensor parallelism across two Arc Pro B70s,
an fp8 KV cache and the full 256K context:

vllm serve /models/Qwen3.8-27B-TURBO-Fable-AWQ-INT4 \
  --served-model-name qwen3.8-27b-turbo-fable-awq \
  --quantization awq --dtype bfloat16 \
  --max-model-len 262144 --kv-cache-dtype fp8_e4m3 \
  --block-size 64 --max-num-seqs 8 --max-num-batched-tokens 16384 \
  --gpu-memory-utilization 0.91 \
  --enable-prefix-caching \
  --tool-call-parser qwen3_coder --enable-auto-tool-choice \
  --reasoning-parser qwen3 \
  --speculative-config '{"method":"mtp","num_speculative_tokens":3}' \
  --trust-remote-code --tensor-parallel-size 2

Container settings that matter on Arc: --device=/dev/dri, --group-add for the render and video
groups, --ipc=host, --security-opt seccomp=unconfined, and ZE_AFFINITY_MASK set to the cards you
want to use. Without the device, group and seccomp combination Level Zero reports zero devices.

MTP speculative decoding

The MTP head ships at 16-bit beside the weights, and quantization_config declares it unquantized
through the mtp entry in modules_to_not_convert.

That entry is required. AutoRound quantizes the module tree it loads, and the MTP head sits outside it,
so without the exclusion vLLM builds the head as a packed 4-bit linear and the draft model fails at load:

ValueError: There is no module or parameter named 'fc.weight' in Qwen3_5MultiTokenPredictor

With the exclusion in place the head loads alongside the target model, and speculative decoding
drafts. Measured on this checkpoint: about 77 percent draft acceptance (77,712 draft tokens against
59,609 accepted) with 3 speculative tokens.

Thinking and tool calling

Tested with the qwen3_coder tool parser and the qwen3 reasoning parser, with thinking controlled
through the chat template:

--default-chat-template-kwargs '{"enable_thinking": true, "reasoning_preserve": true, "reasoning_effort": "low"}'

enable_thinking turns thinking on or off, and reasoning_effort accepts the OpenAI scale that vLLM
forwards to the template.

What was tested, and what was not

Tested:

  • 256K context loads and serves, at tensor parallel 2 with an fp8 KV cache.
  • MTP speculative decoding loads and drafts at about 77 percent acceptance, from the retained MTP head.
  • Tool calling returns a structured call, and thinking separates from the answer: vLLM reports the
    reasoning in message.reasoning_content and the reply in message.content.
  • Text generation returns correct answers, and the vision tower (333 tensors) plus the Qwen2VL
    processor config load, so image input remains available.

Not tested:

  • No accuracy or throughput numbers, and no image-input evaluation, are published with this repository.
    The quantizer's own measurements come from one home-lab setup (Arc Pro B70, XPU vLLM nightly, fp8 KV
    cache) and are not a fair comparison for anyone running different hardware or software. Model quality
    is the base model's, so read its card.

Licence and attribution

Apache License 2.0. The base model declares Apache 2.0 but its repository ships no licence file, so
LICENSE carries the Apache 2.0 text published with the upstream model, and NOTICE.txt states the
position in full.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.