← back to catalog · registered 2026-08-22 13:56

Noobito45/Qwen3.8-9B-heretic-uncensored-NVFP4-GGUF

Noobito45 Qwen 9B GGUF 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/Noobito45%2FQwen3.8-9B-heretic-uncensored-NVFP4-GGUF"
Response includes
  • classification m3
  • files 5
  • hub_downloads_all_time 74,781
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · lifetime
75K
24K last 30d - stable
Likes
37
Model age
7w ago
created 2026-08-18
Downloads over time
Now82.4K→from10K↑727%
030.1K60.3K90.4K10K on Aug 1982.4K on Oct 11AugSepOct
Aug 19 → Oct 11 · 49 snapshots · spans 53 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en
Tags
gguf qwen3.8 qwen3.5 heretic uncensored abliteration distillation reasoning text-generation en base_model:empero-ai/Qwen3.8-9B-Distill base_model:quantized:empero-ai/Qwen3.8-9B-Distill

Related

Total size
10.5 GB
Files
5
Quantizations
2
Registered
2026-08-22 13:56
Last updated on HF
2026-08-21 10:49

Files by quantization

BF16 1 file 879 MB
mmproj-BF16.gguf 879 MB 853698ce download
Auxiliary files 4 files 10.5 GB
Qwen3.8-9B-distill-heretic_nvfp4_q8_0.gguf 5.65 GB 74ae68a6 download
Qwen3.8-9B-distill-heretic_nvfp4_q4_k_m.gguf 4.82 GB c609376c download
README.md 5.55 KB 3755bcbe download
.gitattributes 1.85 KB 9be7d15d download

README current version from Hugging Face


license: apache-2.0
language:

  • en
    base_model:
  • empero-ai/Qwen3.8-9B
    pipeline_tag: text-generation
    tags:
    • qwen3.8
    • qwen3.5
    • heretic
    • uncensored
    • abliteration
    • distillation
    • reasoning

Qwen3.8-9B — Heretic / Uncensored

This is a decensored version of empero-ai/Qwen3.8-9B, created using Heretic v1.4.0.

The original model is a full-parameter distillation of Qwen3.8 2.4T A95B into the Qwen3.5-9B architecture. This repository does not reproduce the original model documentation;
please refer to the original model card for details about the model architecture, training, distillation dataset, and recommended usage.

Decensoring

The optimization run produced the following results:

Metric Heretic model Original model
KL divergence 0.0171 0 (by definition)
Refusals 22/100 100/100

Lower KL divergence indicates that the resulting model stays closer to the original model's behavior, while the refusal score measures how often the model refused the evaluation prompts.

Note: The model is relatively resistant to abliteration, making it difficult to reduce refusals without significantly increasing KL divergence.

The following parameters were obtained during the Heretic optimization:

Parameter Value
direction_index 17.82
attn.o_proj.max_weight 1.48
attn.o_proj.max_weight_position 19.00
attn.o_proj.min_weight 1.36
attn.o_proj.min_weight_distance 13.84
mlp.down_proj.max_weight 1.45
mlp.down_proj.max_weight_position 20.64
mlp.down_proj.min_weight 1.24
mlp.down_proj.min_weight_distance 10.59

Quantization

Quantized versions were produced from the resulting Heretic model.

The repository includes a quantized version using NVFP4 + Q8_0. The quantization process was evaluated separately from the BF16 model to measure the effect of quantization on general benchmark performance.

Note: The BF16 model is the reference version. The quantized version may exhibit small changes in benchmark scores and generation behavior due to reduced numerical precision.

Environment

Component Version / Specification
GPU NVIDIA RTX PRO 5000 48 GB
CUDA 12.8 (12.8.93)
PyTorch 2.9.1+cu128
Heretic v1.4.0
gguf-eval commit 87b8d31
llama.cpp b9968 + 8 commits (e3546c794)

Evaluation

General benchmark evaluation was performed using gguf-eval.

The original model and the Heretic model were evaluated in BF16, while the quantized model was evaluated separately.

Benchmark results

Test \ Model Original BF16 Heretic BF16 Heretic NVFP4 + Q8_0 Heretic NVFP4 + Q4_K_M
HellaSwag 77.75 78.75 77.25 78.00
Winogrande 72.38 72.53 70.40 70.96
MMLU 39.66 39.47 39.79 39.34
MMLU-Redux-2.0-Thinking 0.90 0.90 0.88 0.87
ARC-Challenge 52.84 52.51 52.17 52.84
PIQA 79.30 79.30 79.30 79.30
BoolQ 86.03 82.29 84.04 82.29
FLORES200* 50.19 50.25 49.71 49.96

Delta relative to the Original BF16 model:

Test \ Model Original BF16 Heretic BF16 Heretic NVFP4 + Q8_0 Heretic NVFP4 + Q4_K_M
HellaSwag 0.00 +1.00 −0.50 +0.25
Winogrande 0.00 +0.15 −1.98 −1.42
MMLU 0.00 −0.19 +0.13 −0.32
MMLU-Redux-2.0-Thinking 0.00 0.00 −0.02 −0.03
ARC-Challenge 0.00 −0.33 −0.67 0.00
PIQA 0.00 0.00 0.00 0.00
BoolQ 0.00 −3.74 −1.99 −3.74
FLORES200* 0.00 +0.06 −0.48 −0.23

Note: Delta represents the change in benchmark score relative to the Original BF16 baseline, which is 0 by definition.
Benchmark results may vary depending on the evaluation framework version, inference backend, hardware, and evaluation settings.
Results from other sources should therefore not be considered directly comparable unless the evaluation setup is equivalent.

FLORES200* — average over 5 language pairs, 101 sentences each: zh → en, kr → ru, it → fr, jp → de, en → ar

Reproducibility

The decensoring process is reproducible using Heretic v1.4.0 and the parameters listed above.

The important optimization parameters are included in this model card so that the transformation can be reproduced rather than treating the resulting weights as a black box.

For exact reproduction, use the original model as the starting point and apply the listed Heretic parameters with the corresponding Heretic version.


Usage

The model is provided as a quantized GGUF version of the Heretic BF16 model and can be used with GGUF-compatible inference engines such as llama.cpp.

For recommended generation settings and model-specific behavior, refer to the original Qwen3.8-9B model card.

Links

License

This model is released under the Apache-2.0 license, following the licensing of the underlying model.

README history 8 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-21Update README.md6a9d3ae5.6 KB
    Loading...
  2. 2026-08-20Update README.md286e0065.4 KB
    Loading...
  3. 2026-08-18Update README.mdd50b9d73.3 KB
    Loading...
  4. 2026-08-18Update README.md8741c063.4 KB
    Loading...
  5. 2026-08-18Update README.md1b43f993.4 KB
    Loading...
  6. 2026-08-18Update README.md9c20ec9245 B
    Loading...
  7. 2026-08-18Update README.md5d63538274 B
    Loading...
  8. 2026-08-18Create README.md4896e9f143 B
    Loading...

Discussions 4 threads

  1. 2026-08-27Qwen 3.5, not 3.8open2 💬#4
    Loading...
  2. 2026-08-27Is this any good?open2 💬#3
    Loading...
  3. 2026-08-22Thank you!open1 💬#2
    Loading...
  4. 2026-08-20It's not very uncensored.closed2 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration