← back to catalog · registered 2026-08-25 16:02

sheppo/Qwen3.8-27B-Heretic-Abliterated-W4A16-A100

sheppo Qwen 24B GGUF
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/sheppo%2FQwen3.8-27B-Heretic-Abliterated-W4A16-A100"
Response includes
  • classification m3
  • files 17
  • hub_downloads_all_time 722
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · lifetime
722
312 last 30d - stable
Likes
0
Model age
6w ago
created 2026-08-25
Downloads over time
Now811→from54↑1,402%
1630659788754 on Aug 26811 on Oct 11AugSepOct
Aug 26 → Oct 11 · 47 snapshots · spans 46 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
transformers safetensors qwen3_5_text text-generation qwen3 qwen3.8 vllm compressed-tensors autoround gptq w4a16 a100

Related

Total size
16.5 GB
Files
17
Quantizations
1
Registered
2026-08-25 16:02
Last updated on HF
2026-08-26 11:32

Files by quantization

Auxiliary files 17 files 16.5 GB
model-00001-of-00006.safetensors 2.99 GB 92f88d9f download
model-00002-of-00006.safetensors 2.98 GB e4136fb0 download
model-00003-of-00006.safetensors 2.98 GB cb6ba954 download
model-00004-of-00006.safetensors 2.79 GB 2e2b3d4e download
model-00005-of-00006.safetensors 2.37 GB 84867c3e download
model-00006-of-00006.safetensors 2.37 GB 45c264e8 download
tokenizer.json 19.1 MB 06b95093 download
model.safetensors.index.json 142 KB c9c7fc3d download
LICENSE 11.1 KB 2f961e65 download
chat_template.jinja 8.74 KB c0c686f9 download
config.json 7.91 KB b17c56c8 download
README.md 5.67 KB e8ccf058 download
quantization_config.json 4.95 KB 91da6e05 download
artifact-manifest.json 1.92 KB 839babad download
.gitattributes 1.53 KB 52373fe2 download
tokenizer_config.json 1.10 KB d1a20cc3 download
generation_config.json 205 B 247fe8e6 download

README current version from Hugging Face


license: apache-2.0
base_model:

  • Qwen/Qwen3.8-27B
  • 0bserverx/Qwen3.8-27B-Heretic-Abliterated-Uncensored-GGUF
    library_name: transformers
    pipeline_tag: text-generation
    tags:
  • qwen3
  • qwen3.8
  • text-generation
  • vllm
  • compressed-tensors
  • autoround
  • gptq
  • w4a16
  • a100
  • speculative-decoding

Qwen3.8-27B Heretic Abliterated W4A16 — A100 tested

W4A16 AutoRound checkpoint derived from the exact RVN BF16 GGUF and tested on
one A100-SXM4-80GB. It measured 73.15 tok/s raw and 138.38 tok/s with the
companion DFlash2 runtime on the frozen sampled-prose benchmark, a 1.89x
speedup.

This Hugging Face repo contains the W4A16 target. The accelerated runtime,
DFlash2 drafter, exact launch settings, and base-Qwen A100 results live in
elsheppo/qwen38-a100-fastpath.

What this model is

This is a text-only, vLLM-compatible W4A16 AutoRound derivative of the exact
RVN-BF16.gguf file from
0bserverx/Qwen3.8-27B-Heretic-Abliterated-Uncensored-GGUF.

  • Source revision: 8581a3e4cd8cdeca9bb6709d81ebcbcdfc93fe43
  • Source file: RVN-BF16.gguf
  • Source size: 53,808,272,896 bytes
  • Source SHA-256:
    fe3cb9c7d067f0016fb8ab0150e2b226c4dfb4e63b497e46ada1e7744b035cdc
  • Quantization: symmetric W4A16, group size 128
  • AutoRound revision: 96ce448039b3c36fa879b9f4c740a8ee50c0f9ba
  • Seed: 42
  • Serving payload: 13 model/tokenizer/config files, 17,702,015,479 bytes
  • Base-model fast overlay applied: no

That final line matters. The convenient base-Qwen overlay replaces tensors.
Applying it here could erase the RVN modifications we were trying to preserve.

How the GGUF became this checkpoint

The first attempt matched all 851 tensor names and shapes, loaded in vLLM, and
still produced corrupted multilingual symbols.

The failure was upstream of quantization: the Qwen3.8 GGUF converter changes
several tensor representations. The corrected bridge inverted the non-GDN norm
shift, continuous-time A_log transform, and Gated DeltaNet value-head layout
before AutoRound saw the model. It then mapped all 851 source and target tensors,
checked that every value was finite, and produced the expected deterministic
BF16 answer before quantization resumed.

The user-facing model is Qwen3.8-27B. Qwen3_5ForCausalLM labels inside the
config and runtime are the architecture names used by the official release.

The A100 result

The primary cell used one A100-SXM4-80GB, concurrency one, temperature 0.8,
top-p 0.95, top-k 20, thinking off, and a fixed eight-prompt prose workload.

Serving mode Median committed tok/s Mean committed tok/s Median TTFT
Raw RVN W4A16 73.15 73.31 132 ms
RVN + DFlash2, k=7 138.38 148.89 155 ms

The median speedup is 1.89x. We had frozen a 140 tok/s target before the
run; this landed 1.62 tok/s short, so the experiment record says
promising-below-target.

A separate greedy story fixture measured 152.30 tok/s at C1 and 609.87 tok/s
aggregate at C8. Different workload, different measurement. The primary prose
number stays 138.38.

Does it still behave like the model?

The checkpoint passed:

  • six-shard vLLM load and finite-logit semantic smoke
  • exact chat-template response
  • structured JSON output
  • automatic tool-call parsing
  • completion and token-accounting checks
  • clean-process restart
  • a 36-pair blind sampled-writing comparison

Claude Fable judged all 36 pairs without knowing which side was raw or DFlash2,
then repeated the judgments after every A/B position was swapped. The same
content side won 35 of 36 calls across the swap. After unblinding, the preferred
responses were nearly evenly split between serving modes: raw led by two pairs
in the first presentation and three in the swapped presentation.

We observed no material writing-quality regression from the tested DFlash2
path. That conclusion belongs to this sampling profile and this bounded
battery.

Run it

git clone https://github.com/elsheppo/qwen38-a100-fastpath
cd qwen38-a100-fastpath
./scripts/bootstrap.sh rvn
./scripts/serve.sh rvn dflash2
./scripts/smoke.sh

The fast profile uses
syvai/Qwen3.8-27B-DFlash2-W4A16
at revision 4d30ec736ffc6b8688dc2ae2b502d9b48bdec279, seven draft
tokens, lookup drafting disabled, and a fixed 8 GiB state/KV pool.

Run the target without DFlash2 for comparison:

./scripts/serve.sh rvn raw

Caveats

  • This exact RVN artifact was verified on A100 80 GB. The companion project
    separately verified regular Qwen on A100 40 GB.
  • The serving profile is text-only and launches with --language-model-only.
  • Code and repetitive structured output accept more speculative drafts than
    sampled prose. Throughput moves with the workload.
  • Stochastic responses can differ between raw, speculative, and dynamically
    batched serving. The quality battery looks for degradation, while byte
    identity falls outside its job.
  • The conversion reproduces functionally. A second print matched source
    identity, tensor mapping, file paths, sizes, vLLM load, and semantic smoke.
    Four of fourteen total output-path hashes changed after AutoRound warned that
    Flash Attention was nondeterministic.
  • Abliteration changes model behavior. Test the checkpoint against your own
    prompts and requirements.

Credit

The underlying work comes from Qwen, the RVN/Heretic source authors,
AutoRound, vLLM, the DFlash2 authors, and the syv-ai serving project. The
companion repository preserves exact revisions and links in its NOTICE.

Apache-2.0, following the Qwen base model and the source model card.

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-26docs: add optional vision fast path72e48cf6.4 KB
    Loading...
  2. 2026-08-25release: A100-tested W4A16 checkpointa13afb35.7 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration