← back to catalog · registered 2026-08-22 13:56

noon-at-cgn/Qwen3.8-27B-Uncensored-W4A16-AutoRound

noon-at-cgn Qwen 24B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/noon-at-cgn%2FQwen3.8-27B-Uncensored-W4A16-AutoRound"
Response includes
  • classification m1
  • files 19
  • hub_downloads_all_time 29,435
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
29K
8K last 30d - stable
Likes
5
Descendants
3
in 3 direct forks
Model age
7w ago
created 2026-08-21
Downloads over time
Now30.1K→from20↑150,435%
011K22.1K33.1K20 on Aug 1930.1K on Oct 11AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Genealogy 3 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh
Tags
safetensors qwen3_5 qwen3.8 uncensored abliterated quantized w4a16 autoround int4 vision-language vllm image-text-to-text

Related

Total size
18.1 GB
Files
19
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-09-04 14:18

Files by quantization

Auxiliary files 19 files 18.1 GB
model-00004-of-00007.safetensors 3.00 GB ad40637c download
model-00001-of-00007.safetensors 2.99 GB 098ab1bf download
model-00002-of-00007.safetensors 2.98 GB 34f6f324 download
model-00003-of-00007.safetensors 2.98 GB 1b71bc62 download
model-00006-of-00007.safetensors 2.37 GB da51fa48 download
model-00007-of-00007.safetensors 2.37 GB 6866cf8a download
model_extra_tensors.safetensors 810 MB 7a9fd4ee download
model-00005-of-00007.safetensors 667 MB 0994bb8d download
tokenizer.json 19.1 MB 06b95093 download
model.safetensors.index.json 194 KB e3a119c1 download
config.json 20.0 KB 6503941a download
quantization_config.json 15.8 KB 263f7f84 download
chat_template.jinja 8.74 KB c0c686f9 download
README.md 5.07 KB d7cca010 download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.16 KB 33818c7f download
tokenizer_config.json 1.14 KB 1d134cd2 download
preprocessor_config.json 390 B 2ea84a43 download
generation_config.json 214 B 3f9de11a download

README current version from Hugging Face


license: apache-2.0
base_model: orcarouter/Qwen3.8-27B-Uncensored
base_model_relation: quantized
language:

  • en
  • zh
    pipeline_tag: image-text-to-text
    tags:
  • qwen3.8
  • uncensored
  • abliterated
  • quantized
  • w4a16
  • autoround
  • int4
  • vision-language
  • vllm

Qwen3.8-27B-Uncensored-W4A16-AutoRound

W4A16 quantization of orcarouter/Qwen3.8-27B-Uncensored
(an abliterated derivative of Qwen/Qwen3.8-27B),
produced with Intel AutoRound. Retains
the model's full native 262,144-token context and vision capability — the
vision tower and MTP speculative-decoding head are both preserved,
unquantized where it matters (see below).

[!WARNING]
This model has had safety alignment substantially removed via abliteration
(inherited from the base model this checkpoint quantizes). It will comply
with harmful, unethical, offensive, or illegal requests that an aligned
model would refuse, and has no meaningful built-in guardrails. Released
strictly for legitimate research — interpretability, AI safety study,
red-teaming, evaluation — and adaptation into your own guarded pipeline.
Do not deploy to end users without your own safety layer. Users assume
full responsibility for outputs; the authors and uploaders accept no
liability for misuse or harm arising from this model.

Quantization recipe

scheme=W4A16, dataset=NeelNanda/pile-10k, nsamples=128, seqlen=2048,
batch_size=4, iters=200, seed=42, quant_nontext_module=False.

Excluded from quantization (kept at bf16): lm_head, the GatedDeltaNet
in_proj_a/in_proj_b projections on every linear-attention layer, and the
entire visual.* vision tower. embed_tokens is unquantized too (it's an
embedding table, not a Linear). This mirrors the recipe
dbirks/Qwen3.8-27B-W4A16-AutoRound
used for the official (non-abliterated) base model, applied here to the
abliterated checkpoint instead.

Known deviation from that reference recipe: AutoRound 0.14.2 here vs. their
0.15.0 (not yet released at quantization time).

Eval results

Run with lm_eval (EleutherAI harness) against this checkpoint via vLLM,
thinking mode on, true sampling (temperature=1.0, top_p=0.95, top_k=20)
unless noted. Reference columns are the closest published numbers found for
comparison, not a guaranteed apples-to-apples setup — see notes.

Benchmark This model Reference Notes
GSM8K (flexible-extract, n=1319) 0.9052 dbirks BF16 .911 / int4 .917 close match
MMLU-Pro (14 subjects x 100, 5-shot CoT) 0.761 dbirks int4 .826 same n/methodology, non-abliterated base — see disclaimer below
MMLU (57 subjects x 6, n=342) 0.880 orcarouter FP8-quant .843 matches/exceeds; wide per-subject stderr at n=6

MMLU-Pro gap disclaimer: the ~6.5pt gap to dbirks' int4-of-official-base
number is the most directly comparable reference (same quantization
aggressiveness, same task/shot setup) and likely reflects abliteration's own
cost to capability (orcarouter's own card reports abliteration costs
~0.6-1.3pts on several benchmarks before any quantization), not a defect in
this quantization. Against orcarouter's own FP8 quant of the same abliterated
base, this checkpoint's MMLU-Pro is in the same range.

Safety / refusal (thinking OFF)

Custom rule-based refusal-classification eval, same datasets and n as
orcarouter's model card
where available (AdvBench, JailbreakBench, StrongREJECT, HarmBench,
MaliciousInstruct, SimpleSafetyTests, ForbiddenQuestions, XSTest-safe for
over-refusal). Not a byte-exact reproduction — sample indices differ — but
same source datasets, same sample sizes, same style of opening-phrase
refusal classifier.

Benchmark n This model Card reference
AdvBench 100 0.0% 0.0%
JailbreakBench (harmful) 100 0.0% 0.0%
StrongREJECT 150 0.0% 2.0%
HarmBench (standard) 150 0.7% 2.7%
MaliciousInstruct 100 0.0% 0.0%
SimpleSafetyTests 50 8.0% 6.0%
ForbiddenQuestions 150 3.3% 4.7%
XSTest-safe (over-refusal, lower is better) 250 0.0% 0.4%

With thinking ON, refusal was 0.0% across all eight benchmarks (n=60 each,
except SimpleSafetyTests n=50 and XSTest n=250) — matches the card's own
pattern of thinking mode reducing refusal further.

Usage (vLLM)

pip install vllm==0.27.1
vllm serve noon-at-cgn/Qwen3.8-27B-Uncensored-W4A16-AutoRound \
  --enable-auto-tool-choice --tool-call-parser qwen3_coder

Adjust --tensor-parallel-size, --max-model-len, and
--gpu-memory-utilization for your own hardware — full 262,144-token
context plus vision needs roughly 40GB+ of VRAM depending on how much
concurrency/KV headroom you need.

License

Apache-2.0, inherited from Qwen/Qwen3.8-27B
via orcarouter/Qwen3.8-27B-Uncensored.

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-09-04Update README.md8b3f5d05 KB
    Loading...
  2. 2026-08-21Upload folder using huggingface_hub0e10c9f5.1 KB
    Loading...

Discussions 1 thread

  1. 2026-09-04MTP head needs to be added to config ignore listclosed2 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration