← back to catalog · registered 2026-08-23 17:02

berkerdooo/Qwen3.8-27B-Uncensored-INT8-AutoRound

berkerdooo Qwen 8.6B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/berkerdooo%2FQwen3.8-27B-Uncensored-INT8-AutoRound"
Response includes
  • classification m-uncensored
  • files 24
  • hub_downloads_all_time 1,310
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
1K
565 last 30d - stable
Likes
0
Model age
6w ago
created 2026-08-23
Downloads over time
Now1.6K→from176↑781%
1076341.2K1.7K176 on Aug 261.6K on Oct 11AugSepOct
Aug 26 → Oct 11 · 47 snapshots · spans 46 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
transformers safetensors qwen3_5 image-text-to-text int8 autoround w8a16 conversational base_model:orcarouter/Qwen3.8-27B-Uncensored base_model:quantized:orcarouter/Qwen3.8-27B-Uncensored license:apache-2.0 endpoints_compatible

Related

Total size
34.3 GB
Files
24
Quantizations
1
Registered
2026-08-23 17:02
Last updated on HF
2026-08-23 15:47

Files by quantization

Auxiliary files 24 files 34.3 GB
model-00005-of-00012.safetensors 2.97 GB c22b26c4 download
model-00008-of-00012.safetensors 2.97 GB ded270ed download
model-00004-of-00012.safetensors 2.95 GB 7a8ea412 download
model-00007-of-00012.safetensors 2.95 GB 1d2fcec5 download
model-00001-of-00012.safetensors 2.92 GB e86246ba download
model-00003-of-00012.safetensors 2.92 GB 713724b0 download
model-00006-of-00012.safetensors 2.92 GB e96732ef download
model-00009-of-00012.safetensors 2.92 GB ae94e5a7 download
model-00002-of-00012.safetensors 2.91 GB 44c0814a download
model-00010-of-00012.safetensors 2.71 GB 18fb7416 download
model-00011-of-00012.safetensors 2.37 GB da51fa48 download
model-00012-of-00012.safetensors 2.37 GB 6866cf8a download
model_extra_tensors.safetensors 463 MB 39744749 download
tokenizer.json 19.1 MB 06b95093 download
model.safetensors.index.json 159 KB 3d7093b5 download
config.json 32.3 KB b62cf4bf download
quantization_config.json 26.6 KB f244a6c3 download
chat_template.jinja 8.74 KB c0c686f9 download
README.md 3.03 KB e6733e68 download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.16 KB 33818c7f download
tokenizer_config.json 1.14 KB 875b7b0a download
preprocessor_config.json 443 B 8ed39680 download
generation_config.json 214 B 3f9de11a download

README current version from Hugging Face


library_name: transformers
license: apache-2.0
base_model: orcarouter/Qwen3.8-27B-Uncensored
base_model_relation: quantized
pipeline_tag: image-text-to-text
tags:

  • qwen3_5
  • int8
  • autoround
  • w8a16

Qwen3.8-27B-Uncensored INT8 AutoRound (W8A16, linear attention BF16, group size 128)

INT8 weight-only quantization of orcarouter/Qwen3.8-27B-Uncensored
with AutoRound (SignRound), following the recipe of
Minachist/Qwen3.8-27B-INT8-AutoRound branch linear-attn-bf16-gs128,
with two changes: linear attention is excluded from tuning (not swapped back to BF16 after the fact), and 500 iters instead of 250.

Tensors Precision
self_attn.{q,k,v,o}_proj (16 full-attention layers), mlp.{gate,up,down}_proj (64 layers), MTP block projections INT8 symmetric, group_size 128
linear_attn.{in_proj_qkv,in_proj_z,out_proj,in_proj_a,in_proj_b} (48 GDN layers), embed_tokens, lm_head, mtp.fc, norms, vision tower BF16

263 INT8 linears / 354 BF16 linears. Format: auto_round:auto_gptq packing (vLLM loads it via GPTQ-Marlin with BF16 activations).

Recipe

AutoRound main @ b9f3d0079d014c73a1ff009800c597b9bc3f2a36 (version string 0.15.0), transformers 5.15.1, torch 2.13.0+cu130, one RTX PRO 6000 Blackwell.
scheme="W8A16" (bits 8, group_size 128, sym), iters=500, nsamples=1024, seqlen=2048, batch_size=4, gradient_accumulate_steps=2, low_gpu_mem_usage=False, seed=42.
Calibration: 256 samples built from NeelNanda/pile-10k + 768 from codeparrot/github-code-clean (documents concatenated so every sample is >= 2048 tokens, then truncated to 2048).
Every layer is named in full in layer_config (avoids AutoRound's shared-dict regex aliasing bug). Tuning took 1.26 h.

KL divergence vs the BF16 source

Teacher-forced top-24 logprobs on one 128,000-token wikitext-103 stream (rows 100k+ of the train split), one sequence, BF16 KV cache, vLLM 0.27.1, KL(P_bf16 || Q_int8) in nats over the truncated top-24.
These numbers are only comparable to other models scored with the same script, stream and teacher.

depth n KL mean KL p50 KL p99 top-1 agreement ΔNLL
0k-4k 3,999 0.00189 0.00056 0.0239 97.67% +0.0038
4k-16k 12,000 0.00363 0.00078 0.0338 97.51% +0.0011
16k-48k 32,000 0.00264 0.00085 0.0293 97.22% +0.0019
48k-128k 80,000 0.00320 0.00088 0.0338 97.28% +0.0022

Own NLL: BF16 1.8244, INT8 1.8265. For reference, the same script on Qwen/Qwen3.8-27B gives FP8 (Qwen/Qwen3.8-27B-FP8) KL 0.0048 / top-1 96.5% and Minachist's INT8 0.0029 / 97.2%.

Serving

vllm serve <this-repo> --tensor-parallel-size 2 --max-model-len 262144 --trust-remote-code

Tested with vLLM 0.27.1 (Using MarlinLinearKernel for AutoGPTQLinearMethod). MTP speculative decoding: --speculative-config '{"method":"mtp","num_speculative_tokens":3}'.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-23Upload folder using huggingface_hub71241053 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration