base_model:
- Avesed/Qwopus3.6-27B-v2-abliterated
pipeline_tag: text-generation
library_name: transformers
tags: - abliterated
- uncensored
- qwen3_5
- compressed-tensors
- int4
Qwopus3.6-27B-v2-abliterated-int4
int4 weight-only (W4A16) quantization of Avesed/Qwopus3.6-27B-v2-abliterated. Deployable with vLLM (compressed-tensors / Marlin), ~26 GB.
Recipe
- AWQ activation smoothing + int4 weight quantization,
group_size=32, symmetric, MSE observer - Quantized:
self_attn+mlpLinear layers - Kept higher precision (ignored):
linear_attn, vision tower, MTP head,embed_tokens,lm_head
Evaluation
| Benchmark | Score |
|---|---|
| HumanEval pass@1 | 95.1% |
| GSM8K | 86.0% |
| MMLU-Pro | 83.2% |
| Refusal rate | 8% |
MTP head
The Multi-Token-Prediction (mtp) head is included (for speculative decoding). Its residual-write matrices (self_attn.o_proj, mlp.down_proj) are abliterated with the same refusal direction as the main layers.