license: apache-2.0
language:
- en
- zh
base_model: - orcarouter/Qwen3.8-27B-Uncensored
pipeline_tag: image-text-to-text
Qwen3.8-27B-Uncensored · INT8 W8A16 · BF16 MTP
A high-fidelity, Ampere-optimized quantization of the abliterated Qwen3.8-27B for dual RTX 3090 / single RTX 4090 inference.
Base model · llm-compressor · vLLM
This is a numerical W8A16 quantization of orcarouter/Qwen3.8-27B-Uncensored, an abliterated (refusal-removed) version of Qwen3.8-27B. All model credit belongs to Qwen and OrcaRouter; refer to the upstream model cards for architecture, capabilities, and usage guidance.
Quantization Design
| Component | Precision | Reason |
|---|---|---|
| MLP projections | INT8 W8A16 | Largest dense GEMMs |
| Full-attention projections | INT8 W8A16 | Low output-distribution error |
GDN in_proj_qkv, in_proj_z, out_proj |
INT8 W8A16 | ~4 GB memory recovery |
GDN in_proj_a, in_proj_b |
BF16 | Tiny recurrent gates; precision safeguard |
| Vision tower | BF16 | Preserve multimodal fidelity |
lm_head |
BF16 | Preserve final-logit fidelity |
| MTP head | BF16 | Keep speculative drafter close to target |
Norms, conv1d, A_log, dt_bias |
BF16/FP32 | Non-Linear; never packed |
400 quantized Linear GEMMs: 192 MLP + 64 full-attention + 144 GDN projections.
Recipe
# recipe.yaml
default_stage:
default_modifiers:
QuantizationModifier:
targets: [Linear]
ignore:
- lm_head
- re:.*visual.*
- re:.*mtp.*
- re:.*linear_attn[.]in_proj_a$
- re:.*linear_attn[.]in_proj_b$
scheme: W8A16
Why W8A16 on Ampere
RTX 3090/4090 GPUs are Ampere/Ada sm_86. They do not provide native FP8 tensor-core execution. W8A16 uses INT8 weights (Marlin kernel) with BF16 activations — the correct format for Ampere/Ada inference with predictable behavior and near-lossless quality.
Do not use FP8 checkpoints on RTX 3090 — they silently fall back to BF16 dispatch, negating memory savings.
Memory Requirements
| Configuration | BF16 Source | This Quant |
|---|---|---|
| Weights (disk) | 55.6 GB | 29.4 GB |
| Loaded VRAM (1 GPU) | 56+ GB | ~29 GB |
| Minimum GPU | × H100 80 GB | 1× RTX 4090 / A6000 |
| Dual GPU | 2× 48 GB | 2× RTX 3090 24 GB |
262K context fits on a single RTX 4090 with FP8 KV cache.
Serving
Single GPU (RTX 4090 / A6000 48 GB)
vllm serve morikomorizz/Qwen3.8-27B-Uncensored-INT8-W8A16-MTP \
--served-model-name qwen3.8-27b-uncensored-w8a16 \
--dtype bfloat16 \
--gpu-memory-utilization 0.92 \
--max-model-len 262144 \
--trust-remote-code \
--reasoning-parser qwen3 \
--tool-call-parser qwen3_coder \
--speculative-config '{"method":"mtp","num_speculative_tokens":3}'
Dual RTX 3090 (2×24 GB, tensor-parallel)
export NCCL_P2P_DISABLE=1
export VLLM_WORKER_MULTIPROC_METHOD=spawn
vllm serve morikomorizz/Qwen3.8-27B-Uncensored-INT8-W8A16-MTP \
--tensor-parallel-size 2 \
--dtype bfloat16 \
--gpu-memory-utilization 0.92 \
--max-model-len 262144 \
--max-num-batched-tokens 8192 \
--kv-cache-dtype fp8_e4m3 \
--enable-chunked-prefill \
--trust-remote-code \
--reasoning-parser qwen3 \
--tool-call-parser qwen3_coder \
--speculative-config '{"method":"mtp","num_speculative_tokens":3}'
Python Client
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="none")
resp = client.chat.completions.create(
model="qwen3.8-27b-uncensored-w8a16",
messages=[{"role": "user", "content": "Hello!"}],
max_tokens=512,
)
print(resp.choices[0].message.content)
Checkpoint Profile
| Property | Value |
|---|---|
| Quantization | Data-free symmetric RTN W8A16, group size 128 |
| Runtime format | compressed-tensors / pack-quantized |
| Kernel dispatch | CompressedTensorsWNA16 → MarlinLinearKernel |
| Preserved precision | BF16 vision tower, lm_head, MTP, GDN gates |
| MTP | BF16 draft model; embeddings and lm_head shared with target |
| Runtime | vLLM; this is not a GGUF checkpoint |
Files
| File | Purpose |
|---|---|
model-00001-of-00002.safetensors |
Packed W8A16 language + GDN weights |
model-00002-of-00002.safetensors |
Packed W8A16 + BF16 vision tower |
model_mtp.safetensors |
BF16 MTP head, 15 tensors, ~0.79 GB |
model.safetensors.index.json |
Shard-to-tensor mapping |
recipe.yaml |
Exact llm-compressor W8A16 recipe |
⚠️ Disclaimer
This model is quantized from an abliterated (refusal-removed) source. It has had its safety alignment substantially removed and will comply with harmful, unethical, or offensive requests. It is released strictly for legitimate research. You assume full responsibility for its use.
Acknowledgements
- Qwen — Qwen/Qwen3.8-27B
- OrcaRouter — orcarouter/Qwen3.8-27B-Uncensored
- lued — Qwen3.8-27B-INT8-W8A16-MTP for the quantization recipe reference
- llm-compressor — vllm-project/llm-compressor
License: Apache 2.0, inherited from the base model.