← back to catalog · registered 2026-08-22 13:56

morikomorizz/Qwen3.8-27B-Uncensored-INT8-W8A16-MTP

morikomorizz Qwen 24B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/morikomorizz%2FQwen3.8-27B-Uncensored-INT8-W8A16-MTP"
Response includes
  • classification m-uncensored
  • files 15
  • hub_downloads_all_time 2,303
  • author_summary 3 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
2K
2K last 30d - active
Likes
3
Model age
7w ago
created 2026-08-19
Downloads over time
Now2.7K→from129↑1,996%
09872K3K129 on Aug 192.7K on Oct 11AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh
Tags
safetensors qwen3_5 image-text-to-text conversational en zh base_model:orcarouter/Qwen3.8-27B-Uncensored base_model:quantized:orcarouter/Qwen3.8-27B-Uncensored license:apache-2.0 compressed-tensors region:us

Related

Total size
29.4 GB
Files
15
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-20 17:38

Files by quantization

Auxiliary files 15 files 29.5 GB
model-00001-of-00002.safetensors 18.6 GB 8bd0502a download
model-00002-of-00002.safetensors 10.0 GB ac96c26c download
model_mtp.safetensors 810 MB 7a9fd4ee download
tokenizer.json 19.1 MB 06b95093 download
model.safetensors.index.json 193 KB f63cd2c1 download
config.json 20.3 KB 73897926 download
chat_template.jinja 8.74 KB c0c686f9 download
README.md 5.78 KB 7e7c231e download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.16 KB 33818c7f download
tokenizer_config.json 1.14 KB b4acebe0 download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
recipe.yaml 379 B bb05b251 download
generation_config.json 214 B 0bc3addd download

README current version from Hugging Face


license: apache-2.0
language:

  • en
  • zh
    base_model:
  • orcarouter/Qwen3.8-27B-Uncensored
    pipeline_tag: image-text-to-text

Qwen3.8-27B-Uncensored · INT8 W8A16 · BF16 MTP

A high-fidelity, Ampere-optimized quantization of the abliterated Qwen3.8-27B for dual RTX 3090 / single RTX 4090 inference.

Base model · llm-compressor · vLLM

Format W8A16 Weights INT8 Activations BF16 License Apache 2.0

This is a numerical W8A16 quantization of orcarouter/Qwen3.8-27B-Uncensored, an abliterated (refusal-removed) version of Qwen3.8-27B. All model credit belongs to Qwen and OrcaRouter; refer to the upstream model cards for architecture, capabilities, and usage guidance.


Quantization Design

Component Precision Reason
MLP projections INT8 W8A16 Largest dense GEMMs
Full-attention projections INT8 W8A16 Low output-distribution error
GDN in_proj_qkv, in_proj_z, out_proj INT8 W8A16 ~4 GB memory recovery
GDN in_proj_a, in_proj_b BF16 Tiny recurrent gates; precision safeguard
Vision tower BF16 Preserve multimodal fidelity
lm_head BF16 Preserve final-logit fidelity
MTP head BF16 Keep speculative drafter close to target
Norms, conv1d, A_log, dt_bias BF16/FP32 Non-Linear; never packed

400 quantized Linear GEMMs: 192 MLP + 64 full-attention + 144 GDN projections.

Recipe

# recipe.yaml
default_stage:
  default_modifiers:
    QuantizationModifier:
      targets: [Linear]
      ignore:
        - lm_head
        - re:.*visual.*
        - re:.*mtp.*
        - re:.*linear_attn[.]in_proj_a$
        - re:.*linear_attn[.]in_proj_b$
      scheme: W8A16

Why W8A16 on Ampere

RTX 3090/4090 GPUs are Ampere/Ada sm_86. They do not provide native FP8 tensor-core execution. W8A16 uses INT8 weights (Marlin kernel) with BF16 activations — the correct format for Ampere/Ada inference with predictable behavior and near-lossless quality.

Do not use FP8 checkpoints on RTX 3090 — they silently fall back to BF16 dispatch, negating memory savings.


Memory Requirements

Configuration BF16 Source This Quant
Weights (disk) 55.6 GB 29.4 GB
Loaded VRAM (1 GPU) 56+ GB ~29 GB
Minimum GPU × H100 80 GB 1× RTX 4090 / A6000
Dual GPU 2× 48 GB 2× RTX 3090 24 GB

262K context fits on a single RTX 4090 with FP8 KV cache.


Serving

Single GPU (RTX 4090 / A6000 48 GB)

vllm serve morikomorizz/Qwen3.8-27B-Uncensored-INT8-W8A16-MTP \
  --served-model-name qwen3.8-27b-uncensored-w8a16 \
  --dtype bfloat16 \
  --gpu-memory-utilization 0.92 \
  --max-model-len 262144 \
  --trust-remote-code \
  --reasoning-parser qwen3 \
  --tool-call-parser qwen3_coder \
  --speculative-config '{"method":"mtp","num_speculative_tokens":3}'

Dual RTX 3090 (2×24 GB, tensor-parallel)

export NCCL_P2P_DISABLE=1
export VLLM_WORKER_MULTIPROC_METHOD=spawn

vllm serve morikomorizz/Qwen3.8-27B-Uncensored-INT8-W8A16-MTP \
  --tensor-parallel-size 2 \
  --dtype bfloat16 \
  --gpu-memory-utilization 0.92 \
  --max-model-len 262144 \
  --max-num-batched-tokens 8192 \
  --kv-cache-dtype fp8_e4m3 \
  --enable-chunked-prefill \
  --trust-remote-code \
  --reasoning-parser qwen3 \
  --tool-call-parser qwen3_coder \
  --speculative-config '{"method":"mtp","num_speculative_tokens":3}'

Python Client

from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="none")
resp = client.chat.completions.create(
    model="qwen3.8-27b-uncensored-w8a16",
    messages=[{"role": "user", "content": "Hello!"}],
    max_tokens=512,
)
print(resp.choices[0].message.content)

Checkpoint Profile

Property Value
Quantization Data-free symmetric RTN W8A16, group size 128
Runtime format compressed-tensors / pack-quantized
Kernel dispatch CompressedTensorsWNA16 → MarlinLinearKernel
Preserved precision BF16 vision tower, lm_head, MTP, GDN gates
MTP BF16 draft model; embeddings and lm_head shared with target
Runtime vLLM; this is not a GGUF checkpoint

Files

File Purpose
model-00001-of-00002.safetensors Packed W8A16 language + GDN weights
model-00002-of-00002.safetensors Packed W8A16 + BF16 vision tower
model_mtp.safetensors BF16 MTP head, 15 tensors, ~0.79 GB
model.safetensors.index.json Shard-to-tensor mapping
recipe.yaml Exact llm-compressor W8A16 recipe

⚠️ Disclaimer

This model is quantized from an abliterated (refusal-removed) source. It has had its safety alignment substantially removed and will comply with harmful, unethical, or offensive requests. It is released strictly for legitimate research. You assume full responsibility for its use.


Acknowledgements

License: Apache 2.0, inherited from the base model.

README history 4 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-19Update README.mdc53e8ec5.8 KB
    Loading...
  2. 2026-08-19Update README.md9cf43815.9 KB
    Loading...
  3. 2026-08-19Update README.md0d06a295.8 KB
    Loading...
  4. 2026-08-19Create README.md52fa90a95 B
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration