← back to catalog · registered 2026-08-22 13:56

lhca521/Qwen3.5-122B-A10B-abliterated-AWQ

lhca521 Qwen 117B MoE second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/lhca521%2FQwen3.5-122B-A10B-abliterated-AWQ"
Response includes
  • classification m1
  • files 12
  • hub_downloads_all_time 2,989
  • author_summary 23 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
3K
65 last 30d - cooling
Likes
1
Model age
6mo ago
created 2026-04-12
Downloads over time
Now3K→from42↑7,112%
01.1K2.2K3.3K42 on Apr 153K on Oct 11AprMayJunJulAugSepOct
Apr 15 → Oct 11 · 65 snapshots · spans 179 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
other
Languages
en ko zh ja
Tags
transformers safetensors qwen3_5_moe image-text-to-text qwen3.5 moe awq int4 quantized abliterated compressed-tensors text-generation

Related

Total size
65.8 GB
Files
12
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-04-12 12:04

Files by quantization

Auxiliary files 12 files 65.9 GB
model-00001-of-00002.safetensors 46.6 GB 42238226 download
model-00002-of-00002.safetensors 19.3 GB 72699962 download
tokenizer.json 19.1 MB 639e352c download
model.safetensors.index.json 11.9 MB acd5ee88 download
config.json 23.4 KB 8bd9fd28 download
chat_template.jinja 7.57 KB a585dec8 download
README.md 4.34 KB 47046a33 download
.gitattributes 1.60 KB aa7aacd0 download
recipe.yaml 1.35 KB ae2d3cbb download
processor_config.json 1.33 KB a999a522 download
tokenizer_config.json 1.17 KB c748bf6c download
generation_config.json 218 B 333a71b6 download

README current version from Hugging Face


license: other
license_name: tongyi-qianwen
license_link: https://hf.co/Qwen/Qwen3.5-122B-A10B/blob/main/LICENSE
base_model:

  • wangzhang/Qwen3.5-122B-A10B-abliterated
    tags:
  • qwen3.5
  • moe
  • awq
  • int4
  • quantized
  • abliterated
  • compressed-tensors
    language:
  • en
  • ko
  • zh
  • ja
    library_name: transformers
    pipeline_tag: text-generation

Qwen3.5-122B-A10B-abliterated-AWQ

AWQ INT4 (W4A16) quantized version of wangzhang/Qwen3.5-122B-A10B-abliterated, a Mixture-of-Experts model with 122B total parameters and 10B active parameters per token.

Model Details

Property Value
Base Model wangzhang/Qwen3.5-122B-A10B-abliterated
Architecture Qwen3.5 MoE (256 routed experts, 10B active)
Quantization AWQ INT4 (W4A16, symmetric, group_size=128)
Quantization Tool llm-compressor 0.10.1.dev (main branch)
Quantization Format compressed-tensors, pack-quantized
Original Size 228 GB (BF16)
Quantized Size 66 GB (71% reduction)
Format safetensors (2 shards)
Calibration WikiText-103, 8 samples, seq_len=256

Quantization Details

What is Quantized

Component Format Notes
Routed experts (gate/up/down_proj) INT4 packed 256 experts x 48 layers x 3 projections = 36,864 quantized tensors
Self-attention (q/k/v/o_proj) INT4 packed 12 full-attention layers
Shared experts BF16 Kept at full precision for quality
Linear attention BF16 Kept at full precision (36 layers)
Embeddings, norms, gates BF16 Kept at full precision

Quantization Method

This model was quantized using llm-compressor (main branch, commit e48353f8) with AWQModifier:

  1. Fused Expert Unfusing: llm-compressor's CalibrationQwen3_5MoeSparseMoeBlock unfuses the 3D fused expert parameters (Qwen3_5MoeExperts) into individual nn.Linear modules, enabling standard AWQ quantization
  2. AWQ Smoothing: Activation-aware weight quantization with grid search (n_grid=10) for optimal scale factors
  3. INT4 Packing: Weights packed into int32 tensors (8 INT4 values per int32) with per-group scales (group_size=128)

Compatibility

Serving Requirements

Important: This model uses compressed-tensors format with WNA16 (Weight N-bit Activation 16-bit) quantization. The required inference kernels have specific GPU architecture requirements.

GPU Architecture Compute Capability Compatible? Notes
NVIDIA Hopper (H100, H200) SM90 Yes CutlassW4A8 + MacheteLinearKernel
NVIDIA Ada (L40S, RTX 4090) SM89 Yes Marlin kernel
NVIDIA Blackwell (B200) SM100 Yes Full support
NVIDIA DGX Spark (GB10) SM121 No WNA16 kernels require SM90+
NVIDIA Ampere (A100) SM80 Untested May work with Marlin fallback

vLLM Serving Example

vllm serve bjk110/Qwen3.5-122B-A10B-abliterated-AWQ \
    --served-model-name Qwen3.5-122B-A10B-abliterated-AWQ \
    --quantization compressed-tensors \
    --max-model-len 32768 \
    --trust-remote-code \
    --enable-chunked-prefill \
    --reasoning-parser qwen3

Note: This model requires the TextOnlyShim patch for vLLM since the base architecture is multimodal (Qwen3.5 MoE) but this checkpoint contains only text weights. The patch is included in the vllm_patches/ directory.

Referenced Models

License

This model inherits the license from the base model: Tongyi Qianwen License.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-04-12Duplicate from bjk110/Qwen3.5-122B-A10B-abliterated-AWQaed31fb4.3 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration