← back to catalog · registered 2026-08-22 13:56

w341e/Qwen3.5-122B-A10B-abliterated-NVFP4

w341e Qwen 59B MoE second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/w341e%2FQwen3.5-122B-A10B-abliterated-NVFP4"
Response includes
  • classification m1
  • files 12
  • hub_downloads_all_time 66,996
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
67K
102 last 30d - cooling
Likes
1
Model age
6mo ago
created 2026-04-13
Downloads over time
Now67K→from9.3K↑618%
6.5K28.6K50.7K72.8K9.3K on Apr 1567K on Oct 1167K on Oct 10AprMayJunJulAugSepOct
Apr 15 → Oct 11 · 65 snapshots · spans 179 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en ko zh ja
Tags
transformers safetensors qwen3_5_moe image-text-to-text qwen3.5 moe nvfp4 quantized abliterated compressed-tensors vllm text-generation

Related

Total size
71.2 GB
Files
12
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-04-13 09:16

Files by quantization

Auxiliary files 12 files 71.2 GB
model-00001-of-00002.safetensors 46.6 GB 0009eef4 download
model-00002-of-00002.safetensors 24.6 GB fff0528a download
tokenizer.json 19.1 MB dd6b8cf7 download
model.safetensors.index.json 15.8 MB eda5b07b download
vocab.json 6.41 MB 0aa0ce06 download
config.json 25.5 KB 2b83433a download
chat_template.jinja 7.57 KB a585dec8 download
README.md 5.01 KB 9a0e131c download
.gitattributes 1.60 KB aa7aacd0 download
tokenizer_config.json 1.07 KB 6be6ce17 download
preprocessor_config.json 390 B 2ea84a43 download
generation_config.json 213 B 1068c09f download

README current version from Hugging Face


license: apache-2.0
license_name: apache-2.0
license_link: https://hf.co/Qwen/Qwen3.5-122B-A10B/blob/main/LICENSE
base_model:

  • wangzhang/Qwen3.5-122B-A10B-abliterated
    tags:
  • qwen3.5
  • moe
  • nvfp4
  • quantized
  • abliterated
  • compressed-tensors
  • vllm
    language:
  • en
  • ko
  • zh
  • ja
    library_name: transformers
    pipeline_tag: text-generation

Qwen3.5-122B-A10B-abliterated-NVFP4

NVFP4 (4-bit floating point) quantized derivative of wangzhang/Qwen3.5-122B-A10B-abliterated, which itself is derived from Qwen/Qwen3.5-122B-A10B.

This repository provides a modified derivative checkpoint for local inference and serving. The primary changes in this repository are NVFP4 quantization, weight repacking / export formatting, and serving compatibility adjustments.

Model Details

Property Value
Intermediate Base Model wangzhang/Qwen3.5-122B-A10B-abliterated
Original Base Model Qwen/Qwen3.5-122B-A10B
Architecture Qwen3.5 MoE (256 routed experts, 10B active)
Quantization NVFP4 (compressed-tensors, nvfp4-pack-quantized)
Original Size 228 GB (BF16)
Quantized Size 71.2 GB (69% reduction)
Format safetensors (2 shards)

Quantization Method

This model was quantized using a template-based weight replacement approach:

  1. Reference Template: RedHatAI/Qwen3.5-122B-A10B-NVFP4 — a calibrated NVFP4 checkpoint of the original (non-abliterated) Qwen3.5-122B-A10B, produced by llm-compressor with proper calibration data.
  2. Weight Replacement: Each quantized tensor (weight_packed and weight_scale) was regenerated from the abliterated BF16 weights using the reference checkpoint's weight_global_scale and input_global_scale values.
  3. Format Preservation: The reference checkpoint's config.json, quantization_config, global scales, and all metadata were preserved unchanged, ensuring full compatibility with vLLM's CUTLASS NVFP4 MoE kernel.

What is Quantized

Component Format Notes
Routed experts (gate/up/down_proj) NVFP4 256 experts × 48 layers × 3 projections
Shared experts NVFP4 48 layers × 3 projections
Self-attention (q/k/v/o_proj) NVFP4 12 full-attention layers
Linear attention BF16 36 layers, kept at full precision
Embeddings, norms, gates BF16 Kept at full precision

Serving with vLLM

This model requires a text-only compatibility patch for vLLM since Qwen3.5 MoE is a multimodal architecture but this checkpoint contains only text weights.

Quick Start

# 1. Download the model
huggingface-cli download bjk110/Qwen3.5-122B-A10B-abliterated-NVFP4

# 2. Apply the text-only patch before starting vLLM
python vllm_patches/patch_qwen35_moe_text.py

# 3. Serve with vLLM
vllm serve /path/to/model \
    --served-model-name Qwen3.5-122B-A10B-abliterated-NVFP4 \
    --max-model-len 131072 \
    --max-num-seqs 4 \
    --gpu-memory-utilization 0.90 \
    --trust-remote-code \
    --enable-prefix-caching \
    --enable-chunked-prefill \
    --reasoning-parser qwen3

### Docker Compose (Recommended)

A complete Docker Compose setup is provided in the `serving/` directory:

```bash
# Copy serving files
cp -r serving/ /path/to/your/vllm-setup/

# Edit .env to set MODEL_PATH
vim serving/.env

# Start
cd serving && docker compose --profile head up -d

See serving/ directory for:

  • docker-compose.yml — Full vLLM serving configuration
  • .env.example — Environment variables template
  • entrypoint.sh — Entrypoint with automatic patch application

Hardware Requirements

Configuration Memory max_model_len Notes
1× NVIDIA DGX Spark (GB10) 121 GiB unified 131,072 (128K) Tested and verified
1× GPU with 80+ GB VRAM 80 GiB ~65,536 Estimated

Performance (DGX Spark, TP=1)

Metric Value
Throughput 14.5 tok/s average, 16.8 tok/s peak
KV Cache 222K tokens (20.4 GiB)
Max Concurrency 6.16× at 128K context
Model Loading ~13 min (2 shards)

Referenced Models

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-04-13Duplicate from bjk110/Qwen3.5-122B-A10B-abliterated-NVFP4ddb9bef5 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration