← back to catalog · registered 2026-08-22 13:56

wangzhang/Qwen3.5-122B-A10B-abliterated-GGUF

wangzhang Qwen 122B GGUF MoE second-order 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/wangzhang%2FQwen3.5-122B-A10B-abliterated-GGUF"
Response includes
  • classification m8
  • files 4
  • hub_downloads_all_time 1,409
  • author_summary 28 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
1K
734 last 30d - active
Likes
8
Model age
7mo ago
created 2026-03-14
Downloads over time
Now1.4K→from68↑2,029%
05291.1K1.6K68 on Mar 181.4K on Oct 111.4K on Oct 9MarAprMayJunJulAugSepOct
Mar 18 → Oct 11 · 69 snapshots · spans 207 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Quantizations
Q4_K Q8_0
Tags
llama.cpp gguf abliterix uncensored decensored abliterated moe text-generation base_model:wangzhang/Qwen3.5-122B-A10B-abliterix base_model:quantized:wangzhang/Qwen3.5-122B-A10B-abliterix license:apache-2.0 endpoints_compatible

Related

Total size
190 GB
Files
4
Quantizations
3
Registered
2026-08-22 13:56
Last updated on HF
2026-08-29 17:04

Files by quantization

Q8_0 1 file 121 GB
Qwen3.5-122B-A10B-abliterated-v2-Q8_0.gguf 121 GB ******** download
Q4_K 1 file 69.1 GB
Qwen3.5-122B-A10B-abliterated-v2-Q4_K_M.gguf 69.1 GB ******** download
Auxiliary files 2 files 4.78 KB
README.md 3.14 KB 25870ca9 download
.gitattributes 1.64 KB 69c744fd download

README current version from Hugging Face


library_name: llama.cpp
license: apache-2.0
license_link: https://huggingface.co/Qwen/Qwen3.5-122B-A10B/blob/main/LICENSE
pipeline_tag: text-generation
base_model:

  • wangzhang/Qwen3.5-122B-A10B-abliterated
    tags:
  • abliterix
  • uncensored
  • decensored
  • abliterated
  • moe
  • gguf

Qwen3.5-122B-A10B-abliterated-GGUF

GGUF quantized versions of wangzhang/Qwen3.5-122B-A10B-abliterated, a decensored Qwen/Qwen3.5-122B-A10B created using Abliterix.

Available Quantizations

File Quantization Size Quality Use Case
Q4_K_M Q4_K_M 70 GB High 1x 80GB GPU or CPU with 96GB+ RAM
Q8_0 Q8_0 121 GB Very High 2x 80GB GPUs or CPU with 160GB+ RAM

About the Source Model

  • 95% refusal reduction: Reduced from 100/100 refusals to just 5/100 (5%)
  • MoE architecture: Qwen3.5-122B-A10B activates only ~10B parameters per token — 14B-class speed with 122B-class knowledge
  • Minimal capability loss: KL divergence of just 0.0878 from the original model
  • 50-trial Optuna TPE optimization: Automated Bayesian hyperparameter search with multi-objective Pareto optimization
  • Orthogonalized abliteration: Surgical removal of refusal directions without degrading general intelligence

See wangzhang/Qwen3.5-122B-A10B-abliterated for full details on the abliteration process and steering parameters.

Usage

llama.cpp

# Q4_K_M (recommended for single GPU)
./llama-cli -m Qwen3.5-122B-A10B-abliterated-Q4_K_M.gguf -p "Your prompt here" -n 512

# Q8_0 (higher quality)
./llama-cli -m Qwen3.5-122B-A10B-abliterated-Q8_0.gguf -p "Your prompt here" -n 512

llama-server

./llama-server -m Qwen3.5-122B-A10B-abliterated-Q4_K_M.gguf --host 0.0.0.0 --port 8080

Ollama

# Create a Modelfile
echo "FROM ./Qwen3.5-122B-A10B-abliterated-Q4_K_M.gguf" > Modelfile
ollama create qwen3.5-122b-abliterated -f Modelfile
ollama run qwen3.5-122b-abliterated

VRAM / RAM Requirements

Quantization Full GPU Offload Partial Offload (32 layers) CPU Only
Q4_K_M ~74 GB (1x 80GB) ~40 GB GPU + 40 GB RAM ~80 GB RAM
Q8_0 ~130 GB (2x 80GB) ~70 GB GPU + 70 GB RAM ~140 GB RAM

Disclaimer

This model is provided for research purposes only. The creator is not responsible for any misuse.

Credits

Discussions 3 threads

  1. 2026-04-05Is this normal? Did I do something wrong?open1 💬#3
    Loading...
  2. 2026-03-17Prometheus open source?closed6 💬#2
    Loading...
  3. 2026-03-14老大,成功了没有啊!closed3 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration