← back to catalog · registered 2026-08-22 13:56

SC117/Qwen3.5-122B-A10B-Uncensored-APEX-Compact-GGUF

SC117 Qwen 122B GGUF MoE second-order 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/SC117%2FQwen3.5-122B-A10B-Uncensored-APEX-Compact-GGUF"
Response includes
  • classification m-uncensored
  • files 5
  • hub_downloads_all_time 29,932
  • author_summary 18 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
30K
1K last 30d - cooling
Likes
11
Model age
4mo ago
created 2026-06-07
Downloads over time
Now30.5K→from17.8K↑71%
17.2K22.1K26.9K31.8K17.8K on Jun 2430.5K on Oct 11JunJulAugSepOct
Jun 24 → Oct 11 · 56 snapshots · spans 109 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh multilingual
Tags
gguf quantized apex moe mixture-of-experts qwen3.5 uncensored en zh multilingual base_model:HauhauCS/Qwen3.5-122B-A10B-Uncensored-HauhauCS-Aggressive base_model:quantized:HauhauCS/Qwen3.5-122B-A10B-Uncensored-HauhauCS-Aggressive

Related

Total size
55.1 GB
Files
5
Quantizations
2
Registered
2026-08-22 13:56
Last updated on HF
2026-10-03 10:22

Files by quantization

F16 1 file 867 MB
mmproj-Qwen3.5-122B-A10B-Uncensored-HauhauCS-Aggressive-f16.gguf 867 MB 9d3c48d5 download
Auxiliary files 4 files 55.1 GB
Qwen3.5-122B-A10B-Uncensored-APEX-Compact.gguf 55.1 GB f7036d3e download
README_zh.md 23.1 KB a0c12120 download
README.md 23.0 KB bc90de89 download
.gitattributes 1.75 KB ebd13aa6 download

README current version from Hugging Face


license: apache-2.0
base_model: HauhauCS/Qwen3.5-122B-A10B-Uncensored-HauhauCS-Aggressive
tags:

  • gguf
  • quantized
  • apex
  • moe
  • mixture-of-experts
  • qwen3.5
  • uncensored
    language:
  • en
  • zh
  • multilingual

⚡ Qwen3.5-122B-A10B Uncensored — APEX I-Compact GGUF

English | 📖 中文文档

MoE Mixed-Precision Quantization · Uncensored · 55.1 GB

APEX I-Compact MoE 122B Uncensored 55.1 GB Multimodal

Qwen3.5-122B-A10B (uncensored by HauhauCS) quantized with APEX I-Compact — a MoE-aware mixed-precision strategy that applies layer-wise precision gradients. Edge layers get higher precision, middle layers get aggressive compression. Quantized from Q8_K_P using the APEX project.

📊 Benchmark Results

Measurements from APEX project on 8×RTX PRO 6000 Blackwell (768 GB VRAM). Perplexity on wikitext-2-raw (ctx 512). Accuracy via llama.cpp (400 tasks each).

Profile Size PPL HellaSwag Wino MMLU ARC t/s
Q8_0 (ref)121 GB4.81985.5%77.3%44.1957.1985.5
APEX I-Balanced83.4 GB4.83185.5%77.8%43.8657.8696.7
APEX I-Compact ★55.1 GB4.97884.5%77.5%44.0657.86106.3
APEX I-Mini44.9 GB5.30684.0%75.3%42.8356.52110.0

★ This quantization. I-Compact achieves 84.5% HellaSwag and 57.86 ARC at 55% less size than Q8_0, fastest standard APEX profile at 106 t/s. Quantized from Q8_K_P (137 GB → 55.1 GB).

🔬 Quantization Strategy

APEX I-Compact applies layer-wise mixed-precision with MoE-aware tensor classification. Edge layers (L0–4, L43–47) get higher precision, middle layers (L10–29) get more aggressive compression. I-variants use diverse imatrix calibration (chat, code, reasoning, tool-calling, agentic traces) for better real-world accuracy.

Component Edge (L0-4, L43-47) Middle (L10-29) Role
Routed Experts (exps)Q3_K_MQ3_K_S256 experts, 8 active
Shared Experts (shexp)Q4_K_SQ4_K_SAlways active
Attention (QKV)Q4_K_SQ3_K_MPer-layer attention
Router (gate_inp)F32F32Precision-critical

Router weights kept in F32 (lossless) to preserve routing accuracy. Shared experts at Q4_K_S across all layers for stable token processing.

🏗️ Architecture
Base ModelHauhauCS/Qwen3.5-122B-A10B-Uncensored-HauhauCS-Aggressive
Parameters122B total, ~10B active per token
Experts256 routed + 1 shared (8 active per token)
ArchitectureHybrid: Gated DeltaNet linear attention + softmax attention (3:1)
Layers48 (12 × (3 DeltaNet-MoE + 1 Attention-MoE))
Context262K native
ModalitiesText, Image, Video (natively multimodal)
Vocabulary248K tokens, 201 languages
UncensoredAggressive variant — 0/465 refusals
Quantized FromQ8_K_P (137 GB → 55.1 GB, 3.88 BPW effective)
⚙️ Recommended Settings

Thinking mode (default):

Generaltemp=1.0, top_p=0.95, top_k=20, min_p=0, presence_penalty=1.5
Codingtemp=0.6, top_p=0.95, top_k=20, min_p=0, presence_penalty=0

Non-thinking mode:

Generaltemp=0.7, top_p=0.8, top_k=20, min_p=0, presence_penalty=1.5
Reasoningtemp=1.0, top_p=1.0, top_k=40, min_p=0, presence_penalty=2.0

Use --jinja flag with llama.cpp. Thinking mode is on by default — disable with --chat-template-kwargs '{"enable_thinking":false}'. Vision support requires mmproj file.

📝 Usage

Works with llama.cpp, LM Studio, Jan, koboldcpp, and other GGUF-compatible runtimes.

# Text only
llama-cli -m Qwen3.5-122B-A10B-Uncensored-APEX-Compact.gguf \
  --jinja -c 131072 -ngl 99

# With vision
llama-cli -m Qwen3.5-122B-A10B-Uncensored-APEX-Compact.gguf
--mmproj mmproj-Qwen3.5-122B-A10B-Uncensored-f16.gguf
--jinja -c 131072 -ngl 99



For CPU inference: this 55 GB model fits in 64 GB+ RAM systems. For GPU offload, adjust -ngl based on available VRAM.




README history 4 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-10-03Add optional Ko-fi support banner33b025e23.2 KB
    Loading...
  2. 2026-06-26Update documentation1c2952022.9 KB
    Loading...
  3. 2026-06-07Upload README.md with huggingface_hub40dad0722.9 KB
    Loading...
  4. 2026-06-07Upload README.md with huggingface_hub928a61f22.6 KB
    Loading...

Discussions 1 thread

  1. 2026-06-29does it have MTP head included?open2 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration