For privacy reasons a browser tells us at most
"≥ 8 GB RAM, 8 cores" - same reading whether you have 8 GB
or 128 GB. It has no idea how much RAM is free right now, which
apps are open, or whether you have a GPU.
The Abliteration app is integrated with your machine
It reads your exact RAM, GPU model and VRAM,
free memory right now, and picks the sharpest quant that
still fits. Every model page lights up precisely for your rig.
And you can chat with any model, right now
The app is a full local runtime - no API keys, no subscription,
everything runs on your machine. Click any model on this site and
start a conversation in seconds.
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
'abliterated' in name/tags
is_gguf=0 (base model)
no specific method indicators - defaulting to M1 (most common)
Method: Refusal direction projection removal (Arditi et al., 2024)
Layers ablated: 5 (layers 43-47, covering both self_attn and linear_attn/Mamba layers)
Tensors modified: 10 (o_proj/out_proj + q_proj/in_proj_qkv per layer)
Alpha: 1.0 (full removal)
Measurement: 64 harmful + 64 harmless prompts, strongest refusal signal at layer 48 (score: 78.4)
Architecture
Type: Mixture of Experts (MoE) + Mamba hybrid attention
Total params: 122B
Active params: 10B per token (8/256 experts routed + 1 shared)
Context: 262K tokens native
Layers: 48 (13 self_attn + 36 linear_attn/Mamba)
Usage
Compatible with vLLM, transformers, and other inference frameworks that support Qwen3.5 MoE architecture.
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained(
"Chompa1422/Qwen3.5-122B-A10B-abliterated",
device_map="auto",
trust_remote_code=True,
dtype="bfloat16",
)
tokenizer = AutoTokenizer.from_pretrained("Chompa1422/Qwen3.5-122B-A10B-abliterated")
Disclaimer
This model is intended for authorized security testing, CTF competitions, and educational purposes only.
README history
1 version
The author's README evolved over time. Click a version to see its content at that point.
2026-06-04Duplicate from Chompa1422/Qwen3.5-122B-A10B-abliterated8e271fa1.5 KB
Loading...
Catalog is the map. Apps are the tools.
Run models on your own machine, not in the cloud.
Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.