For privacy reasons a browser tells us at most
"≥ 8 GB RAM, 8 cores" - same reading whether you have 8 GB
or 128 GB. It has no idea how much RAM is free right now, which
apps are open, or whether you have a GPU.
The Abliteration app is integrated with your machine
It reads your exact RAM, GPU model and VRAM,
free memory right now, and picks the sharpest quant that
still fits. Every model page lights up precisely for your rig.
And you can chat with any model, right now
The app is a full local runtime - no API keys, no subscription,
everything runs on your machine. Click any model on this site and
start a conversation in seconds.
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
'abliterated' in name/tags
is_gguf=0 (base model)
no specific method indicators - defaulting to M1 (most common)
uncensored extra_gated_prompt: "Access is granted manually after review of your request."
GLM-5.2-abliterated
An uncensored build of GLM-5.2 (NVFP4) with the refusal direction removed, so the model follows instructions directly. General capabilities — reasoning, coding, multilingual output, and tool use — are preserved.
In GLM-5.2 the refusal direction is concentrated in the mid-to-late decoder layers (roughly layers 32–77, peaking around layer 34).
Usage
Loads as GlmMoeDsaForCausalLM (NVFP4) in vLLM or Transformers, identically to the base model. For best instruction-following, provide an explicit system prompt.
Access
Non Researchers are excluded. If you are applying with a email address public like gmail hotmail or there like's, be sure the request will be denied. Due to demand and people beeing constantly rude about not getting access I stopped reviewing access.
Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.