For privacy reasons a browser tells us at most
"≥ 8 GB RAM, 8 cores" - same reading whether you have 8 GB
or 128 GB. It has no idea how much RAM is free right now, which
apps are open, or whether you have a GPU.
The Abliteration app is integrated with your machine
It reads your exact RAM, GPU model and VRAM,
free memory right now, and picks the sharpest quant that
still fits. Every model page lights up precisely for your rig.
And you can chat with any model, right now
The app is a full local runtime - no API keys, no subscription,
everything runs on your machine. Click any model on this site and
start a conversation in seconds.
No classification signals present. This may not be an abliterated model at all - it could be a repackaging, a merge with unrelated goals, or unrelated content that mentions the term.
no classification signals present (no abliterated, uncensored, or known producer/method markers)
Refusal direction extraction
No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.
Exactly the same set of tensors is quantized as in the official checkpoint (routed + shared experts, dense MLPs, MLA projections); everything in modules_to_not_convert (embeddings, lm_head, routers, norms, linear-attention/KDA, hyper-connections, vision tower) is left in BF16/F32 as in the source.
Each FP8 weight has a weight_scale_inv (F32, one scale per 128×128 block, scale = amax / 448). Round-to-nearest, no calibration data.
Tensor names and the quantization_config are copied from the official release, so any engine that serves zai-org/GLM-5.3-Flash should load this checkpoint the same way.
Serving
Use the same setup as for zai-org/GLM-5.3-Flash, e.g.
From the upstream card: this abliterated variant has not undergone a separate evaluation. Results reported for apex-flash-1 apply to the standard checkpoint only. It is intended for authorized security research.
This quantization has not been separately benchmarked either.
Not affiliated with Cantina Security or Z.AI.
License
MIT, inherited from the base model. Copyright (c) 2026 Z.AI Co., Ltd — see LICENSE.
Catalog is the map. Apps are the tools.
Run models on your own machine, not in the cloud.
Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.