For privacy reasons a browser tells us at most
"≥ 8 GB RAM, 8 cores" - same reading whether you have 8 GB
or 128 GB. It has no idea how much RAM is free right now, which
apps are open, or whether you have a GPU.
The Abliteration app is integrated with your machine
It reads your exact RAM, GPU model and VRAM,
free memory right now, and picks the sharpest quant that
still fits. Every model page lights up precisely for your rig.
And you can chat with any model, right now
The app is a full local runtime - no API keys, no subscription,
everything runs on your machine. Click any model on this site and
start a conversation in seconds.
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
'abliterated' in name/tags
is_gguf=0 (base model)
no specific method indicators - defaulting to M1 (most common)
Only the routed MoE experts (mlp.experts.*.{gate,up,down}_proj) are quantized: E2M1 values packed two per byte (U8), FP8 E4M3 weight_scale per 16 input elements, FP32 per-tensor weight_scale_2 (= amax / (6 * 448)).
Weight-only: no input_scale, input_activations: null. Use a W4A16 MoE path (e.g. vLLM Marlin / b12x W4A16), not native W4A4 kernels.
Everything else (attention, KDA, shared experts, dense MLPs, routers, MTP, vision) is BF16, unchanged from the source.
Round-to-nearest, no calibration. Expert tensor relative error ~9% (typical for NVFP4); passthrough tensors are bit-identical to the source.
Notes
From the upstream card: this abliterated variant has not undergone a separate evaluation; it is intended for authorized security research.
This quantization has not been separately benchmarked. Not affiliated with Cantina Security or Z.AI.
License
MIT, inherited from the base model. Copyright (c) 2026 Z.AI Co., Ltd — see LICENSE.
Catalog is the map. Apps are the tools.
Run models on your own machine, not in the cloud.
Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.