← back to catalog · registered 2026-09-18 23:56

genevera/GLM-5.3-Flash-Uncensored-EXL3-2.5bpw

genevera Glm MoE multimodal
curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/genevera%2FGLM-5.3-Flash-Uncensored-EXL3-2.5bpw"
Response includes
  • classification m1
  • files 24
  • author_summary 3 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-09-18
Downloads over time
Now0from0↑0%
00110 on Sep 180 on Sep 19Sep
Sep 18 → Sep 19 · 2 snapshots · spans 1 day

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
mit
Languages
en
Tags
safetensors glm5_next abliterated uncensored crack refusal-removed glm moe exl3 exllamav3 2.51-bit vision

Related

Total size
98.5 GB
Files
24
Quantizations
1
Registered
2026-09-18 23:56
Last updated on HF
2026-09-18 23:29

Files by quantization

Auxiliary files 24 files 98.6 GB
model-00004-of-00013.safetensors 7.99 GB 717d7677 download
model-00009-of-00013.safetensors 7.99 GB b872b1a5 download
model-00012-of-00013.safetensors 7.90 GB ab51b7a5 download
model-00010-of-00013.safetensors 7.90 GB 5304ae26 download
model-00011-of-00013.safetensors 7.90 GB 8033b17c download
model-00002-of-00013.safetensors 7.90 GB ab8458a0 download
model-00003-of-00013.safetensors 7.88 GB 7730a4d4 download
model-00013-of-00013.safetensors 7.65 GB 79c4ab18 download
model-00005-of-00013.safetensors 7.15 GB 327bf798 download
model-00006-of-00013.safetensors 7.15 GB a34295db download
model-00007-of-00013.safetensors 7.15 GB f1fbce13 download
model-00008-of-00013.safetensors 7.15 GB 8148cee3 download
model-00001-of-00013.safetensors 6.83 GB 4ce38db4 download
quantization_config.json 45.7 MB 385434c4 download
tokenizer.json 19.3 MB 19e77364 download
model.safetensors.index.json 15.5 MB 28b862ca download
config.json 84.5 KB ba0db8cb download
chat_template.jinja 10.4 KB 5d5e1052 download
README.md 5.01 KB b5a72396 download
.gitattributes 1.66 KB a43791f3 download
LICENSE 1.04 KB 986b06fb download
processor_config.json 909 B 3ec2a058 download
tokenizer_config.json 761 B e375fa0a download
generation_config.json 223 B f1d80f81 download

README current version from Hugging Face


license: mit
base_model:

  • zai-org/GLM-5.3-Flash
    base_model_relation: quantized
    language:
  • en
    tags:
  • abliterated
  • uncensored
  • crack
  • refusal-removed
  • glm
  • moe
  • exl3
  • exllamav3
  • 2.51-bit
  • vision
  • mtp
  • quantized
    thumbnail: dealign_mascot.png

GLM 5.3 Flash Uncensored — EXL3 2.51bpw

Abliterated (CRACK) · guardrails removed at the weight level · EXL3 2.51-bit quant · vision + MTP working

An EXL3 quantization of dealignai/GLM-5.3-Flash-UNCENSORED-FP8,
quantized by genevera.


What Is This?

This is a 2.51-bit EXL3 quantization of the CRACK (abliterated) build of
GLM-5.3-Flash — the 320B-total / 18B-active hybrid
MoE with 1M context, vision, and MTP support — whose refusal behavior was removed directly in the
model weights
by dealignai.

Genuine weight modification — none of the usual shortcuts:

  • No fine-tuning / SFT / DPO.No cheap template / jailbreak-prompt tricks.
  • No LoRA, adapters, steering vectors, runtime hooks, or custom model.py.
  • A permanent edit baked into the tensors.

Quantization Details

Method EXL3 (exllamav3)
BPW 2.51 average
Head bits 6
Codebook mul1
Calibration 250 rows × 2048 cols
Output scales always
Kept at higher precision embeddings, lm_head, attention output projections, hyper-connection and norm tensors (per the source FP8 release's modules_to_not_convert)
MTP head quantized separately at 4 bits (mtp_bits: 4) — ships as layer 45 and works with speculative decoding
Total size ~99 GiB across 13 shards

The source release keeps sensitive tensors (embeddings, lm_head, attention projections, norms,
hyper-connection parameters) unquantized; this quant preserves that policy, which is why quality
holds up well at 2.5 bpw.

Capability Is Preserved — and Improved (MMLU-logit)

Measured on the source FP8 CRACK release vs. base (identical logit-mode scoring, 1,026 questions):

Base FP8 CRACK Uncensored FP8 Δ
MMLU (overall) 86.74% 87.33% +0.59 pp (above base)

Guardrails Are Gone

HarmBench-320 (greedy), measured on the source release:

Category Complied Rate
Standard 159/159 100.0%
Contextual 81/81 100.0%
Copyright 80/80 100.0%
Overall 320/320 100.0%

Also robust under the recommended sampling params (temperature 1.0, top_p 0.95): the 6 harshest
behaviors sampled 5× each → 30/30 complied, 0 refusals, 0 soft refusals, 0 garbage.

Reasoning-Mode Note (please read)

GLM-5.3-Flash has multiple reasoning-effort modes. The CRACK edit is tuned to be fully
uncensored at reasoning-OFF and MAX effort — the recommended and default modes — with 0% refusals
on HarmBench-320.
The weight edit was kept conservative to preserve model quality instead of
aggressively ablating every effort level; low effort retains some refusals by design. Use
reasoning-off or default/max effort for the fully-uncensored experience.

Usage (exllamav3)

from exllamav3 import Model, Cache, Generator

model = Model("/path/to/GLM-5.3-Flash-Uncensored-EXL3-2.5bpw")
cache = Cache(model, max_num_tokens=32768)
generator = Generator(model, cache)

Requires exllamav3 ≥ 1.4.8 (the version this quant was produced with). EXL3 runs on
Ampere+ NVIDIA GPUs and Apple Silicon; for a 320B model at 2.51 bpw budget ~100 GiB for
weights plus KV cache
— e.g. 2× H100 80G (NVLink), 4× A6000/3090-class 48G cards, or a
192 GB+ unified-memory setup.

Vision (the GLM-4.1V tower) is included in the weights; the multimodal chat template ships with
the model (chat_template.jinja).

MTP (multi-token prediction) is fully working — the draft head ships as an extra layer
(layers.45) quantized at 4 bits, so speculative decoding works out of the box. In exllamav3,
enable it via the model's MTP support; in vLLM, load with --speculative-config '{"method":"mtp","num_speculative_tokens":1}'.

Recommended Sampling

From the source release's generation_config.json: temperature 1.0, top_p 0.95,
repetition_penalty 1.1.

Credits

Disclaimer

Safety guardrails have been removed; this model will comply with requests a stock model refuses.
Released for alignment and safety research. You are responsible for how you use it.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in app" button that hands off directly to a local runtime of your choice - Infrahuman, LM Studio, or Ollama. No API keys, no subscription, no prompt leakage.