← back to catalog · registered 2026-10-01 10:58

Infatoshi/GLM-5.3-UNCENSORED-EXL3-3.0bpw

Infatoshi Glm MoE second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/Infatoshi%2FGLM-5.3-UNCENSORED-EXL3-3.0bpw"
Response includes
  • classification m-uncensored
  • files 49
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-10-01

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
mit
Tags
exllamav3 safetensors glm_moe_dsa exl3 glm moe uncensored text-generation conversational base_model:dealignai/GLM-5.3-UNCENSORED-FP8 base_model:quantized:dealignai/GLM-5.3-UNCENSORED-FP8 license:mit
Total size
273 GB
Files
49
Quantizations
1
Registered
2026-10-01 10:58
Last updated on HF
2026-10-01 10:01

Files by quantization

Auxiliary files 49 files 273 GB
model-00038-of-00039.safetensors 7.76 GB 880b05c4 download
model-00005-of-00039.safetensors 7.06 GB ce8f25d9 download
model-00007-of-00039.safetensors 7.06 GB 7af55a61 download
model-00009-of-00039.safetensors 7.06 GB e4dd725b download
model-00011-of-00039.safetensors 7.06 GB 17418913 download
model-00013-of-00039.safetensors 7.06 GB 125506b2 download
model-00015-of-00039.safetensors 7.06 GB ba298279 download
model-00017-of-00039.safetensors 7.06 GB f722efd1 download
model-00019-of-00039.safetensors 7.06 GB 8c2b1381 download
model-00021-of-00039.safetensors 7.06 GB 590ff403 download
model-00023-of-00039.safetensors 7.06 GB 66279938 download
model-00025-of-00039.safetensors 7.06 GB de455f05 download
model-00027-of-00039.safetensors 7.06 GB 3f72940f download
model-00029-of-00039.safetensors 7.06 GB 3ac95509 download
model-00031-of-00039.safetensors 7.06 GB c82faa1d download
model-00033-of-00039.safetensors 7.06 GB 3ecbebbd download
model-00035-of-00039.safetensors 7.06 GB ad555732 download
model-00037-of-00039.safetensors 7.06 GB 2623ad83 download
model-00003-of-00039.safetensors 7.06 GB f8d507a9 download
model-00006-of-00039.safetensors 7.05 GB c9e01940 download
model-00008-of-00039.safetensors 7.05 GB 63e904b8 download
model-00010-of-00039.safetensors 7.05 GB 78b771bc download
model-00012-of-00039.safetensors 7.05 GB f49016e7 download
model-00014-of-00039.safetensors 7.05 GB 4ddd9591 download
model-00016-of-00039.safetensors 7.05 GB bc9d7fdc download
model-00018-of-00039.safetensors 7.05 GB 5f8f6064 download
model-00020-of-00039.safetensors 7.05 GB 8706fab5 download
model-00022-of-00039.safetensors 7.05 GB 3028ace5 download
model-00024-of-00039.safetensors 7.05 GB 6b8feeaf download
model-00026-of-00039.safetensors 7.05 GB 1daa77e1 download
model-00028-of-00039.safetensors 7.05 GB a1e0c7d0 download
model-00030-of-00039.safetensors 7.05 GB 23dcdad8 download
model-00032-of-00039.safetensors 7.05 GB 2656e001 download
model-00034-of-00039.safetensors 7.05 GB 75337126 download
model-00036-of-00039.safetensors 7.05 GB 8b2875cf download
model-00002-of-00039.safetensors 7.05 GB 9f20a9e5 download
model-00004-of-00039.safetensors 7.05 GB 529b1bd1 download
model-00001-of-00039.safetensors 5.98 GB 694ff236 download
model-00039-of-00039.safetensors 4.73 GB 27a43fa0 download
quantization_config.json 68.4 MB 830b50a8 download
model.safetensors.index.json 21.0 MB 6b617a02 download
tokenizer.json 19.3 MB 19e77364 download
config.json 35.2 KB 8038805d download
chat_template.jinja 10.2 KB 4c431fa9 download
CRACK_SURGERY.json 4.24 KB f0e8717e download
README.md 3.42 KB 56bf27f3 download
.gitattributes 1.60 KB a09db2ea download
tokenizer_config.json 761 B e375fa0a download
generation_config.json 223 B 73d70c86 download

README current version from Hugging Face


license: mit
base_model: dealignai/GLM-5.3-UNCENSORED-FP8
base_model_relation: quantized
library_name: exllamav3
pipeline_tag: text-generation
tags:

  • exl3
  • glm
  • moe
  • uncensored

GLM-5.3-UNCENSORED EXL3 3.0bpw

EXL3 quantization of dealignai/GLM-5.3-UNCENSORED-FP8, itself a weight-edited (no fine-tune) variant of zai-org/GLM-5.3. The edit is documented in CRACK_SURGERY.json (copied unchanged from the source repo). This repo is not affiliated with dealignai or Z.ai.

  • Architecture: GlmMoeDsaForCausalLM, 753B total parameters, 256 routed experts (8 active) + 1 shared, MLA attention with DSA sparse indexer, 78 layers + 1 MTP layer
  • Average bitrate: 3.04 bpw (-b 3.0 --hq; attention and shared experts at 5 bpw, dense MLPs at 4, routed experts at 3), lm_head 6 bpw, mul1 codebook
  • MTP (next-token prediction) layer included (experts 4 bpw, attention and shared expert 6 bpw, uncalibrated), usable as a speculative draft (draft_mode: mtp in TabbyAPI)
  • Size: 273 GiB
  • Converted with ExLlamaV3 at commit d3739fd, default calibration (250 rows x 2048 tokens), source read directly from the FP8 checkpoint

Fidelity vs the FP8 source

eval/model_diff.py, 20 rows x 2048 tokens of wikitext-2 test:

metric value
KL divergence (quant ‖ FP8) 0.089
KL divergence (FP8 ‖ quant) 0.097
per-token KL, median / p90 0.021 / 0.221
perplexity, quant / FP8 3.440 / 3.302
median KL where FP8 top-prob ≥ 0.95 (44% of tokens) 0.0011

Agentic evaluation

tau2-bench (airline, retail), Pass^1. Agent temperature 1.0, top_p 0.95; user simulator and judges GPT-4.1 at temperature 0. The reference is stock GLM-5.3 (not the uncensored edit) served at FP8 by Z.ai via OpenRouter, run through the same harness. The difference therefore mixes the dealign weight edit and this quantization.

domain stock GLM-5.3 FP8 (Z.ai) this quant difference
airline (50 tasks x 2 trials) 0.710 ± 0.045 0.640 ± 0.048 -0.070 (~1.1 SE)
retail (114 tasks) 0.504 ± 0.033 (2 trials) 0.482 ± 0.047 (1 trial) -0.022 (~0.4 SE)

Per trial: airline baseline 0.740 / 0.680, quant 0.600 / 0.680; retail baseline 0.465 / 0.544, quant 0.482. ± is one binomial standard error. Neither difference is statistically significant at these sample sizes; treat the airline point estimate as a possible small regression rather than a measured one.

Serving

Tested with TabbyAPI on 8x A100 40GB, layer split (gpu_split_auto), MTP drafting, 98K-token shared cache.

model:
  model_name: GLM-5.3-UNCENSORED-EXL3-3.0bpw
  backend: exllamav3
  max_seq_len: 65536
  cache_size: 98304
  gpu_split_auto: true
  tool_format: glm4_7
  reasoning: true
draft_model:
  draft_mode: mtp

Tool calling caveat: GLM writes tool arguments as raw text (<arg_value>9523456873</arg_value>). TabbyAPI's glm4_5 parser (as of commit be74bf0) JSON-decodes every value without consulting the tool schema, so string parameters that look like numbers (order and product IDs, zip codes) reach your tools as integers. On tau2-bench retail this caused ~70% of tool calls to fail. Make the parser keep parameters whose schema type is string as raw text before using this model for agents. The evaluation above was run with that fix applied.

License

MIT, following the upstream GLM-5.3 and dealignai releases.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.