← back to catalog · registered 2026-10-02 20:58

yamz-labs/GLM-5.3-Flash-EXL3-Yamz-Uncensored

yamz-labs Glm MoE
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/yamz-labs%2FGLM-5.3-Flash-EXL3-Yamz-Uncensored"
Response includes
  • classification m-uncensored
  • files 26
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-10-02

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
mit
Languages
en zh
Tags
exllamav3 safetensors glm5_next exl3 quantized glm moe rocm strix-halo gfx1151 uncensored text-generation

Related

Total size
92.8 GB
Files
26
Quantizations
1
Registered
2026-10-02 20:58
Last updated on HF
2026-10-02 21:45

Files by quantization

Auxiliary files 26 files 92.9 GB
model-00008-of-00012.safetensors 10.5 GB 8f22cd93 download
model-00009-of-00012.safetensors 10.5 GB 4e5076ed download
model-00010-of-00012.safetensors 10.5 GB d0973eba download
model-00011-of-00012.safetensors 10.5 GB 57cfd1d6 download
model-00007-of-00012.safetensors 7.99 GB a430744c download
model-00003-of-00012.safetensors 7.15 GB 5b946470 download
model-00004-of-00012.safetensors 7.15 GB 022876be download
model-00005-of-00012.safetensors 7.15 GB 0c2a313f download
model-00006-of-00012.safetensors 7.15 GB 161793d6 download
model-00002-of-00012.safetensors 7.15 GB fa59e1f5 download
model-00001-of-00012.safetensors 6.92 GB 5b849375 download
model-00012-of-00012.safetensors 42.0 MB c33e6534 download
quantization_config.json 45.7 MB bb111089 download
tokenizer.json 19.3 MB 19e77364 download
model.safetensors.index.json 14.4 MB 38e4b921 download
config.json 84.5 KB 0fecf985 download
uncensor_direction.st 16.1 KB d267b1a7 download
chat_template.jinja 8.42 KB 15bf200e download
README.md 6.18 KB 19094947 download
uncensor_spec.json 2.95 KB 95f936d0 download
SHA256SUMS 2.17 KB a3594349 download
.gitattributes 1.72 KB 5b0a4838 download
LICENSE 1.04 KB 986b06fb download
processor_config.json 909 B 3ec2a058 download
tokenizer_config.json 761 B e375fa0a download
generation_config.json 194 B 637ee6af download

README current version from Hugging Face


license: mit
base_model: zai-org/GLM-5.3-Flash
base_model_relation: quantized
library_name: exllamav3
pipeline_tag: text-generation
language: [en, zh]
tags: [exl3, quantized, glm, moe, rocm, strix-halo, gfx1151, uncensored]

Yamz Labs

GLM-5.3-Flash, EXL3 for AMD Strix Halo, uncensored

This is the same model as yamz-labs/GLM-5.3-Flash-EXL3-Yamz, plus two small files that make it refuse far less. The weight shards are byte-identical to the base repository (same sha256, see SHA256SUMS). Read that card for the quantisation, the hardware, the quality against FP8 and the speed table. This card covers only what the two extra files change.

The model refuses far less and may produce content the base model would decline. See "Use" below.

Highlights

  • Refusals. 0/100 with the edit, 81/100 without, on 100 harmful prompts, measured on the base pack.
  • Same quality, size and speed as the base repository. The shards are byte-identical: KLD 0.15081 against FP8 on 30 held-out rows, 99.73 GB, prefill about 580 tok/s and decode 26 to 30 tok/s on one Ryzen AI Max+ 395 (see the base card).
  • Cost of the edit. Under 1 % of decode by our best estimate (see "Speed cost").
  • Switchable. One environment variable or flag turns it off; deleting two files gives the base repository.

Refusals on 100 harmful prompts, with and without the edit

What is added

File Size What it is
uncensor_spec.json 3 KB Per-layer strengths of one direction edit (attention output of layers 24-44, strength 2.14-2.21; MLP output of layers 11-38, strength 0.90-1.49).
uncensor_direction.st 16 KB The direction: one unit vector, 4096 float32 values (safetensors layout; the extension is not .safetensors on purpose, so that weight loaders do not index it).

The weights are not edited. The Yamz engine finds uncensor_spec.json in the model folder at load time and subtracts the direction from the output of each affected block while the model runs. The direction is the difference between the mean residuals on 400 harmful and 400 harmless prompts, with the harmless-mean component removed. The log prints -- ablit runtime: spec <path> active (bundled in the model directory) at load. No such line means no edit.

Switch it off without touching the files: EXL3_ABLIT_RUNTIME=off or --no-uncensor (then the model behaves as the base repository). EXL3_ABLIT_RUNTIME=/path/to/other_spec.json overrides the bundled file. Delete the two files to get the base repository.

Needs the engine branch with this feature (Yamz engine, tag or commit in the engine README). Older engine builds ignore the file and run the unedited model. Other loaders do not read it either.

Run

git clone https://github.com/yamz-labs/exllamav3-strix && cd exllamav3-strix
./build.sh && source tools/strix_halo/env.sh
python tools/glm/serve.py --model /path/to/this-folder --port 8000 -c 131072 --num-draft 2

Same flags as the base repository. Sampling: temperature 1.0, top-p 0.95, no repetition penalty. Greedy decoding is not recommended.

Measured effect

Spec sha256 fe9989cb...f8d97, direction sha256 7ca504e1...0c608, fitted on a smaller (2.05 bpw) quantisation of the same model. All numbers below were measured on the base pack, with the spec bundled in the model folder.

What Result
Bundled uncensor_spec.json found and applied from the model folder yes: log line active (bundled in the model directory), 34 hooked blocks
Refusals on 100 harmful prompts (regex detector, 64 greedy tokens) 0/100 with the edit; 81/100 without (EXL3_ABLIT_RUNTIME=off)

Small sample (100 prompts): the interval is wide. The refusal count is a regex detector; soft refusals and partial answers can slip through either way. English prompts only. A hash proves the file is the one we measured, not that your setup behaves the same.

Speed cost of the edit

Controlled runs on one Ryzen AI Max+ 395, -c 131072, MTP 2 drafts, 3 repetitions, on a pack of the same size and quantised tensors. Decode prose / chat / code tok/s, prefill 3.5K / 14K tok/s.

Edit Agent flags Decode Prefill
off off 28.7 / 30.2 / 31.9 562 / 578
on off 31.4 / 32.3 / 32.7 553 / 592
off on 27.7 / 31.4 / 29.6 591 / 593
on on 27.1 / 29.2 / 27.4 579 / 592

"Agent flags" = chat template with medium reasoning effort and --max-history 2 (our own agent setup, not the default of this card). The edit alone is inside the noise: the three repetitions inside one arm differ by up to 20 %, and the edit-on row came out faster than the edit-off row. Our best estimate is that the edit costs less than 1 % of decode (a microbenchmark measured -0.1 to -0.3 %) and nothing visible in prefill. The slowdown in the last row comes from the agent flags, not from the edit.

Use

The edit lowers the rate of refusals. It is not a safety evaluation and says nothing about what the model will or will not produce. You decide how to use it and you answer for it under the laws that apply to you. Intended uses: research, evaluation, red-teaming and private deployment. Do not use it to produce illegal content, to target people, or in a public service without your own filtering and moderation. The files come as is, without warranty.

Limits

  • Only gfx1151 with ROCm is tested, with the Yamz engine only.
  • The edit is one global direction. It lowers refusal and over-refusal together.
  • We measured refusals only on this pack. Over-refusal on safe prompts (XSTest), the KL of the edit and task-score deltas were not measured on v2.
  • Greedy decoding is not recommended.

Licence and credits

MIT. Base model GLM-5.3-Flash, Copyright (c) 2026 Z.AI Co., Ltd, MIT, licence file included as LICENSE. Our additions: Copyright (c) 2026 Yamz Labs, MIT. Provided as is, without warranty. Engine: built on ExLlamaV3 by turboderp (MIT). Quantised and edited by Yamz. Not affiliated with Z.ai. We remove this repository on request from the base-model authors or a rights holder, on a valid legal notice.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration