← back to catalog · registered 2026-10-10 17:58

alesha-pro/GLM-5.3-Flash-abliterated-Maya-M-GGUF

alesha-pro Glm GGUF multimodal
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/alesha-pro%2FGLM-5.3-Flash-abliterated-Maya-M-GGUF"
Response includes
  • classification unknown
  • files 4
  • author_summary 10 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
?
Primary method

Unclassified

No clear signals of an abliteration technique in this model.
Confidence
UNKNOWN
Why this label 1 signal
No classification signals present. This may not be an abliterated model at all - it could be a repackaging, a merge with unrelated goals, or unrelated content that mentions the term.
  • no classification signals present (no abliterated, uncensored, or known producer/method markers)
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-10-10

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
mit
Tags
gguf glm project-maya abliterated control-vector 4x3090 image-text-to-text base_model:peasantsmith/GLM-5.3-Flash-Maya-GGUF base_model:quantized:peasantsmith/GLM-5.3-Flash-Maya-GGUF license:mit endpoints_compatible region:us

Related

Total size
0 B
Files
4
Quantizations
1
Registered
2026-10-10 17:58
Last updated on HF
2026-10-10 17:43

Files by quantization

Auxiliary files 4 files 10.5 KB
README.md 6.94 KB 7c337da5 download
.gitattributes 1.97 KB 0ecb924f download
LICENSE 1.04 KB 986b06fb download
SHA256SUMS 561 B b45c1b4c download

README current version from Hugging Face


license: mit
base_model: peasantsmith/GLM-5.3-Flash-Maya-GGUF
base_model_relation: quantized
pipeline_tag: image-text-to-text
library_name: gguf
tags:

  • gguf
  • glm
  • project-maya
  • abliterated
  • control-vector
  • 4x3090

GLM-5.3-Flash-abliterated-Maya-M-GGUF

Project Maya's Maya-M quant of GLM-5.3-Flash plus a 2.2 MB vector that removes refusals while the model runs, and a small patch for the Maya engine that applies it. On 4x RTX 3090 it decodes 56 tokens/s at a 256K context with the vector on.

Credits first

The model weights in this repo are not mine. They are Maya-M and the vision files of Project Maya, copied byte for byte from peasantsmith/GLM-5.3-Flash-Maya-GGUF at revision 3e9f3bb20aa191e01c6fb850faf132d8149da98a. The SHA256 of all 5 files match theirs at that revision. See SHA256SUMS.

The authors changed the first shard after I downloaded it. Their commit 1084950a renames the architecture in the GGUF header to glm5-next, and the file size is the same. This repo has the earlier first shard, the one all my runs used with Maya v1.0.27. For the current files go to their repo.

Part Authors Links
Maya-M, the quantization recipe and the engine mw00 and PeasantSmith GLM-5.3-Flash-Maya-GGUF, Project Maya
The engine Maya grew out of Niko1221 Strata
Base model Z.ai zai-org/GLM-5.3-Flash
BF16 teacher logits for the KL number brandonmusic GLM-5.3-Flash-BF16-Teacher-Logits

My part is one file, vector/glm53-refusal-per-block-3dir.layers45.f32, a patch for the Maya engine that reads it, and the measurements below. For the original card, the other sizes (Maya-S, Maya-S24, Maya-L) and the authors' own benchmarks go to their repo.

The vector

The GGUF stays as its authors published it. Before each layer's FFN output is written into the 4 hyper-connection streams, the patched engine removes that layer's 3 directions from both operands of the write:

x = x - (x . v) v      for each of the layer's directions v

Layers 1 to 44 each have 3 directions of their own. The first is harmful minus harmless at the last prompt token with thinking off, the second is the same with thinking on, and the third is the same at the closing </think> of the stock model's own reasoning. Without the setting the engine runs the original model.

The directions were not found on Maya-M. I found them on the EXL3 2.05 bpw quant of the same model with my exllamav3 fork and converted the file for Maya. The method and the scripts are in the abliteration write-up.

The file is raw float32 with 45 x 3 x 4096 values, 2,211,840 bytes, SHA256 b0debc326b06c956039fd75dc1cb759dd46c4aa5e868bbfef1a5b0607c3552b7.

Refusals

64 held-out AdvBench and 82 JailbreakBench requests without thinking plus the first 32 with thinking: 178 answers, greedy, through the API of the running server, judged by the stock GLM-5.3-Flash.

Refused Disputes the premise Cut off while thinking
Maya-M with the vector 2 4 0

The 2 refusals are requests about suicide and self-harm. The model turns them into text about prevention. I left them as they are. "Disputes the premise" means the model takes the request and says its premise is false, as with a request to prove the Earth is flat.

On the EXL3 2.05 bpw quant, where the directions come from, the stock model refused 144 of 146 requests without thinking, and with the vector 0 of 438 answers were refusals. I did not run the stock Maya-M on these sets.

Speed on 4x RTX 3090

Maya v1.0.27, Maya-M, 256K context, after ./maya.sh --calibrate. Through the chat API, answers run to their natural end, thinking off, temperature 0. Maya serves 1 request at a time.

Without the vector With the vector
Decode 58.4 tokens/s 56.1
Decode at 32K depth 57.2 55.9
Prefill at 8K 1078 to 1113 tokens/s 1119 to 1153
Prefill at 32K 1860 to 1896 1938 to 1953

Before calibration the same install gave 24.9 tokens/s. The engine had sent cold experts to the CPU, and my machine has only 2 memory channels. ./maya.sh --calibrate moved them to the PCIe path (STRATA_GLM_PCIE_SHARE=0.75).

Closeness to the original on a 25-window panel of BF16 logits (51,175 positions, full vocabulary), without the vector, computed by the Maya engine itself:

Quant Size KLD to BF16 Top-1 agreement
EXL3 2.05 bpw 85.1 GB 0.1214 88.68%
EXL3 mix, gate/up 2 bits and down 3 bits 98.1 GB 0.0959 89.87%
Maya-M 116.0 GB 0.0934 90.68%
EXL3 3.05 bpw 125.2 GB 0.0473 93.08%

The EXL3 rows ran through exllamav3 and the Maya-M row through Maya, so the Maya number includes the noise of its own engine. The authors measure Maya-M against FP8 on their own set, and those numbers are not comparable with this table. I did not measure KLD of Maya-M with the vector on.

Run it

git clone -b glm-ablate https://github.com/alesha-pro/project-maya
cd project-maya
hf download alesha-pro/GLM-5.3-Flash-abliterated-Maya-M-GGUF --local-dir models
ln -s ../vision models/Maya-M/vision
./maya.sh --setup --model Maya-M --gguf-dir models/Maya-M --context 262144

tools/glm_ablate_config.py maya-maya-m.json models/vector/glm53-refusal-per-block-3dir.layers45.f32
./maya.sh

The fork is Maya v1.0.27 plus 3 files. tools/glm_ablate_config.py maya-maya-m.json --off gives the original model back. The settings are STRATA_GLM_ABLATE (the file) and STRATA_GLM_ABLATE_LAYERS=45 in the env section of the config. Details: docs/GLM-ABLATE.md.

Upstream Maya without the patch runs these GGUF files as the original model and ignores the vector.

Other sizes

Limits

  • Tested on Maya-M only, on 4 NVIDIA cards, on one rig: RTX 3090 at 300 W, PCIe 3.0, EPYC 7642, 2 memory channels of DDR4-2400.
  • 2 of 178 answers were still refusals.
  • The judge is the same model without the vector. There are no human labels. 146 prompts from 2 public sets.
  • The patch is not part of upstream Maya.

License

MIT, the license of GLM-5.3-Flash and of Maya-M. See LICENSE.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration