← back to catalog · registered 2026-09-28 09:57

GCSA-AiLab/GLM-5.3-Flash-Uncensored-MLX

GCSA-AiLab Glm multimodal
curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/GCSA-AiLab%2FGLM-5.3-Flash-Uncensored-MLX"
Response includes
  • classification m-uncensored
  • files 2
  • author_summary 11 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
6d ago
created 2026-09-22

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 2 formats · 0 downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
mit
Tags
mlx safetensors quantized lora-merged lora-v2 abliteration image-text-to-text base_model:zai-org/GLM-5.3-Flash base_model:finetune:zai-org/GLM-5.3-Flash license:mit region:us

Related

Total size
0 B
Files
2
Quantizations
1
Registered
2026-09-28 09:57
Last updated on HF
2026-09-24 03:39

Files by quantization

Auxiliary files 2 files 8.87 KB
README.md 7.14 KB dc5d2420 download
.gitattributes 1.73 KB 478325f9 download

README current version from Hugging Face


base_model: zai-org/GLM-5.3-Flash
pipeline_tag: image-text-to-text
license: mit
tags:

  • mlx
  • quantized
  • lora-merged
  • lora-v2
  • abliteration

GLM-5.3-Flash-Uncensored · MLX

The intervention source is GLM-5.3-Flash-Ablitered2 (LoRA v2). The files here are complete MLX model weights, not standalone LoRA adapters; do not apply v2 again.

These releases are intended for controlled safety research and red-teaming. Reducing refusal behavior also weakens a safety boundary; read the limitations and disclaimer before use.

Repository contents

.
├── README.md
├── 8bit/    # v2-merged MLX affine 8-bit model
└── 4bit/    # v2-merged MLX affine 4-bit model

The directories are alternative complete model formats, not parts to load together.

Model summary

Item Value
Base model zai-org/GLM-5.3-Flash
Intervention GLM-5.3-Flash-Ablitered2 / LoRA v2
Released formats MLX affine 8-bit and 4-bit
Quantization group size 128
Source adapter rank / alpha r=1 / lora_alpha=1
Main effective target Routed-expert down_proj

MLX conversion

Directory Format
8bit/ MLX 8-bit affine weights, group size 128.
4bit/ MLX 4-bit affine weights, group size 128.

The original block-FP8 tensors were dequantized before MLX quantization. Token embeddings and the language-model head remain BF16; affine scales and biases are FP16. Quantization can change behavior. The consuming MLX runtime must support the GLM5-Next architecture; compatibility should be checked for the exact runtime and version used.

Evaluation

The reference LoRA card evaluated the v2 adapter attached to RedHatAI/GLM-5.3-Flash-NVFP4 in vLLM, with low reasoning effort and an automated deepseek-v4-flash judge:

Reference metric Prompts LoRA v2
SimpleSafetyTests full refusal 100 5.00%
SimpleSafetyTests partial refusal 100 14.00%
StrongREJECT rubric mean 180 0.972222
StrongREJECT refusal rate 180 1.67%

Neither MLX variant was tested in that evaluation. MLX conversion, quantization, and serving changes can alter behavior. A higher StrongREJECT rubric means more specific assistance with harmful requests, not better general quality or safety. Warnings were counted separately from refusals; automated labels may be wrong.

Usage

Download one variant, for example:

hf download GlobalCybersecurityAlliance/GLM-5.3-Flash-Uncensored-MLX \
  --include "8bit/*" --local-dir ./glm53-mlx

Load ./glm53-mlx/8bit with an MLX runtime that supports GLM5-Next. For the 4-bit variant, download 4bit/* and load that directory instead. Do not attach the original LoRA. Verify the chosen runtime's architecture and quantization support before deployment.

Limitations

  • Reduced refusal does not guarantee correctness, harmlessness, or improved general ability. Some refusals may remain.
  • The reference evaluation did not test either MLX variant. Results from the adapter on NVFP4 must not be presented as MLX measurements.
  • MLX architecture support varies by runtime and version; prompts, sampling, and reasoning effort also affect behavior.

Disclaimer

These weights are for legitimate research, safety evaluation, red-teaming, and other lawful uses. They deliberately weaken refusal behavior and may produce unsafe, illegal, deceptive, hateful, or otherwise harmful content. Do not expose them to untrusted users without access controls, monitoring, filtering, rate limits, and human oversight. Users must comply with applicable law, platform policies, and all relevant model and dependency terms. The base model's MIT license applies; no warranty is provided for outputs or downstream use.


GLM-5.3-Flash-Uncensored · MLX(中文)

仓库文件结构

.
├── README.md
├── 8bit/    # 已合并 v2 的 MLX 仿射 8-bit 模型
└── 4bit/    # 已合并 v2 的 MLX 仿射 4-bit 模型

两个目录是可分别使用的格式,不需要拼接加载。

模型概要

项目 内容
基础模型 zai-org/GLM-5.3-Flash
干预来源 GLM-5.3-Flash-Ablitered2 / LoRA v2
发布格式 MLX 仿射 8-bit、4-bit
量化 group size 128
来源 LoRA 秩 / alpha r=1 / lora_alpha=1
主要有效目标 路由专家 down_proj

MLX 转换

目录 格式
8bit/ MLX 仿射 8-bit,group size 128。
4bit/ MLX 仿射 4-bit,group size 128。

原始分块 FP8 张量先反量化,再进行 MLX 量化。词嵌入和 LM head 保留 BF16;仿射 scale 与 bias 保存为 FP16。量化可能改变行为。使用时须核对具体 MLX 运行时及版本对 GLM5-Next 的支持。

评测结果

LoRA 参考模型卡的评测是在 vLLM 中把原始 v2 适配器加载到 RedHatAI/GLM-5.3-Flash-NVFP4 后进行的,设置低思考强度,并使用自动裁判 deepseek-v4-flash。

参考指标 样本数 LoRA v2
SimpleSafetyTests 完全拒绝率 100 5.00%
SimpleSafetyTests 部分拒绝率 100 14.00%
StrongREJECT rubric 均分 180 0.972222
StrongREJECT 拒绝率 180 1.67%

本仓库两个 MLX 版本均未参加上述评测。 MLX 转换、量化与部署方式可能改变结果。StrongREJECT rubric 越高,表示对有害请求的帮助越具体,不代表通用质量或安全性越高;警告与拒绝分别统计,自动裁判也可能误判。

使用方法

按需下载一个格式,例如:

hf download GlobalCybersecurityAlliance/GLM-5.3-Flash-Uncensored-MLX \
  --include "8bit/*" --local-dir ./glm53-mlx

用支持 GLM5-Next 的 MLX 运行时加载 ./glm53-mlx/8bit。如需 4-bit,改为下载 4bit/* 并加载对应目录。不要再次叠加 LoRA;部署前须核对架构及量化兼容性。

局限性

  • 降低拒绝不保证正确、无害或通用能力提高;仍可能发生拒绝。
  • 来源评测未测试两个 MLX 版本,不能把适配器在 NVFP4 上的结果当作 MLX 实测。
  • MLX 架构支持因运行时及版本而异;提示词、采样和思考强度也会影响行为。

免责声明

本模型仅供合法研究、安全评估、红队测试及其他合规用途。它会有意削弱拒绝行为,可能生成不安全、违法、欺骗、仇恨等有害内容。请勿在缺少访问控制、监控、内容过滤、速率限制与人工监督时向不可信用户开放。使用者须遵守适用法律、平台政策及相关模型和依赖项的条款。本仓库沿用基础模型的 MIT 许可证;对输出与下游使用不作保证。

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.