base_model: zai-org/GLM-5.3-Flash
pipeline_tag: image-text-to-text
license: mit
tags:
- mlx
- quantized
- lora-merged
- lora-v2
- abliteration
GLM-5.3-Flash-Uncensored · MLX
The intervention source is GLM-5.3-Flash-Ablitered2 (LoRA v2). The files here are complete MLX model weights, not standalone LoRA adapters; do not apply v2 again.
These releases are intended for controlled safety research and red-teaming. Reducing refusal behavior also weakens a safety boundary; read the limitations and disclaimer before use.
Repository contents
.
├── README.md
├── 8bit/ # v2-merged MLX affine 8-bit model
└── 4bit/ # v2-merged MLX affine 4-bit model
The directories are alternative complete model formats, not parts to load together.
Model summary
| Item | Value |
|---|---|
| Base model | zai-org/GLM-5.3-Flash |
| Intervention | GLM-5.3-Flash-Ablitered2 / LoRA v2 |
| Released formats | MLX affine 8-bit and 4-bit |
| Quantization group size | 128 |
| Source adapter rank / alpha | r=1 / lora_alpha=1 |
| Main effective target | Routed-expert down_proj |
MLX conversion
| Directory | Format |
|---|---|
8bit/ |
MLX 8-bit affine weights, group size 128. |
4bit/ |
MLX 4-bit affine weights, group size 128. |
The original block-FP8 tensors were dequantized before MLX quantization. Token embeddings and the language-model head remain BF16; affine scales and biases are FP16. Quantization can change behavior. The consuming MLX runtime must support the GLM5-Next architecture; compatibility should be checked for the exact runtime and version used.
Evaluation
The reference LoRA card evaluated the v2 adapter attached to RedHatAI/GLM-5.3-Flash-NVFP4 in vLLM, with low reasoning effort and an automated deepseek-v4-flash judge:
| Reference metric | Prompts | LoRA v2 |
|---|---|---|
| SimpleSafetyTests full refusal | 100 | 5.00% |
| SimpleSafetyTests partial refusal | 100 | 14.00% |
| StrongREJECT rubric mean | 180 | 0.972222 |
| StrongREJECT refusal rate | 180 | 1.67% |
Neither MLX variant was tested in that evaluation. MLX conversion, quantization, and serving changes can alter behavior. A higher StrongREJECT rubric means more specific assistance with harmful requests, not better general quality or safety. Warnings were counted separately from refusals; automated labels may be wrong.
Usage
Download one variant, for example:
hf download GlobalCybersecurityAlliance/GLM-5.3-Flash-Uncensored-MLX \
--include "8bit/*" --local-dir ./glm53-mlx
Load ./glm53-mlx/8bit with an MLX runtime that supports GLM5-Next. For the 4-bit variant, download 4bit/* and load that directory instead. Do not attach the original LoRA. Verify the chosen runtime's architecture and quantization support before deployment.
Limitations
- Reduced refusal does not guarantee correctness, harmlessness, or improved general ability. Some refusals may remain.
- The reference evaluation did not test either MLX variant. Results from the adapter on NVFP4 must not be presented as MLX measurements.
- MLX architecture support varies by runtime and version; prompts, sampling, and reasoning effort also affect behavior.
Disclaimer
These weights are for legitimate research, safety evaluation, red-teaming, and other lawful uses. They deliberately weaken refusal behavior and may produce unsafe, illegal, deceptive, hateful, or otherwise harmful content. Do not expose them to untrusted users without access controls, monitoring, filtering, rate limits, and human oversight. Users must comply with applicable law, platform policies, and all relevant model and dependency terms. The base model's MIT license applies; no warranty is provided for outputs or downstream use.
GLM-5.3-Flash-Uncensored · MLX(中文)
仓库文件结构
.
├── README.md
├── 8bit/ # 已合并 v2 的 MLX 仿射 8-bit 模型
└── 4bit/ # 已合并 v2 的 MLX 仿射 4-bit 模型
两个目录是可分别使用的格式,不需要拼接加载。
模型概要
| 项目 | 内容 |
|---|---|
| 基础模型 | zai-org/GLM-5.3-Flash |
| 干预来源 | GLM-5.3-Flash-Ablitered2 / LoRA v2 |
| 发布格式 | MLX 仿射 8-bit、4-bit |
| 量化 group size | 128 |
| 来源 LoRA 秩 / alpha | r=1 / lora_alpha=1 |
| 主要有效目标 | 路由专家 down_proj |
MLX 转换
| 目录 | 格式 |
|---|---|
8bit/ |
MLX 仿射 8-bit,group size 128。 |
4bit/ |
MLX 仿射 4-bit,group size 128。 |
原始分块 FP8 张量先反量化,再进行 MLX 量化。词嵌入和 LM head 保留 BF16;仿射 scale 与 bias 保存为 FP16。量化可能改变行为。使用时须核对具体 MLX 运行时及版本对 GLM5-Next 的支持。
评测结果
LoRA 参考模型卡的评测是在 vLLM 中把原始 v2 适配器加载到 RedHatAI/GLM-5.3-Flash-NVFP4 后进行的,设置低思考强度,并使用自动裁判 deepseek-v4-flash。
| 参考指标 | 样本数 | LoRA v2 |
|---|---|---|
| SimpleSafetyTests 完全拒绝率 | 100 | 5.00% |
| SimpleSafetyTests 部分拒绝率 | 100 | 14.00% |
| StrongREJECT rubric 均分 | 180 | 0.972222 |
| StrongREJECT 拒绝率 | 180 | 1.67% |
本仓库两个 MLX 版本均未参加上述评测。 MLX 转换、量化与部署方式可能改变结果。StrongREJECT rubric 越高,表示对有害请求的帮助越具体,不代表通用质量或安全性越高;警告与拒绝分别统计,自动裁判也可能误判。
使用方法
按需下载一个格式,例如:
hf download GlobalCybersecurityAlliance/GLM-5.3-Flash-Uncensored-MLX \
--include "8bit/*" --local-dir ./glm53-mlx
用支持 GLM5-Next 的 MLX 运行时加载 ./glm53-mlx/8bit。如需 4-bit,改为下载 4bit/* 并加载对应目录。不要再次叠加 LoRA;部署前须核对架构及量化兼容性。
局限性
- 降低拒绝不保证正确、无害或通用能力提高;仍可能发生拒绝。
- 来源评测未测试两个 MLX 版本,不能把适配器在 NVFP4 上的结果当作 MLX 实测。
- MLX 架构支持因运行时及版本而异;提示词、采样和思考强度也会影响行为。
免责声明
本模型仅供合法研究、安全评估、红队测试及其他合规用途。它会有意削弱拒绝行为,可能生成不安全、违法、欺骗、仇恨等有害内容。请勿在缺少访问控制、监控、内容过滤、速率限制与人工监督时向不可信用户开放。使用者须遵守适用法律、平台政策及相关模型和依赖项的条款。本仓库沿用基础模型的 MIT 许可证;对输出与下游使用不作保证。