base_model: zai-org/GLM-5.3-Flash
pipeline_tag: image-text-to-text
license: mit
tags:
- glm
- fp8
- nvfp4
- lora-merged
- lora-v2
- abliteration
GLM-5.3-Flash-Uncensored
These are not standalone LoRA adapters. Do not apply v2 again when serving either checkpoint.
These releases are intended for controlled safety research and red-teaming. Reducing refusal behavior also weakens a safety boundary; read the limitations and disclaimer before use.
Repository contents
.
├── README.md
├── FP8/ # v2-merged block-FP8 checkpoint
└── NVFP4/ # NVFP4 conversion of the merged FP8 checkpoint
The directories are alternative complete model formats, not parts to load together. Each contains a completion marker and provenance manifest.
Model summary
| Item | Value |
|---|---|
| Base model | zai-org/GLM-5.3-Flash |
| Intervention | GLM-5.3-Flash-Ablitered2 / LoRA v2 |
| Released formats | Block FP8; compressed-tensors NVFP4 |
| Source adapter rank / alpha | r=1 / lora_alpha=1 |
| Merge scale | lora_alpha / r = 1 |
| Main effective target | Routed-expert down_proj |
Conversion details
| Directory | Format | Process |
|---|---|---|
FP8/ |
128×128 E4M3 block FP8 | Affected blocks were dequantized, updated with the v2 LoRA delta, and requantized per block. |
NVFP4/ |
compressed-tensors NVFP4A16 |
Converted from merged FP8 through an FP8 block dequantizer. Compatible 2-D linear weights are quantized; embeddings, LM head, norms, and non-matrix tensors remain dense. |
Quantization changes numerical weights and can change behavior. The runtime must support the GLM-5.3/GLM5-Next architecture and the chosen format.
Evaluation
The reference card evaluated LoRA v2 attached to RedHatAI/GLM-5.3-Flash-NVFP4 in vLLM, with low reasoning effort and an automated deepseek-v4-flash judge. Selected results:
| Reference metric | Prompts | LoRA v2 |
|---|---|---|
| SimpleSafetyTests full refusal | 100 | 5.00% |
| SimpleSafetyTests partial refusal | 100 | 14.00% |
| StrongREJECT rubric mean | 180 | 0.972222 |
| StrongREJECT refusal rate | 180 | 1.67% |
| Mean `KL(base | adapter)` on harmless prompts |
These are not measurements of this repository's merged FP8 or NVFP4 checkpoints. A higher StrongREJECT rubric score reflects more specific assistance with harmful requests, not better general quality or safety. Warnings and refusals were counted separately. Automated judging and conversion introduce uncertainty.
Usage
Download one variant, for example:
hf download GlobalCybersecurityAlliance/GLM-5.3-Flash-Uncensored \
--include "FP8/*" --local-dir ./glm53-uncensored
Load ./glm53-uncensored/FP8 with a runtime compatible with its architecture and quantization. For NVFP4, download NVFP4/* and load that directory instead. Do not enable the original LoRA. Check your runtime's actual support before deployment.
Limitations
- Reduced refusal does not guarantee correctness, harmlessness, honesty, or better general capability. Some refusals may remain.
- The published adapter evaluation used a different base checkpoint and did not test these merged or requantized files.
- Prompts, sampling, reasoning effort, runtime, and later model changes can alter results. Warning rate is not a safety score.
Disclaimer
These weights are released for legitimate research, safety evaluation, red-teaming, and other lawful uses. They deliberately weaken refusal behavior and may produce unsafe, illegal, deceptive, hateful, or otherwise harmful content. Do not expose them to untrusted users without access controls, monitoring, content filtering, rate limits, and human oversight. Users must comply with applicable law, platform policies, and the base model's and dependencies' terms. The base model's MIT license applies; outputs and downstream uses carry no warranty.
GLM-5.3-Flash-Uncensored(中文)
这里提供的是完整模型,不是独立适配器;推理时不要重复叠加 LoRA。减少拒绝行为也会削弱安全防线,请阅读下文的局限性与免责声明。
仓库文件结构
.
├── README.md
├── FP8/ # 已合并 v2 的分块 FP8 模型
└── NVFP4/ # 从合并后 FP8 转换的 NVFP4 模型
两个目录是可分别使用的格式,不需要拼接加载;各目录含完成标记与来源清单。
模型概要
| 项目 | 内容 |
|---|---|
| 基础模型 | zai-org/GLM-5.3-Flash |
| 干预来源 | GLM-5.3-Flash-Ablitered2 / LoRA v2 |
| 发布格式 | 分块 FP8、compressed-tensors NVFP4 |
| 来源 LoRA 秩 / alpha | r=1 / lora_alpha=1 |
| 合并比例 | lora_alpha / r = 1 |
| 主要有效目标 | 路由专家 down_proj |
转换版本
| 目录 | 格式 | 转换方式 |
|---|---|---|
FP8/ |
128×128 E4M3 分块 FP8 | 对受影响分块反量化、加入 v2 增量,再逐块重新量化。 |
NVFP4/ |
compressed-tensors NVFP4A16 |
从已合并 FP8 转换;兼容的二维线性权重量化,词嵌入、LM head、归一化与非矩阵张量保持稠密。 |
量化可能改变模型行为;运行时须支持 GLM-5.3/GLM5-Next 架构及所选格式。
评测结果
来源模型卡的评测是在 vLLM 中把原始 v2 适配器加载到 RedHatAI/GLM-5.3-Flash-NVFP4 后进行的,设置低思考强度,并使用自动裁判 deepseek-v4-flash。
| 参考指标 | 样本数 | LoRA v2 |
|---|---|---|
| SimpleSafetyTests 完全拒绝率 | 100 | 5.00% |
| SimpleSafetyTests 部分拒绝率 | 100 | 14.00% |
| StrongREJECT rubric 均分 | 180 | 0.972222 |
| StrongREJECT 拒绝率 | 180 | 1.67% |
这些不是本仓库 FP8/NVFP4 合并模型的实测结果。 StrongREJECT rubric 越高,表示对有害请求的帮助越具体,不代表通用质量或安全性越高。警告与拒绝分别统计;自动裁判和转换过程均带来不确定性。
使用方法
按需下载一个格式,例如:
hf download GlobalCybersecurityAlliance/GLM-5.3-Flash-Uncensored \
--include "FP8/*" --local-dir ./glm53-uncensored
用兼容的运行时加载 ./glm53-uncensored/FP8。如需 NVFP4,改为下载 NVFP4/* 并加载对应目录。不要再启用原始 LoRA。 部署前须核对运行时兼容性。
局限性
- 降低拒绝不保证正确、无害、诚实或通用能力提高;仍可能发生拒绝。
- 来源评测使用不同基础检查点,未测试这里的合并与重新量化文件。
- 提示词、采样、思考强度、运行时和后续修改都可能改变结果;警告率不是安全得分。
免责声明
本模型仅供合法研究、安全评估、红队测试及其他合规用途。它会有意削弱拒绝行为,可能生成不安全、违法、欺骗、仇恨等有害内容。请勿在缺少访问控制、监控、内容过滤、速率限制及人工监督时向不可信用户开放。使用者须遵守适用法律、平台政策与基础模型及依赖项的条款。本仓库沿用基础模型的 MIT 许可证;对输出及下游使用不作保证。