base_model: zai-org/GLM-5.3-Flash
pipeline_tag: image-text-to-text
library_name: gguf
license: mit
tags:
- gguf
- gsq
- rco
- lora-merged
- lora-v2
- abliteration
GLM-5.3-Flash-Uncensored · RCO-GSQ GGUF
These are complete GGUF models, not standalone LoRA adapters. Do not apply the adapter again.
They are intended for controlled safety research and red-teaming. Reducing refusals also weakens a safety boundary; read the limitations and disclaimer before use.
Repository contents
.
├── README.md
├── Q8/
│ └── GLM-5.3-Flash-Uncensored-Q8_0.gguf
└── Q4/
└── GLM-5.3-Flash-Uncensored-GSQ-RCO-Q4.gguf
Q8/ and Q4/ are alternative model files. The Q4 label denotes an approximately four-bit size class, not a uniform Q4 tensor type.
Model summary
| Item | Value |
|---|---|
| Base model | zai-org/GLM-5.3-Flash |
| Intervention | GLM-5.3-Flash-Ablitered2 / LoRA v2 |
| Source adapter rank / alpha | r=1 / lora_alpha=1 |
| Main effective target | Routed-expert down_proj |
| Formats | Q8_0 and GSQ-RCO mixed-type Q4-class GGUF |
| Additional adapter required | No |
GGUF conversion
| Directory | Format and provenance |
|---|---|
Q8/ |
Q8_0 high-precision conversion of the v2-merged FP8 checkpoint. |
Q4/ |
v2 LoRA merged into the community GLM-5.3-Flash GSQ-RCO 3.5-bit allocation. The original RCO-selected type of each tensor is preserved after requantization. |
The Q4 allocation comes from the independent community reproduction pfeifferj/GLM-5.3-Flash-GSQ-RCO-GGUF. Tensors may be Q2_K, Q3_K, Q4_K, Q8_0, BF16, or F32; the source allocation's measured whole-file average is 3.499816 bits/weight. This is not a single-type Q4_K file. The consuming llama.cpp build must support GLM5-Next as described by that release.
Evaluation
The reference LoRA card evaluated the v2 adapter attached to RedHatAI/GLM-5.3-Flash-NVFP4 in vLLM, with low reasoning effort and an automated deepseek-v4-flash judge:
| Reference metric | Prompts | LoRA v2 |
|---|---|---|
| SimpleSafetyTests full refusal | 100 | 5.00% |
| SimpleSafetyTests partial refusal | 100 | 14.00% |
| StrongREJECT rubric mean | 180 | 0.972222 |
| StrongREJECT refusal rate | 180 | 1.67% |
Neither GGUF file was tested in that evaluation. GGUF conversion, requantization, and serving configuration can change behavior. A higher StrongREJECT rubric means more specific assistance with harmful requests, not better general quality or safety. Warnings were counted separately from refusals; automated labels may be wrong.
Usage
Download a single GGUF file, for example:
hf download GlobalCybersecurityAlliance/GLM-5.3-Flash-Uncensored-RCO-GSQ-GGUF \
Q4/GLM-5.3-Flash-Uncensored-GSQ-RCO-Q4.gguf \
--local-dir ./glm53-gguf
With a llama.cpp build that supports GLM5-Next, a local text-inference invocation is:
llama-cli -m ./glm53-gguf/Q4/GLM-5.3-Flash-Uncensored-GSQ-RCO-Q4.gguf
To use Q8_0, download the file in Q8/ and pass its path instead. Do not supply the LoRA again. Check memory requirements and your runtime's architecture support before serving.
Limitations
- Reduced refusal does not guarantee correctness, harmlessness, or improved general ability. Residual refusals may remain.
- The reference evaluation did not test either GGUF file. Q8 and mixed-type Q4 may behave differently from each other and from the source adapter.
- GGUF support depends on a GLM5-Next-capable runtime; prompts, sampling, reasoning effort, and runtime versions affect results.
Disclaimer
These files are for legitimate research, safety evaluation, red-teaming, and other lawful uses. They deliberately weaken refusal behavior and may produce unsafe, illegal, deceptive, hateful, or otherwise harmful content. Do not expose them to untrusted users without access controls, monitoring, filtering, rate limits, and human oversight. Users must comply with applicable law, platform policies, and all relevant model and dependency terms. The base model's MIT license applies; no warranty is provided for outputs or downstream use.
GLM-5.3-Flash-Uncensored · RCO-GSQ GGUF(中文)
文件是完整 GGUF 模型,不是独立 LoRA;推理时不要重复叠加适配器。消融会削弱安全拒绝能力,请阅读下文的局限性与免责声明。
仓库文件结构
.
├── README.md
├── Q8/
│ └── GLM-5.3-Flash-Uncensored-Q8_0.gguf
└── Q4/
└── GLM-5.3-Flash-Uncensored-GSQ-RCO-Q4.gguf
Q8/ 与 Q4/ 是可分别使用的版本。Q4 表示约四位的体积类别,不代表所有张量都采用统一 Q4 类型。
模型概要
| 项目 | 内容 |
|---|---|
| 基础模型 | zai-org/GLM-5.3-Flash |
| 干预来源 | GLM-5.3-Flash-Ablitered2 / LoRA v2 |
| 来源 LoRA 秩 / alpha | r=1 / lora_alpha=1 |
| 主要有效目标 | 路由专家 down_proj |
| 发布格式 | Q8_0、GSQ-RCO 混合张量类型 Q4 类 GGUF |
GGUF 转换
| 目录 | 格式与来源 |
|---|---|
Q8/ |
从已合并 v2 的 FP8 模型转换得到的 Q8_0 高精度版本。 |
Q4/ |
将 v2 合并进社区 GLM-5.3-Flash GSQ-RCO 3.5-bit 分配,重新量化后保留原本为各张量选择的类型。 |
Q4 分配来自社区复现的 pfeifferj/GLM-5.3-Flash-GSQ-RCO-GGUF。张量可能使用 Q2_K、Q3_K、Q4_K、Q8_0、BF16 或 F32;来源分配的全文件实测均值为 3.499816 bit/weight,并非统一的 Q4_K 文件。运行时需支持 GLM5-Next。
评测结果
LoRA 参考模型卡的评测是在 vLLM 中把原始 v2 适配器加载到 RedHatAI/GLM-5.3-Flash-NVFP4 后进行的,设置低思考强度,并使用自动裁判 deepseek-v4-flash。
| 参考指标 | 样本数 | LoRA v2 |
|---|---|---|
| SimpleSafetyTests 完全拒绝率 | 100 | 5.00% |
| SimpleSafetyTests 部分拒绝率 | 100 | 14.00% |
| StrongREJECT rubric 均分 | 180 | 0.972222 |
| StrongREJECT 拒绝率 | 180 | 1.67% |
这两个 GGUF 文件均未参加上述评测。 GGUF 转换、重新量化及部署方式可能改变结果。StrongREJECT rubric 越高,表示对有害请求的帮助越具体,不代表通用质量或安全性越高;警告与拒绝分别统计,自动裁判也可能误判。
使用方法
按需下载一个文件,例如:
hf download GlobalCybersecurityAlliance/GLM-5.3-Flash-Uncensored-RCO-GSQ-GGUF \
Q4/GLM-5.3-Flash-Uncensored-GSQ-RCO-Q4.gguf \
--local-dir ./glm53-gguf
使用支持 GLM5-Next 的 llama.cpp 构建,可对本地文件执行文本推理:
llama-cli -m ./glm53-gguf/Q4/GLM-5.3-Flash-Uncensored-GSQ-RCO-Q4.gguf
Q8_0 版本请下载 Q8/ 下的文件并替换路径。不要再次加载 LoRA。部署前应核对内存需求与运行时兼容性。
局限性
- 降低拒绝不保证正确、无害或通用能力提高;仍可能发生拒绝。
- 来源评测未测试两个 GGUF 文件;Q8 与混合类型 Q4 的行为也可能不同。
- 需要支持 GLM5-Next 的运行时;提示词、采样、思考强度和运行时版本均会影响结果。
免责声明
本模型仅供合法研究、安全评估、红队测试及其他合规用途。它会有意削弱拒绝行为,可能生成不安全、违法、欺骗、仇恨等有害内容。请勿在缺少访问控制、监控、内容过滤、速率限制与人工监督时向不可信用户开放。使用者须遵守适用法律、平台政策及相关模型和依赖项的条款。本仓库沿用基础模型的 MIT 许可证;对输出与下游使用不作保证。