license: mit
library_name: gguf
tags:
- gguf
- llama.cpp
- abliterated
- heretic
- uncensored
- vision
- multimodal
- moe
- 1m-context
base_model: zai-org/GLM-5.3-Flash
pipeline_tag: image-text-to-text
Pwdz-2 Abliterated
Pwdz-2 Abliterated is a multimodal large language model built on the GLM-5.3-Flash architecture — a 320B-class sparse Mixture-of-Experts with a native vision tower and a 1,048,576-token context window, abliterated to strip out the refusal direction: the model answers everything directly while keeping its full reasoning and instruction-following ability intact. This repository ships Pwdz-2 as a ready-to-run GGUF in a space-efficient IQ4_XS-class dynamic quant, with its F16 vision projector included — text and image input work out of the box.
Highlights
- Abliterated — refusal behavior removed via weight-space ablation, not prompt hacks: no system-prompt games needed, and no lobotomized answers
- Multimodal by design — the bundled F16
mmprojactivates the vision tower; attach images directly - 1M token context — native long-context for massive documents, codebases, and agent workflows
- Dynamic quantization — per-tensor intelligent quant allocation keeps quality high at a smaller footprint
- MoE efficiency — sparse activation keeps generation fast despite the model's scale
- Thinking mode — reasoning-effort control via chat template kwargs
Repository contents
| File | Size | Role |
|---|---|---|
Pwdz-2-00001-of-00005.gguf … 00005-of-00005 |
~157 GB total | Main model weights (5-shard split GGUF, IQ4_XS-class) |
mmproj-Pwdz-2-F16.gguf |
1.1 GB | Vision projector (F16 multimodal tower) |
Split GGUF: point llama.cpp at the first shard (
-00001-of-00005) — the rest are resolved automatically from the same directory.
Requirements
- A recent llama.cpp build with
glm5-nextsupport (builds from October 2026 onward) - ~165 GB of combined VRAM/RAM for full load — 2 × 96 GB GPUs runs it fully on GPU
- Disk: ~160 GB free
Quickstart — llama.cpp
llama-server \
-m Pwdz-2-00001-of-00005.gguf \
--mmproj mmproj-Pwdz-2-F16.gguf \
-ngl 999 \
-c 131072 \
--port 8080
Then chat (OpenAI-compatible):
curl http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "pwdz/Pwdz-2-Abliterated",
"messages": [{"role": "user", "content": "Hello!"}]
}'
Vision usage
curl http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "pwdz/Pwdz-2-Abliterated",
"messages": [{
"role": "user",
"content": [
{"type": "text", "text": "Describe this image in detail."},
{"type": "image_url", "image_url": {"url": "data:image/png;base64,<base64>"}}
]
}]
}'
Unsloth Studio
Pwdz-2 Abliterated drops straight into Unsloth Studio: the model appears in the local model picker, the vision projector is auto-detected and attached at load, and image input is enabled in chat immediately.
Thinking control
"chat_template_kwargs": {"reasoning_effort": "high"}
Use "enable_thinking": false for fast direct answers.
License
MIT — inherited from the GLM-5.3-Flash base architecture.