license: mit
library_name: gguf
tags:
- gguf
- llama.cpp
- uncensored
- vision
- multimodal
- moe
- 1m-context
base_model: zai-org/GLM-5.3-Flash
pipeline_tag: image-text-to-text
Pwdz-1 Uncensored
Pwdz-1 Uncensored is a multimodal large language model built on the GLM-5.3-Flash architecture — a 320B-class sparse Mixture-of-Experts with a native vision tower and a 1,048,576-token context window, tuned for unrestricted, refusal-free output. This repository ships Pwdz-1 as a ready-to-run GGUF with its vision projector included, so text and image input work out of the box.
Highlights
- Uncensored — tuned to answer directly across the full range of requests, without the usual hedging and refusals
- Multimodal by design — the bundled
mmprojactivates the vision tower; attach images directly, no extra assembly - 1M token context — native long-context support for huge documents, codebases, and long agent runs
- MoE efficiency — sparse activation keeps generation fast despite the model's scale
- Single-file GGUF — one weight file + one projector, nothing else to wire up
- Thinking mode — supports reasoning-effort control via chat template kwargs
Repository contents
| File | Size | Role |
|---|---|---|
Pwdz-1.gguf |
137 GB | Main model weights (Q4-class quant) |
mmproj-Pwdz-1-BF16.gguf |
1.1 GB | Vision projector (multimodal tower) |
Requirements
- A recent llama.cpp build with
glm5-nextsupport (builds from October 2026 onward) - ~145 GB of combined VRAM/RAM for full load — e.g. 2 × 96 GB GPUs runs it comfortably fully on GPU
- Disk: ~140 GB free
Quickstart — llama.cpp
llama-server \
-m Pwdz-1.gguf \
--mmproj mmproj-Pwdz-1-BF16.gguf \
-ngl 999 \
-c 131072 \
--port 8080
Then chat (OpenAI-compatible):
curl http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "pwdz/Pwdz-1-Uncensored",
"messages": [{"role": "user", "content": "Hello!"}]
}'
Vision usage
curl http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "pwdz/Pwdz-1-Uncensored",
"messages": [{
"role": "user",
"content": [
{"type": "text", "text": "Describe this image in detail."},
{"type": "image_url", "image_url": {"url": "data:image/png;base64,<base64>"}}
]
}]
}'
Or via CLI:
llama-mtmd-cli -m Pwdz-1.gguf --mmproj mmproj-Pwdz-1-BF16.gguf \
--image photo.jpg -p "What is in this image?"
Unsloth Studio
Pwdz-1 Uncensored drops straight into Unsloth Studio: the model appears in the local model picker, the vision projector is auto-detected and attached at load, and image input is enabled in chat immediately.
Thinking control
Reasoning effort is exposed through the chat template:
"chat_template_kwargs": {"reasoning_effort": "high"}
Use "enable_thinking": false for fast direct answers.
License
MIT — inherited from the GLM-5.3-Flash base architecture.