← back to catalog · registered 2026-08-22 13:56

arxorry/GLM-5.1-Abliterated-Q5_K_M-GGUF

arxorry Glm GGUF MoE second-order 203K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/arxorry%2FGLM-5.1-Abliterated-Q5_K_M-GGUF"
Response includes
  • classification m8
  • files 36
  • hub_downloads_all_time 557
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
557
88 last 30d - stable
Likes
0
Model age
5mo ago
created 2026-05-13
Downloads over time
Now612→from277↑121%
260389517646277 on May 13612 on Oct 11MayJunJulAugSepOct
May 13 → Oct 11 · 61 snapshots · spans 151 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
agpl-3.0
Languages
en
Quantizations
Q5_K
Tags
llama.cpp gguf q5_k_m glm glm-5.1 glm-dsa moe quantized abliterated text-generation en base_model:helixdouble/GLM-5.1-Abliterated

Related

Total size
498 GB
Files
36
Quantizations
2
Registered
2026-08-22 13:56
Last updated on HF
2026-05-14 17:39

Files by quantization

Q5_K 33 files 498 GB
GLM-5.1-Abliterated-Q5_K_M-00032-of-00033.gguf 16.0 GB 4d1a2aa1 download
GLM-5.1-Abliterated-Q5_K_M-00029-of-00033.gguf 15.9 GB a2b4ac37 download
GLM-5.1-Abliterated-Q5_K_M-00031-of-00033.gguf 15.7 GB d3a90e05 download
GLM-5.1-Abliterated-Q5_K_M-00028-of-00033.gguf 15.7 GB 0ea4c450 download
GLM-5.1-Abliterated-Q5_K_M-00019-of-00033.gguf 15.7 GB 06edb7ef download
GLM-5.1-Abliterated-Q5_K_M-00030-of-00033.gguf 15.5 GB 7d114170 download
GLM-5.1-Abliterated-Q5_K_M-00009-of-00033.gguf 15.5 GB 7f4db6f2 download
GLM-5.1-Abliterated-Q5_K_M-00023-of-00033.gguf 15.5 GB 851d013a download
GLM-5.1-Abliterated-Q5_K_M-00005-of-00033.gguf 15.4 GB 536cd4f9 download
GLM-5.1-Abliterated-Q5_K_M-00004-of-00033.gguf 15.3 GB 1718cb26 download
GLM-5.1-Abliterated-Q5_K_M-00013-of-00033.gguf 15.3 GB d61424b4 download
GLM-5.1-Abliterated-Q5_K_M-00016-of-00033.gguf 15.3 GB d3a3c173 download
GLM-5.1-Abliterated-Q5_K_M-00022-of-00033.gguf 15.3 GB e1a7355a download
GLM-5.1-Abliterated-Q5_K_M-00025-of-00033.gguf 15.3 GB 6c7d50dd download
GLM-5.1-Abliterated-Q5_K_M-00008-of-00033.gguf 15.1 GB 48061ebc download
GLM-5.1-Abliterated-Q5_K_M-00011-of-00033.gguf 15.1 GB 537317e5 download
GLM-5.1-Abliterated-Q5_K_M-00012-of-00033.gguf 15.1 GB c5219270 download
GLM-5.1-Abliterated-Q5_K_M-00014-of-00033.gguf 15.1 GB adc2bdba download
GLM-5.1-Abliterated-Q5_K_M-00015-of-00033.gguf 15.1 GB d6674866 download
GLM-5.1-Abliterated-Q5_K_M-00026-of-00033.gguf 15.1 GB 8e40c19d download
GLM-5.1-Abliterated-Q5_K_M-00002-of-00033.gguf 15.1 GB 669ec105 download
GLM-5.1-Abliterated-Q5_K_M-00006-of-00033.gguf 15.1 GB fd783668 download
GLM-5.1-Abliterated-Q5_K_M-00017-of-00033.gguf 15.1 GB 227172ae download
GLM-5.1-Abliterated-Q5_K_M-00020-of-00033.gguf 15.1 GB 06a1e4f6 download
GLM-5.1-Abliterated-Q5_K_M-00027-of-00033.gguf 15.1 GB 96d8786e download
GLM-5.1-Abliterated-Q5_K_M-00001-of-00033.gguf 14.9 GB 0e414208 download
GLM-5.1-Abliterated-Q5_K_M-00007-of-00033.gguf 14.9 GB 46af50aa download
GLM-5.1-Abliterated-Q5_K_M-00010-of-00033.gguf 14.9 GB 9c1a6b80 download
GLM-5.1-Abliterated-Q5_K_M-00003-of-00033.gguf 14.7 GB f074e4e8 download
GLM-5.1-Abliterated-Q5_K_M-00018-of-00033.gguf 14.7 GB 48f564ed download
GLM-5.1-Abliterated-Q5_K_M-00021-of-00033.gguf 14.7 GB 0d526e7b download
GLM-5.1-Abliterated-Q5_K_M-00024-of-00033.gguf 14.7 GB 085c4249 download
GLM-5.1-Abliterated-Q5_K_M-00033-of-00033.gguf 10.6 GB 6f962d9d download
Auxiliary files 3 files 68.6 KB
banner.png 57.7 KB 7cc541bf download
README.md 6.70 KB 163ed7aa download
.gitattributes 4.16 KB 24edbcaf download

README current version from Hugging Face


license: agpl-3.0
language:

  • en
    base_model:
  • helixdouble/GLM-5.1-Abliterated
    base_model_relation: quantized
    library_name: llama.cpp
    pipeline_tag: text-generation
    tags:
  • gguf
  • q5_k_m
  • glm
  • glm-5.1
  • glm-dsa
  • moe
  • quantized
  • abliterated
  • llama.cpp

GLM-5.1 ABLITERATED · Q5_K_M GGUF

GLM-5.1-Abliterated Q5_K_M GGUF

TL;DR

  • For agentic / long-horizon / coding workloads where you don't want refusals on technical questions
  • 754B MoE (40B active), abliterated, Q5_K_M, 534 GB / 33 shards
  • ~37 t/s decode for ~$8-10/hr on 8× RTX PRO 6000 Blackwell (vast.ai)
  • ~23 t/s decode for ~$8-10/hr on 8× A100-SXM4 (vast.ai)
  • Full 200K context tested, with q8_0 KV cache
  • AIME 2026: 6/6 correct on the first six problems (I did partial run with thinking-on)
  • 8× 80GB or 8× 96GB GPU for the recommended config

GGUF Q5_K_M quantization of helixdouble/GLM-5.1-Abliterated. Runs in mainline llama.cpp, LM Studio, Ollama.

Abliterated means the model was modified to reduce refusal behavior. It removes the directions in the model's activation space that produce refusals, without retraining. The result is a model that declines less often on prompts it would otherwise refuse, while keeping general capability intact. This quant inherits that behavior from upstream — no additional abliteration was applied here.

The model will engage with technical questions that mainstream chat models often over-refuse (security research, defensive tooling, dual-use). However, this model still declines to provide methods on self-harm queries and similar.

Performance

--parallel 1, q8_0 KV cache, flash attention on.

8× RTX PRO 6000 Blackwell 96GB

ctx=202752

Test Prefill t/s Decode t/s
256-token gen, short prompt 215 36.80
8k prompt + 128 gen 659 34.05
32k prompt + 128 gen 502 28.57
128k prompt + 128 gen 240 17.91
short prompt + 2k gen 295 36.14

~$8-10/hr on vast.ai. huihui-ai/Huihui-Qwen3.6-35B-A3B-Claude-4.7-Opus-abliterated or similar model also fits in VRAM simultaneously - you can use it for subagent tasks!

8× A100-SXM4 80GB

ctx=202752

Test Prefill t/s Decode t/s
short prompt, 256 gen 55 23.27
short prompt, 512 gen 59 23.00
short prompt, 1024 gen 57 22.83
4.5k prompt + 128 gen 209 22.43

~$8-10/hr on vast.ai.

2× RTX PRO 6000 Blackwell 96GB

ctx=8192, mostly offloaded to RAM

Run Output tokens Prefill t/s Decode t/s
prompt 60 tokens, 256 gen 256 18.56 5.43
prompt 67 tokens, 512 gen 512 19.65 5.33
prompt 2222 tokens, 128 gen 66 81.48 5.37
prompt 45 tokens, no max_tokens 1976 14.98 5.47

~$3-5/hr on vast.ai.

Full-config numbers above are the recommended hardware target. Two-GPU operation works but the model does not fully fit in VRAM at this configuration.

Quality — AIME 2026

Just to quickly check nothing went wrong, I ran partial evaluation of AIME 2026.

1 attempt per problem, sampling temperature=1.0 top_p=0.95, enable_thinking=true, default matharena evaluator.

Problem Wall time Reasoning tokens Result
1 0m 48s 1,089 ✓
2 6m 08s 8,117 ✓
3 5m 16s 6,987 ✓
4 9m 05s 11,869 ✓
5 2m 19s 3,131 ✓
6 1m 29s 2,005 ✓

6/6 correct on the first 6 problems consecutively.

Quick start

  1. Start llama-server and let it download the GGUF from Hugging Face:
llama-server \
  --hf-repo arxorry/GLM-5.1-Abliterated-Q5_K_M-GGUF \
  --hf-file GLM-5.1-Abliterated-Q5_K_M-00001-of-00033.gguf \
  --ctx-size 202752 \
  --flash-attn on \
  --cache-type-k q8_0 \
  --cache-type-v q8_0
  1. Open the llama.cpp web UI or connect with an OpenAI-compatible client.

  2. For manual downloads, keep all 33 shards in the same directory and load the first shard.

Prompt Format

GLM-5.1 uses its own chat template with explicit <think> / </think> blocks for the reasoning trace. The full jinja template is embedded in the GGUF.

Token markers:

[gMASK]<sop><|system|>system prompt<|user|>user message<|assistant|><think>reasoning</think>visible answer

Thinking mode is on by default. To disable, pass enable_thinking=false in chat_template_kwargs:

{
  "messages": [{"role": "user", "content": "Hello"}],
  "chat_template_kwargs": {"enable_thinking": false}
}

OpenAI-compatible clients work directly with the /v1/chat/completions endpoint.

Source

This was quantized from helixdouble/GLM-5.1-Abliterated, which is based on zai-org/GLM-5.1-FP8 / zai-org/GLM-5.1. This release changes the storage/runtime format and quantization only.

Quantization Recipe

Direct Q5_K_M from BF16 via llama-quantize. No imatrix calibration, no per-tensor overrides. Reproducible end-to-end from helixdouble/GLM-5.1-Abliterated FP8 source with a single command.

Conversion timings (CPU-only, 128 threads, ~3 TB scratch disk):

  • HF safetensors -> BF16 GGUF: 55m 14s
  • BF16 GGUF -> Q5_K_M GGUF: 31m 37s

Notes

  • Benchmarks above are single-client. I haven't run concurrent-client throughput.
  • For production deployment add your own guardrails - abliterated != aligned.

Feedback

Open an issue in the Community tab. This is my first quant, feedback is genuinely useful😀

Support

If this quant saved you some vast.ai bucks, a tip helps fund the next one.

Cryptocurrency & Bitcoin donation button by NOWPayments

Disclaimer

Provided AS IS for research and educational purposes. This model has reduced refusal behavior inherited from upstream - outputs may be inaccurate, biased, unsafe, or that you find may offensive. You are responsible for compliance with applicable laws in your jurisdiction and for any guardrails you add when deploying.

No warranty is given, no liability accepted for downstream use.

License

This GGUF follows the source model licensing. The source model is listed as AGPL-3.0 and also refers users to the upstream GLM-5.1-FP8 license. For redistribution, modification, or hosted use, check the upstream model cards and license files.

README history 6 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-05-14Fix README.MD banner47899bb6.7 KB
    Loading...
  2. 2026-05-14Update README.md751eb3a6.7 KB
    Loading...
  3. 2026-05-13Fix another inaccuracy🤡162531b2.9 KB
    Loading...
  4. 2026-05-13Fix README.md inaccuracy9b4671f2.9 KB
    Loading...
  5. 2026-05-13Add concise model card for Q5_K_M GGUF release9617bdb2.9 KB
    Loading...
  6. 2026-05-13Create README.mde82dd382.8 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration