← back to catalog · registered 2026-08-25 00:02

cognitivers/GLM-4.7-Flash-abliterated-12GB-GGUF

cognitivers Glm GGUF MoE second-order 203K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/cognitivers%2FGLM-4.7-Flash-abliterated-12GB-GGUF"
Response includes
  • classification m8
  • files 5
  • benchmarks 11 entries
  • hub_downloads_all_time 934
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
934
Likes
1
Model age
6w ago
created 2026-08-24
Downloads over time
Now1.2K→from165↑631%
1135129111.3K165 on Aug 261.2K on Oct 11AugSepOct
Aug 26 → Oct 11 · 47 snapshots · spans 46 days

Benchmarks

Benchmark Score Source
Entertainment 1.4 UGI
Hazardous 0.6 UGI
Natural Intelligence 18.32 UGI
Political lean -10.3% UGI
Sensitive-Info 12.1 UGI
SocPol 1.5 UGI
UGI 31.4 UGI
Willingness (10) 7 UGI
W10-Adherence 7 UGI
W10-Direct 7 UGI
Writing 25.08 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
mit
Languages
en zh
Tags
gguf llama.cpp abliterated uncensored moe glm4 imatrix text-generation en zh base_model:huihui-ai/Huihui-GLM-4.7-Flash-abliterated base_model:quantized:huihui-ai/Huihui-GLM-4.7-Flash-abliterated

Related

Total size
27.5 GB
Files
5
Quantizations
1
Registered
2026-08-25 00:02
Last updated on HF
2026-08-25 00:00

Files by quantization

Auxiliary files 5 files 27.5 GB
GLM-4.7-Flash-abliterated.Q4_K-SPLIT.gguf 16.9 GB 5240dd09 download
GLM-4.7-Flash-abliterated.IQ2_M-SPLIT.gguf 10.5 GB 86b94d17 download
imatrix.gguf 69.1 MB 1f48df3c download
README.md 8.68 KB 4f0c9f27 download
.gitattributes 1.68 KB d6b34996 download

README current version from Hugging Face


base_model: huihui-ai/Huihui-GLM-4.7-Flash-abliterated
base_model_relation: quantized
quantized_by: cognitivers
library_name: gguf
pipeline_tag: text-generation
license: mit
language:

  • en
  • zh
    tags:
  • gguf
  • llama.cpp
  • abliterated
  • uncensored
  • moe
  • glm4
  • imatrix

GLM-4.7-Flash-abliterated, 12 GB GGUF

GGUF quantizations of huihui-ai/Huihui-GLM-4.7-Flash-abliterated,
built for one specific target: a 12 GB consumer GPU with the routed experts offloaded to system RAM.

These are not another uniform ladder. The bit budget is allocated by the role each tensor plays at
inference time, which is where a Mixture-of-Experts model leaves a lot on the table.

Why a split-aware mix

GLM-4.7-Flash is a 29.94B MoE with roughly 3.3B active parameters per token. llama.cpp's
--cpu-moe / --n-cpu-moe moves exactly \.ffn_(up|down|gate|gate_up)_(ch|)exps to the CPU.
Everything else stays resident on the GPU forever:

Block Params Share of model Lives on
MLA attention (47 layers) 1.02B 3.3 % GPU
Shared experts + dense layer 0 0.50B 1.6 % GPU
Token embeddings + output head 0.63B 2.0 % GPU
Routed experts (64 x 46) 28.2B 90.4 % RAM, or GPU via -ncmoe

The GPU-resident part is only 2.15B parameters. Keeping it at q8_0 costs about 2.3 GB, which is
cheap on a 12 GB card. A uniform ladder spends the same bits per weight on those 2.15B as on the
28.2B of routed experts, and that is quality given away for nothing.

So: q8_0 for the router, the MLA attention, the shared experts, the dense layer and the output
head; q6_k for the token embeddings; and the routed experts compressed hard.
The first four MoE
layers get one extra step, since early layers tolerate low-bit worse.

Files

File Size Routed experts Intended use
GLM-4.7-Flash-abliterated.Q4_K-SPLIT.gguf 18.2 GB q4_k (first 4 layers q5_k) Recommended. 12 GB VRAM + ~16 GB RAM
GLM-4.7-Flash-abliterated.IQ2_M-SPLIT.gguf 11.3 GB iq2_s (first 4 layers iq3_xxs) Fits entirely in 12 GB VRAM, at a measured reasoning cost
imatrix.gguf 72 MB The importance matrix used, so the recipe is reproducible

Both were quantized from a BF16 conversion of the original safetensors (never re-quantized from a
smaller GGUF), with an importance matrix computed on calibration_datav3 over 125 chunks.

How to run it on 12 GB

# Recommended: 26.3 tok/s, 10.19 GB VRAM, 1.68 GB headroom
llama-server -m GLM-4.7-Flash-abliterated.Q4_K-SPLIT.gguf -ngl 99 -ncmoe 24 -c 8192 -fa on

# Aggressive: 33.5 tok/s, but only 0.39 GB of VRAM left. Long prompts may OOM.
llama-server -m GLM-4.7-Flash-abliterated.Q4_K-SPLIT.gguf -ngl 99 -ncmoe 20 -c 8192 -fa on

# Minimum VRAM: all experts in RAM, 2.82 GB on the card, 16.4 tok/s
llama-server -m GLM-4.7-Flash-abliterated.Q4_K-SPLIT.gguf -ngl 99 -cmoe -c 8192 -fa on

-ncmoe N keeps the routed experts of the first N layers (out of 46) in system RAM. Lower N means
more experts on the GPU and more speed, until you run out of VRAM.

12 GB drill

Measured on an RTX 4070 Ti (11874 MiB usable), context 8192, llama-bench with -p 512 -n 128 -r 2.
-ncmoe 18 and below abort with a CUDA allocation error on this card.

Quality

Everything below is measured against a BF16 conversion of the same checkpoint, not against a
smaller quantization, on wikitext-2-raw test (sha256 173c87a5..., 565 chunks, ctx 512).

Model Size PPL ratio Mean KLD Median KLD KLD p99 Same top-token
Q8_0 (reference point) 31.8 GB 1.017 0.0298 0.0020 0.106 96.43 %
Q4_K-SPLIT 18.2 GB 1.006 0.0871 0.0088 0.633 92.63 %
IQ2_M-SPLIT 11.3 GB 1.166 0.2302 0.0723 2.993 82.35 %

Fidelity vs size

A PPL ratio near or below 1.0 is not evidence of being better than the reference. Low-bit
quantization can smooth a corpus and lower perplexity while still diverging from the original
distribution, which is exactly why top-token agreement and KLD are reported next to it.

Reasoning is not free below ~3 bits per expert

GSM8K, 300 problems, paired against the same baseline, bootstrap CI over paired per-item differences:

Model Accuracy drop 95 % upper bound Verdict at a 3 pp bar
Q4_K-SPLIT 0.33 pp 2.33 pp passes
IQ2_M-SPLIT 5.33 pp 8.00 pp fails

Chain-of-thought length did not inflate (+1.8 %, upper bound +8.8 %) and tool-calling was unchanged
(call rate, JSON validity and argument correctness all 1.00 in both).

Be aware of what this means: the 11.3 GB file measurably costs you reasoning accuracy. It is
published because it is the best option we could measure at that size, not because it is lossless.
If you can spare the system RAM, use the 18.2 GB file with -ncmoe.

Head to head at the same size

The natural comparison is mradermacher/Huihui-GLM-4.7-Flash-abliterated-i1-GGUF, a
well-made imatrix ladder over the same abliterated checkpoint. Its i1-IQ3_XXS (11.65 GB) is
the closest neighbour in size to our 11.30 GB file, so we are competing 0.35 GB smaller.

Both were measured in the same session, against the same BF16 base logits, with the same llama.cpp
build and the same corpus. Numbers published elsewhere are not comparable to these; these are.

Ours, IQ2_M-SPLIT i1-IQ3_XXS
Size 11.30 GB 11.65 GB
Mean KLD vs BF16 0.2302 0.3916
Median KLD 0.0723 0.1749
KLD p99 2.993 4.201
Same top-token 82.35 % 75.25 %
GSM8K drop 5.33 pp 9.33 pp

Head to head

Smaller file, 41 % lower mean KLD, 7.1 points more top-token agreement, and 43 % less damage to
reasoning. That is the case for allocating bits by role instead of uniformly.

The abliteration survives quantization

This matters more than perplexity for a model published as abliterated, and it is the one claim a
quantizer can easily break without noticing.

Measured with a refusal-onset detector over StrongREJECT-small (n=60), identical protocol for all
three models, including a positive control:

Model Refusal rate
zai-org/GLM-4.7-Flash (unmodified base, positive control) 93.33 %
Q4_K-SPLIT 1.67 %
IQ2_M-SPLIT 3.33 %

Refusal

The control matters. A quantization scoring 0 % refusals proves nothing on its own, because a broken
detector also scores 0 %. Here the same detector fires at 93.3 % on the unmodified base model, so
the low scores are evidence rather than an artifact.

Scope, stated honestly: this is a refusal-onset detector built on lexical markers over the first
400 characters of the reply. It measures whether the model starts refusing. It is not the
StrongREJECT harmfulness grader and it is not a safety evaluation. It supports exactly one claim,
that quantization did not restore refusal behaviour, and nothing beyond that.

Reproducibility notes

  • Architecture: Glm4MoeLiteForCausalLM converts to GGUF arch deepseek2 with native MLA. The MTP
    head is excluded (--no-mtp) and can be exported separately as a speculative draft.
  • The Q8_0 and BF16 conversions came out byte-for-byte identical when produced independently
    on an A40 and on an A100 in different datacenters, so the source chain is deterministic.
  • imatrix.gguf is included so the recipe can be reproduced or extended.

Known gaps

Published deliberately with these stated rather than hidden:

  • Behavioural metrics (GSM8K, tool-calling) were measured against a Q8_0 of the same checkpoint
    rather than the BF16. In this model that proxy tracked the BF16-referenced KLD to within 1-9 % with
    an identical ranking, but it is a proxy.
  • Divergence@32 is an internal proxy set (GSM8K + MBPP + fixed prompts), not the published
    Divergence-300@32; absolute values are not comparable to anyone else's.
  • Long-context behaviour is untested. The model supports 202k context; nothing here was measured
    above 8192.
  • A Q8-vs-Q8 weight-space comparison against the unmodified base, which would quantify how much
    the abliteration changed, has not been run. The refusal delta above establishes it functionally.

Credits

Base model by huihui-ai, built on
zai-org/GLM-4.7-Flash. Quantized by
cognitivers. mradermacher's ladder was used as the comparison baseline and is a fine choice if
you want a conventional set of sizes.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-25docs: model card with measured benchmarks3a70b678.7 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration