← back to catalog · registered 2026-10-11 20:59

AI-conn/GLM-4.7-Flash-Uncensored-HauhauCS-Balanced

AI-conn Glm GGUF MoE
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/AI-conn%2FGLM-4.7-Flash-Uncensored-HauhauCS-Balanced"
Response includes
  • classification m-uncensored
  • files 7
  • author_summary 3 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-10-11

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
mit
Quantizations
F16 Q4_K Q6_K Q8_0
Tags
gguf uncensored moe ollama base_model:zai-org/GLM-4.7-Flash base_model:quantized:zai-org/GLM-4.7-Flash license:mit endpoints_compatible region:us conversational

Related

Total size
125 GB
Files
7
Quantizations
5
Registered
2026-10-11 20:59
Last updated on HF
2026-10-11 20:21

Files by quantization

F16 1 file 55.8 GB
GLM-4.7-Flash-Uncensored-HauhauCS-Balanced-FP16.gguf 55.8 GB 064533e4 download
Q8_0 1 file 29.7 GB
GLM-4.7-Flash-Uncensored-HauhauCS-Balanced-Q8_0.gguf 29.7 GB ffb16bb2 download
Q6_K 1 file 22.9 GB
GLM-4.7-Flash-Uncensored-HauhauCS-Balanced-Q6_K.gguf 22.9 GB d397ce6d download
Q4_K 1 file 16.9 GB
GLM-4.7-Flash-Uncensored-HauhauCS-Balanced-Q4_K_M.gguf 16.9 GB 8c1ac6d1 download
Auxiliary files 3 files 5.83 KB
README.md 3.05 KB edb332b2 download
.gitattributes 1.83 KB bdd34fec download
Modelfile.glm-4.7-flash-prime 974 B da02ac1f download

README current version from Hugging Face


license: mit
base_model: zai-org/GLM-4.7-Flash
tags:

  • gguf
  • uncensored
  • moe
  • ollama

GLM-4.7-Flash-Uncensored-HauhauCS-Balanced — AI-conn fork

This is a fork, not our model. All credit for the uncensored build goes to
HauhauCS,
and for the base model to Zhipu AI (zai-org/GLM-4.7-Flash).
This fork exists so the fleet's copy cannot be altered or pulled upstream,
and to carry the operational fix notes below. License: MIT (per the source repo).

Recommended file for a 20 GB card: GLM-4.7-Flash-Uncensored-HauhauCS-Balanced-Q4_K_M.gguf (16.89 GiB).

Ollama pull:

ollama pull hf.co/AI-conn/GLM-4.7-Flash-Uncensored-HauhauCS-Balanced:Q4_K_M

Known bug + fix — read this before benchmarking

Symptom: On Ollama versions before 0.32.1, this model decodes at
dense-model speed (~13.7 tok/s measured) even though it is a 30B-total /
~3B-active MoE. A sibling 3B-active model (Qwen3-Coder-30B) does ~56.7 tok/s
on the same machine under the same Ollama.

Root cause: Pre-0.32.x Ollama loaded glm4moelite GGUFs without mapping
them onto its working DeepSeek2/MLA MoE path, so the routed experts were
CPU-offloaded and the model executed densely. Ollama 0.32.1 translates
glm4moelite GGUFs to DeepSeek2 (MLA) conventions at load (server log:
handle_glm4moelite … translating to deepseek2 (MLA conventions)); the same
machine measured 13.7 → 77.8 tok/s after upgrading 0.24.0 → 0.32.1.

Fix: run Ollama ≥ 0.32.1. No model-side change is needed — this build's
GGUF already declares the deepseek2 architecture (29.94B params).

Second trap — stock context: the model's default context is 202,752
tokens
. At full context, independent measurements show ~51 s prefills and
decode collapsing to ~10 tok/s (vs ~35 tok/s at ~43K context). Mitigate in
your Modelfile (a ready one ships in this repo as
Modelfile.glm-4.7-flash-prime):

PARAMETER num_ctx 32768
PARAMETER num_batch 2048

num_batch 2048 also matters for prefill: a 25K-token prompt measured 665 s
at Ollama's default batch vs 33 s at 2048.

Related, still open upstream: ollama/ollama#14045 documents an AMD-side
cousin — on a Radeon 7900 XTX, one Ollama RC double-enumerated the GPU via
Vulkan (phantom second device), bloating the CPU compute graph. If
ollama ps shows a CPU/GPU split on an AMD card, check your Ollama version
and device enumeration before blaming the model.

Evidence: ollama-herd project trace doc (measured, third-party);
ollama/ollama#14045 (GitHub API, still open at 2026-10-11); practitioner
benchmarks mid-2026 (~53–57 tok/s via Ollama, corroborating).

Prime fit (measured budget, RX 7900 XT — 19.98 GiB VRAM, 32 GB RAM)

  • Q4_K_M weights: 16.89 GiB → fully GPU-resident with ~3 GiB left for KV
    cache + compute at the tuned 32,768-token context.
  • Do not run this at stock context on a 20 GB card (see trap above).
  • Acceptance bar used by our fleet: tuned decode ≥ 45 tok/s and ollama ps
    showing 100% GPU.
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration