license: mit
base_model: zai-org/GLM-4.7-Flash
tags:
- gguf
- uncensored
- moe
- ollama
GLM-4.7-Flash-Uncensored-HauhauCS-Balanced — AI-conn fork
This is a fork, not our model. All credit for the uncensored build goes to
HauhauCS,
and for the base model to Zhipu AI (zai-org/GLM-4.7-Flash).
This fork exists so the fleet's copy cannot be altered or pulled upstream,
and to carry the operational fix notes below. License: MIT (per the source repo).
Recommended file for a 20 GB card: GLM-4.7-Flash-Uncensored-HauhauCS-Balanced-Q4_K_M.gguf (16.89 GiB).
Ollama pull:
ollama pull hf.co/AI-conn/GLM-4.7-Flash-Uncensored-HauhauCS-Balanced:Q4_K_M
Known bug + fix — read this before benchmarking
Symptom: On Ollama versions before 0.32.1, this model decodes at
dense-model speed (~13.7 tok/s measured) even though it is a 30B-total /
~3B-active MoE. A sibling 3B-active model (Qwen3-Coder-30B) does ~56.7 tok/s
on the same machine under the same Ollama.
Root cause: Pre-0.32.x Ollama loaded glm4moelite GGUFs without mapping
them onto its working DeepSeek2/MLA MoE path, so the routed experts were
CPU-offloaded and the model executed densely. Ollama 0.32.1 translatesglm4moelite GGUFs to DeepSeek2 (MLA) conventions at load (server log:handle_glm4moelite … translating to deepseek2 (MLA conventions)); the same
machine measured 13.7 → 77.8 tok/s after upgrading 0.24.0 → 0.32.1.
Fix: run Ollama ≥ 0.32.1. No model-side change is needed — this build's
GGUF already declares the deepseek2 architecture (29.94B params).
Second trap — stock context: the model's default context is 202,752
tokens. At full context, independent measurements show ~51 s prefills and
decode collapsing to ~10 tok/s (vs ~35 tok/s at ~43K context). Mitigate in
your Modelfile (a ready one ships in this repo asModelfile.glm-4.7-flash-prime):
PARAMETER num_ctx 32768
PARAMETER num_batch 2048
num_batch 2048 also matters for prefill: a 25K-token prompt measured 665 s
at Ollama's default batch vs 33 s at 2048.
Related, still open upstream: ollama/ollama#14045 documents an AMD-side
cousin — on a Radeon 7900 XTX, one Ollama RC double-enumerated the GPU via
Vulkan (phantom second device), bloating the CPU compute graph. Ifollama ps shows a CPU/GPU split on an AMD card, check your Ollama version
and device enumeration before blaming the model.
Evidence: ollama-herd project trace doc (measured, third-party);
ollama/ollama#14045 (GitHub API, still open at 2026-10-11); practitioner
benchmarks mid-2026 (~53–57 tok/s via Ollama, corroborating).
Prime fit (measured budget, RX 7900 XT — 19.98 GiB VRAM, 32 GB RAM)
- Q4_K_M weights: 16.89 GiB → fully GPU-resident with ~3 GiB left for KV
cache + compute at the tuned 32,768-token context. - Do not run this at stock context on a 20 GB card (see trap above).
- Acceptance bar used by our fleet: tuned decode ≥ 45 tok/s and
ollama ps
showing 100% GPU.