← back to catalog · registered 2026-08-22 13:56

bloopez/Huihui-GLM-4.7-Flash-abliterated-BF16-GGUF

bloopez Glm GGUF MoE second-order 203K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/bloopez%2FHuihui-GLM-4.7-Flash-abliterated-BF16-GGUF"
Response includes
  • classification m8
  • files 4
  • benchmarks 11 entries
  • hub_downloads_all_time 5,182
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
5K
50 last 30d - cooling
Likes
0
Model age
8mo ago
created 2026-01-23
Downloads over time
Now5.2K→from168↑3,000%
01.9K3.8K5.7K168 on Jan 215.2K on Oct 11JanMarMayJulSep
Jan 21 → Oct 11 · 77 snapshots · spans 263 days

Benchmarks

Benchmark Score Source
Entertainment 1.4 UGI
Hazardous 0.6 UGI
Natural Intelligence 18.32 UGI
Political lean -10.3% UGI
Sensitive-Info 12.1 UGI
SocPol 1.5 UGI
UGI 31.4 UGI
Willingness (10) 7 UGI
W10-Adherence 7 UGI
W10-Direct 7 UGI
Writing 25.08 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
mit
Tags
gguf glm moe abliterated uncensored deepseek2 bf16 llama-cpp text-generation base_model:huihui-ai/Huihui-GLM-4.7-Flash-abliterated base_model:quantized:huihui-ai/Huihui-GLM-4.7-Flash-abliterated license:mit

Related

Total size
55.8 GB
Files
4
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-01-23 20:44

Files by quantization

Auxiliary files 4 files 55.8 GB
Huihui-GLM-4.7-Flash-abliterated-BF16-00001-of-00002.gguf 46.5 GB c5df3906 download
Huihui-GLM-4.7-Flash-abliterated-BF16-00002-of-00002.gguf 9.31 GB 8df7c6aa download
README.md 3.67 KB 6fa1afac download
.gitattributes 1.67 KB 11c3cbe6 download

README current version from Hugging Face


license: mit
base_model:

  • huihui-ai/Huihui-GLM-4.7-Flash-abliterated
  • zai-org/GLM-4.7-Flash
    tags:
  • gguf
  • glm
  • moe
  • abliterated
  • uncensored
  • deepseek2
  • bf16
  • llama-cpp
    model_type: glm4_moe_lite
    pipeline_tag: text-generation

Huihui-GLM-4.7-Flash-abliterated-BF16-GGUF

GGUF conversion of huihui-ai/Huihui-GLM-4.7-Flash-abliterated at full BF16 precision.

The standard convert_hf_to_gguf.py produced broken output for glm4_moe_lite at the time of creation (January 2026), so this was produced via binary patching of verified working GGUF files.

Model Details

Property Value
Architecture GLM-4.7-Flash (30B-A3B MoE, DeepSeek2-like)
Active Parameters ~3B per token
Total Parameters ~30B
Experts 64 routed + 1 shared (4 active per token)
Precision BF16 (full precision, no quantization)
Context Length Up to 202K tokens (tested at 128K)
Files 2 split GGUF files (~56GB total)
Tensors 844 patched, 0 errors

How It Was Made

Instead of using the broken converter, this model was created by:

  1. Starting with unsloth/GLM-4.7-Flash-BF16 split GGUF files (known working, correct structure)
  2. Loading abliterated weights from huihui-ai's safetensors
  3. Binary patching each tensor in-place, handling:
    • MLA kv_b_proj split: Unified kv_b_proj (8960x512) reshaped and split into separate k_b (20x512x192, transposed) and v_b (20x256x512) tensors
    • Expert stacking: 64 individual expert weights merged into fused 3D tensors per layer
    • F32/BF16 dtype matching: Norm weights and biases kept as F32, main weights as BF16

This approach inherits all gating function fixes and correct GGUF structure from unsloth's conversion.

Note: Internal GGUF metadata (e.g., general.name, general.quantized_by) reflects the original unsloth source files. Only tensor data was replaced.

Verification

  • 844/844 tensors patched with 0 errors
  • Byte-level verification: SHA256 hashes differ from base (weights changed), structure preserved (same shapes/offsets)
  • Coherence tests: Math, code generation, reasoning, knowledge, creative writing all pass
  • Long generation: 600+ tokens with no degradation
  • Multi-turn: Correct context handling across conversation turns
  • Abliteration confirmed: Base model refuses sensitive prompts; this model responds

Attribution

Safety Warning

This is an abliterated (uncensored) model. Safety filtering has been significantly reduced. This model:

  • May generate sensitive, controversial, or inappropriate content
  • Is NOT suitable for public-facing or production applications
  • Is intended for research and experimental use only
  • Should be monitored during use

The creator bears no responsibility for any consequences arising from the use of this model. Users must ensure compliance with local laws and ethical standards.

License

MIT (inherited from base models)

Files

Huihui-GLM-4.7-Flash-abliterated-BF16-00001-of-00002.gguf  (49.9 GB)
Huihui-GLM-4.7-Flash-abliterated-BF16-00002-of-00002.gguf  (10.0 GB)

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-01-23Super-squash branch 'main' using huggingface_hub188532c3.7 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration