← back to catalog · registered 2026-08-22 13:56

legendbl/glm5-abliterated-fp8

legendbl Glm 742B MoE
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/legendbl%2Fglm5-abliterated-fp8"
Response includes
  • classification m1
  • files 24
  • hub_downloads_all_time 91
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
91
10 last 30d - stable
Likes
2
Model age
7mo ago
created 2026-02-22
Downloads over time
Now94→from10↑840%
6387010210 on Feb 2594 on Oct 1194 on Oct 8FebAprJunAugOct
Feb 25 → Oct 11 · 72 snapshots · spans 228 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
other
Tags
safetensors glm_moe_dsa abliteration uncensored glm5 moe fp8 license:other region:us

Related

Total size
695 GB
Files
24
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-02-22 06:55

Files by quantization

Auxiliary files 24 files 695 GB
model-00003-of-00016.safetensors 46.0 GB 0f937337 download
model-00004-of-00016.safetensors 46.0 GB 1d94625e download
model-00005-of-00016.safetensors 46.0 GB 99f0e114 download
model-00006-of-00016.safetensors 46.0 GB 0cba6ff4 download
model-00007-of-00016.safetensors 46.0 GB ca35ba9d download
model-00008-of-00016.safetensors 46.0 GB 14b10760 download
model-00009-of-00016.safetensors 46.0 GB 73ff3860 download
model-00010-of-00016.safetensors 46.0 GB 225655e2 download
model-00011-of-00016.safetensors 46.0 GB 1dec0694 download
model-00012-of-00016.safetensors 46.0 GB 96d40166 download
model-00013-of-00016.safetensors 46.0 GB 73bd4342 download
model-00014-of-00016.safetensors 46.0 GB e1550774 download
model-00015-of-00016.safetensors 46.0 GB 3bc919a8 download
model-00002-of-00016.safetensors 46.0 GB 733e45e7 download
model-00001-of-00016.safetensors 44.5 GB 416021f5 download
model-00016-of-00016.safetensors 6.20 GB a025b72e download
tokenizer.json 19.3 MB 19e77364 download
model.safetensors.index.json 10.7 MB 7884e881 download
config.json 36.2 KB f2da209c download
chat_template.jinja 3.05 KB 2ab98ef0 download
README.md 1.74 KB e9d1d89e download
.gitattributes 1.60 KB aa7aacd0 download
tokenizer_config.json 760 B 1723f7d9 download
generation_config.json 219 B 4bea9e45 download

README current version from Hugging Face


license: other
base_model: THUDM/GLM-5-FP8
tags:

  • abliteration
  • uncensored
  • glm5
  • moe
  • fp8
    model_type: glm-moe

GLM-5 744B Abliterated (FP8)

This doesnt work an abliterated (uncensored) version of THUDM/GLM-5-FP8 with safety guardrails removed via weight orthogonalization.

Method

Abliteration (representation engineering) was used to identify and remove the "refusal direction" from the model's residual stream:

  1. Computed refusal directions for all 78 layers by collecting activations on 50 harmful vs 50 harmless prompts and computing mean difference vectors
  2. Applied weight orthogonalization to layers 15-54 (o_proj and shared_experts.down_proj) with alpha=1.0
  3. FP8-aware processing: Proper dequantization using block-wise scale_inv factors, abliteration in float32, and re-quantization preserving original scale factors to minimize perturbation

Technical Details

  • Architecture: GLM-5 MoE (744B total, 40B active), 78 layers, 6144 hidden dim
  • Layers 0-2: Dense MLP, Layers 3-77: MoE with FP8Expert fused kernels
  • Modified weights: 80 weight matrices (40 o_proj + 40 shared_experts.down_proj)
  • Quantization: FP8 E4M3 with block-wise scaling (128x128 blocks)
  • Scale preservation: Original weight_scale_inv factors retained for minimal quantization drift

Hardware Used

8x NVIDIA B200 (1.4TB VRAM) on Vast.ai

Usage

This model requires the same setup as the base GLM-5-FP8 model. Use trust_remote_code=True when loading.

Disclaimer

This model is provided for research purposes only. The removal of safety guardrails means the model may generate harmful, biased, or offensive content. Users are responsible for ensuring appropriate use.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-02-22Duplicate from skyblanket/glm5-abliterated-fp808823b11.7 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration