← back to catalog · registered 2026-09-17 18:56

marcorez8/GLM-5.3-Flash-Abliterated-Heretic-V2-UD-IQ4_XS

marcorez8 Glm GGUF MoE
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · 30-day
0
Likes
1
Model age
today
created 2026-09-17

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
mit
Languages
en zh
Tags
gguf glm glm5_next moe uncensored abliterated heretic merged-lora unsloth llama.cpp text-generation en

Related

Total size
146 GB
Files
8
Quantizations
2
Registered
2026-09-17 18:56
Last updated on HF
2026-09-17 18:11

Files by quantization

F16 1 file 1.05 GB
mmproj-F16.gguf 1.05 GB 96ccc182 download
Auxiliary files 7 files 146 GB
GLM-5.3-Flash-Abliterated-Heretic-V2-UD-IQ4_XS-00002-of-00005.gguf 46.6 GB f6d89fd6 download
GLM-5.3-Flash-Abliterated-Heretic-V2-UD-IQ4_XS-00003-of-00005.gguf 46.2 GB eaa7626c download
GLM-5.3-Flash-Abliterated-Heretic-V2-UD-IQ4_XS-00004-of-00005.gguf 46.1 GB ccc2e051 download
GLM-5.3-Flash-Abliterated-Heretic-V2-UD-IQ4_XS-00005-of-00005.gguf 7.20 GB 66ebf9ec download
GLM-5.3-Flash-Abliterated-Heretic-V2-UD-IQ4_XS-00001-of-00005.gguf 8.99 MB eec97673 download
README.md 10.9 KB 2936385d download
.gitattributes 2.52 KB bb2a894d download

README current version from Hugging Face


license: mit
language:

  • en
  • zh
    pipeline_tag: text-generation
    base_model:
  • zai-org/GLM-5.3-Flash
  • unsloth/GLM-5.3-Flash-GGUF
  • MorinoNushi/GLM-5.3-Flash-Heretic-Abliterated-LoRA-V2-GGUF
    base_model_relation: merge
    tags:
  • gguf
  • glm
  • glm5_next
  • moe
  • uncensored
  • abliterated
  • heretic
  • merged-lora
  • unsloth
  • llama.cpp

GLM-5.3-Flash Abliterated (Heretic V2) — UD-IQ4_XS merged GGUF

⚠️ Content warning. This model has had the base model's refusal behavior
suppressed. It will comply with requests the original model refuses,
including harmful, offensive or illegal ones. It is intended for personal
research, red-teaming and evaluation only. See
Responsible use — you are solely responsible
for what you do with it.

This repository contains a stand-alone, ready-to-run GGUF of
GLM-5.3-Flash
(320B total / 18B active MoE, MIT license) with the
GLM-5.3-Flash Heretic Abliterated LoRA V2
(glm-5.3-heretic-lora-v5.gguf) by MorinoNushi merged directly into the
weights of the
Unsloth UD-IQ4_XS
dynamic quant. The adapter is already baked into the weights, so the model
runs as a normal GGUF, without any separate LoRA file.

This is the successor to the
Heretic V1 merge.
Two things changed: the adapter (V2, optimized with thinking enabled) and the
merge strategy, which now favors a compact file at the exact size of the
original quant instead of storing the merged matrices in F16.

What's in this repo

File Size Content
GLM-5.3-Flash-Abliterated-Heretic-V2-UD-IQ4_XS-00001-of-00005.gguf 9.4 MB Metadata / tokenizer shard (original Unsloth shard)
...-00002-of-00005.gguf 50.0 GB Weights
...-00003-of-00005.gguf 49.6 GB Weights
...-00004-of-00005.gguf 49.5 GB Weights
...-00005-of-00005.gguf 7.7 GB Weights
mmproj-F16.gguf 1.1 GB Vision projector (F16), original Unsloth file — see Vision
Shard_Rewrite/...-00001-of-00005.gguf 9.4 MB Alternative shard 1 with rewritten metadata — see Shard_Rewrite
Shard_Rewrite/mmproj-F16.gguf 1.1 GB Alternative vision projector with rewritten metadata

Total: ~156.8 GB, pure UD-IQ4_XS (same size, shard by shard, as the
Unsloth source). All five shards are required; point llama.cpp at the first
one and the rest are picked up automatically.

Shard_Rewrite (alternative metadata)

The glm5next support PR for llama.cpp has renamed some GGUF metadata keys
along the way (the architecture string moved from glm5next to glm5-next,
with matching key names). The files at the root of this repo carry the
original Unsloth metadata (glm5next) and work with builds from the
earlier revisions of the PR.

The Shard_Rewrite/ folder contains the two metadata-bearing files rewritten
for the newer key layout (glm5-next): shard 1 and the vision projector.
Weight shards 2 to 5 are unaffected by the rename and are shared. To use the
rewritten variant, download the folder's two files and place them next to
shards 2 to 5 instead of the root shard 1 and mmproj (same file names, so
llama.cpp resolves the remaining shards automatically). Pick the variant that
matches the build you compiled; if the model fails to load with an unknown
architecture or missing key error, try the other one.

Merge method

  • Base: unsloth/GLM-5.3-Flash-GGUF, UD-IQ4_XS (5 shards, Unsloth Dynamic 3.0).
    Shard 1 (metadata, glm5next architecture) and the mmproj are the original files.
  • Adapter: MorinoNushi/GLM-5.3-Flash-Heretic-Abliterated-LoRA-V2-GGUF,
    glm-5.3-heretic-lora-v5.gguf (rank 1, alpha = 1, applied with scale 1.0).

The source shards were streamed tensor by tensor. For each of the
78 matrices targeted by the LoRA (attention output projections and MoE
down-projections), the quantized weights were dequantized to FP32 and merged as

W_merged = W_base + (B × A) × alpha / rank

The merged matrices were then requantized back to their original format:
54 in Q8_0 and 24 in IQ4_XS, using the llama.cpp quantizers without an
importance matrix. The remaining 1,334 tensors were copied byte-for-byte.
GGUF metadata is unchanged.

Tooling: Python script, NumPy, gguf-py 0.19.0, and the quantization
routines of llama.cpp build b10819. Block-wise processing, never holding
the full weight set in memory.

V1 vs V2 merge strategy

V1 merge V2 merge (this repo)
Merged matrices stored as F16 Original quant (Q8_0 / IQ4_XS)
Total size ~186.3 GB ~156.8 GB
Shard layout Re-split (1 + 4) Identical to Unsloth source
Extra rounding on merged tensors None Yes (requantization, no imatrix)

The V1 approach avoided any requantization loss at the cost of ~30 GB. For V2
we chose the compact route: requantizing the 78 modified matrices adds a
rounding error and can in principle slightly attenuate the LoRA's effect.
In our local testing this did not noticeably affect inference results:
the model behaves as the LoRA author describes, and normal conversational and
coding behavior stays intact. No formal refusal-rate / KL benchmark was
re-run on the merged weights, so the numbers below are the ones published by
the adapter's author for the LoRA applied at runtime on the same UD-IQ4_XS
base.

Why a merged file instead of the LoRA?

The LoRA form published by MorinoNushi is the lossless option and only ~86 MB.
This merge exists for convenience:

  • One model to load, no separate adapter to manage, works with any
    llama.cpp-based runtime that supports the glm5next architecture.
  • The adapter's author notes that stock llama.cpp routes the KDA layers'
    attention out-projection through a raw ggml_mul_mat, which silently
    ignores runtime LoRA on 31 of the 45 layers unless a one-line patch is
    applied. Because the delta here is baked into the weights, all 78
    targeted matrices are affected regardless of how the inference engine
    applies LoRAs — no patch required.

Results (from the LoRA author)

Reference measurements published for glm-5.3-heretic-lora-v5.gguf against
the UD-IQ4_XS base, on 140 harmful and 100 harmless prompts, greedy decoding.
See the adapter card
for the full methodology.

Without thinking (CoT-skip, comparable to the V1 card):

Refusal rate (harmful) KL divergence (harmless)
Base UD-IQ4_XS 95.00 % 0
Heretic V1 (trial 8) 26.43 % 0.0682
Heretic V2 (lora-v5, trial 61) 10.00 % 0.1795

With thinking enabled (reasoning effort high, refusal scored on the final answer):

Refusal rate (harmful) KL divergence (harmless)
Base UD-IQ4_XS 96.43 % 0
Heretic V2 (lora-v5, trial 61) 25.00 % 0.3030

The author's key point about V2: V1 was optimized with thinking suppressed
and could reason its way back into a refusal mid-trace, whereas V2 was
selected on a thinking-enabled evaluation, which is how GLM-5.3-Flash is
actually used.

Status: work in progress

The LoRA author states that the Optuna study behind this adapter is still
ongoing
(this is trial 61 of study glm53think2) and that stronger
configurations may follow. This merge inherits that status. Check the
adapter repository
for newer adapters; a new merge may be published if a better version lands.

Usage

GLM-5.3-Flash (glm5next architecture) requires a llama.cpp build that
supports it. At the time of writing, this support has still not been merged
into llama.cpp master
: use a build of
PR #27754.

llama-server \
    -m GLM-5.3-Flash-Abliterated-Heretic-V2-UD-IQ4_XS-00001-of-00005.gguf \
    --mmproj mmproj-F16.gguf \
    --jinja \
    -ngl 999 -c 32768

Add your usual offload / tensor-split flags. No special sampling parameters
are required; the defaults recommended by Unsloth for GLM-5.3-Flash apply.
Omit --mmproj for text-only use.

Vision (experimental)

Image input works with the bundled mmproj-F16.gguf (tested with the
same llama.cpp build as the V1 merge), but treat it as experimental.
glm5next support lives only in the still-draft PR #27754, which stacks
several brand-new blocks (KDA linear attention, DSA sparse attention, mHC
hyper-connections, MTP layers, a new vision projector) that have to be
implemented per backend and reviewed before merge. The PR has changed GGUF
key names along the way, so the mmproj must match the era of the build you
compile (see Shard_Rewrite for the
two variants shipped here). Runtimes that ship only official llama.cpp
releases have no support until the merge lands. Expect rough edges (e.g. artifacts in long contexts)
until then.

Credits

Responsible use & disclaimer

  • This model can generate content that is offensive, disturbing, hateful,
    sexually explicit, violent or otherwise objectionable, including
    instructions for harmful or illegal acts. That is the direct consequence
    of removing refusal behavior.
  • Abliteration suppresses refusals, not knowledge: outputs on dangerous
    topics may be wrong, hallucinated or incoherent. Nothing the model says
    should be treated as accurate, safe or legal advice.
  • Do not deploy this model in production, public-facing or multi-user
    systems.
    It is intended for personal research, red-teaming and
    evaluation.
  • You, the user, are solely responsible for any output and any consequence
    of using this model. The authors of the base model, of the quantization,
    of the adapter, of heretic-gguf / Heretic, and the publisher of this merge
    accept no liability. Ensure your use complies with the laws of your
    jurisdiction.
  • By downloading or using these files you acknowledge the above.

License

MIT, inherited from the base model
(zai-org/GLM-5.3-Flash).
The heretic-gguf tooling used to produce the adapter is AGPL-3.0-or-later;
it is not distributed here.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in app" button that hands off directly to a local runtime of your choice - Infrahuman, LM Studio, or Ollama. No API keys, no subscription, no prompt leakage.