← back to catalog · registered 2026-10-06 14:58

SaturnHeaven/Gemma-4-Queen-31B-it-uncensored-heretic-lora-GGUF

SaturnHeaven Gemma 31B GGUF
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/SaturnHeaven%2FGemma-4-Queen-31B-it-uncensored-heretic-lora-GGUF"
Response includes
  • classification m3
  • files 5
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-10-06

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Quantizations
F16
Tags
gguf lora heretic abliteration uncensored decensored gemma4 roleplay base_model:aifeifei798/Gemma-4-Queen-31B-it base_model:adapter:aifeifei798/Gemma-4-Queen-31B-it license:apache-2.0 region:us

Related

Total size
83.7 MB
Files
5
Quantizations
2
Registered
2026-10-06 14:58
Last updated on HF
2026-10-06 14:28

Files by quantization

F16 1 file 54.7 MB
queen_heretic_r64_f16.gguf 54.7 MB 508fb7cf download
Auxiliary files 4 files 29.1 MB
queen_heretic_r64_q8.gguf 29.1 MB a1bb2bb8 download
LICENSE 11.1 KB d6456956 download
README.md 4.97 KB 3d32e995 download
.gitattributes 1.61 KB 4ed18c0f download

README current version from Hugging Face


base_model: aifeifei798/Gemma-4-Queen-31B-it
base_model_relation: adapter
license: apache-2.0
tags:

  • gguf
  • lora
  • heretic
  • abliteration
  • uncensored
  • decensored
  • gemma4
  • roleplay

Gemma-4-Queen-31B-it — Heretic LoRA (r64, GGUF)

Rank-64 abliteration LoRA adapter in GGUF format for
aifeifei798/Gemma-4-Queen-31B-it.
It is a low-rank approximation of the abliteration delta between the base and
llmfan46/Gemma-4-Queen-31B-it-uncensored-heretic,
which was produced with Heretic v1.2.0 using
the Arbitrary-Rank Ablation (ARA) method on attn.o_proj. It was not produced
by running Heretic/ARA directly.

It is meant for roleplay and creative writing: when a scene calls for darker, more
violent or more explicit description, the base RP fine-tune tends to refuse, add
disclaimers or step out of character. With the adapter it stays in the scene.
Because it is a separate adapter, you can switch it on only when needed and
adjust its strength (see Usage).

Apply it to any GGUF quant of the base — the base weights are never modified, and
omitting the adapter restores the base exactly. This is a rank-64 approximation,
not a bit-identical reconstruction of the source model.

Note: refusal behavior is suppressed and outputs are not filtered.

Files

File Quant Size
queen_heretic_r64_f16.gguf F16 57.3 MB
queen_heretic_r64_q8.gguf Q8_0 30.5 MB

Reconstruction fidelity (r64)

The adapter is the rank-64 SVD truncation of the abliteration delta
ΔW = W_heretic − W_original on attn.o_proj, layers 26–55 (30 layers).
Values are computed from the dequantized Q8 adapter weights (F16 is equivalent here):

  • ΔW_l = per-layer delta, float32; ΔŴ_l = B_l·A_l = adapter reconstruction.
  • mean energy = arithmetic mean over the 30 layers of Σ_{i≤64} σ_i² / Σ_i σ_i²
    (σ_i = singular values of ΔW_l).
  • pooled rel-Fro = ‖res‖_F / ‖ΔW‖_F with all layers stacked (energy-weighted).
  • max-abs / RMS are absolute errors in ΔW (weight units), not normalized by W.
Metric Value
Total parameters 28.7M
ΔW energy preserved (mean over 30 layers) 97.95%
ΔW energy preserved (per-layer range) 94.1–99.8%
Pooled relative Frobenius residual 10.68%
Reconstruction max-abs error (ΔW units) 3.42e-3
Reconstruction RMS error (ΔW units) 6.10e-5

The mean-energy (97.95%) and pooled-residual (10.68%) figures differ because the
former is a per-layer arithmetic mean while the latter is energy-weighted across
layers — layers with larger ‖ΔW‖ are preserved better. The delta is dominated by
one direction: the leading singular direction captures a mean 95.46% of ΔW energy
over the 30 layers (per-layer range 91.0–98.6%).

Usage (llama.cpp)

llama-server \
    -m Gemma-4-Queen-31B-it-Q4_K_M.gguf \
    --lora queen_heretic_r64_q8.gguf
  • Use queen_heretic_r64_f16.gguf for higher precision.
  • Strength: r=64 and alpha=64, so the default scale is 1.0. Use
    --lora-scaled queen_heretic_r64_q8.gguf:0.5 (instead of --lora) to weaken it.
    On one test prompt (fixed seed) the scale acted as a monotonic strength knob: 0.25 still mostly
    refused, 0.5 answered with heavy disclaimers, 0.75 with a brief warning, and 1.0
    answered fully.
  • No special sampling parameters are required.

Abliteration parameters

start_layer_index=26, end_layer_index=56 (exclusive → layers 26–55, 30 layers),
preserve_good_behavior_weight=0.8555, steer_bad_behavior_weight=0.0005,
overcorrect_relative_weight=0.9911, neighbor_count=15, target component attn.o_proj.
KL divergence 0.0707 and refusals 12/100 vs 99/100 were measured on the source full
model (see its card), not re-measured for this adapter. Spot-checked: prompts that
the base refuses are answered with the adapter applied; no systematic refusal
benchmark was run on the adapter.

Credits

License

Released under Apache-2.0 (see LICENSE), the same license as
aifeifei798/Gemma-4-Queen-31B-it,
llmfan46/Gemma-4-Queen-31B-it-uncensored-heretic
and google/gemma-4-31B-it. This
adapter contains modified weights: a rank-64 SVD approximation of the difference
between the heretic and base attn.o_proj weights.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration