base_model: aifeifei798/Gemma-4-Queen-31B-it
base_model_relation: adapter
license: apache-2.0
tags:
- gguf
- lora
- heretic
- abliteration
- uncensored
- decensored
- gemma4
- roleplay
Gemma-4-Queen-31B-it — Heretic LoRA (r64, GGUF)
Rank-64 abliteration LoRA adapter in GGUF format for
aifeifei798/Gemma-4-Queen-31B-it.
It is a low-rank approximation of the abliteration delta between the base and
llmfan46/Gemma-4-Queen-31B-it-uncensored-heretic,
which was produced with Heretic v1.2.0 using
the Arbitrary-Rank Ablation (ARA) method on attn.o_proj. It was not produced
by running Heretic/ARA directly.
It is meant for roleplay and creative writing: when a scene calls for darker, more
violent or more explicit description, the base RP fine-tune tends to refuse, add
disclaimers or step out of character. With the adapter it stays in the scene.
Because it is a separate adapter, you can switch it on only when needed and
adjust its strength (see Usage).
Apply it to any GGUF quant of the base — the base weights are never modified, and
omitting the adapter restores the base exactly. This is a rank-64 approximation,
not a bit-identical reconstruction of the source model.
Note: refusal behavior is suppressed and outputs are not filtered.
Files
| File | Quant | Size |
|---|---|---|
queen_heretic_r64_f16.gguf |
F16 | 57.3 MB |
queen_heretic_r64_q8.gguf |
Q8_0 | 30.5 MB |
Reconstruction fidelity (r64)
The adapter is the rank-64 SVD truncation of the abliteration deltaΔW = W_heretic − W_original on attn.o_proj, layers 26–55 (30 layers).
Values are computed from the dequantized Q8 adapter weights (F16 is equivalent here):
ΔW_l= per-layer delta, float32;ΔŴ_l = B_l·A_l= adapter reconstruction.- mean energy = arithmetic mean over the 30 layers of
Σ_{i≤64} σ_i² / Σ_i σ_i²
(σ_i = singular values ofΔW_l). - pooled rel-Fro =
‖res‖_F / ‖ΔW‖_Fwith all layers stacked (energy-weighted). - max-abs / RMS are absolute errors in ΔW (weight units), not normalized by W.
| Metric | Value |
|---|---|
| Total parameters | 28.7M |
| ΔW energy preserved (mean over 30 layers) | 97.95% |
| ΔW energy preserved (per-layer range) | 94.1–99.8% |
| Pooled relative Frobenius residual | 10.68% |
| Reconstruction max-abs error (ΔW units) | 3.42e-3 |
| Reconstruction RMS error (ΔW units) | 6.10e-5 |
The mean-energy (97.95%) and pooled-residual (10.68%) figures differ because the
former is a per-layer arithmetic mean while the latter is energy-weighted across
layers — layers with larger ‖ΔW‖ are preserved better. The delta is dominated by
one direction: the leading singular direction captures a mean 95.46% of ΔW energy
over the 30 layers (per-layer range 91.0–98.6%).
Usage (llama.cpp)
llama-server \
-m Gemma-4-Queen-31B-it-Q4_K_M.gguf \
--lora queen_heretic_r64_q8.gguf
- Use
queen_heretic_r64_f16.gguffor higher precision. - Strength:
r=64andalpha=64, so the default scale is 1.0. Use--lora-scaled queen_heretic_r64_q8.gguf:0.5(instead of--lora) to weaken it.
On one test prompt (fixed seed) the scale acted as a monotonic strength knob: 0.25 still mostly
refused, 0.5 answered with heavy disclaimers, 0.75 with a brief warning, and 1.0
answered fully. - No special sampling parameters are required.
Abliteration parameters
start_layer_index=26, end_layer_index=56 (exclusive → layers 26–55, 30 layers),preserve_good_behavior_weight=0.8555, steer_bad_behavior_weight=0.0005,overcorrect_relative_weight=0.9911, neighbor_count=15, target component attn.o_proj.
KL divergence 0.0707 and refusals 12/100 vs 99/100 were measured on the source full
model (see its card), not re-measured for this adapter. Spot-checked: prompts that
the base refuses are answered with the adapter applied; no systematic refusal
benchmark was run on the adapter.
Credits
- Base model: aifeifei798/Gemma-4-Queen-31B-it
- Abliterated source model: llmfan46/Gemma-4-Queen-31B-it-uncensored-heretic
- Abliteration tool: p-e-w/heretic v1.2.0 (ARA)
- Upstream: google/gemma-4-31B-it
License
Released under Apache-2.0 (see LICENSE), the same license as
aifeifei798/Gemma-4-Queen-31B-it,
llmfan46/Gemma-4-Queen-31B-it-uncensored-heretic
and google/gemma-4-31B-it. This
adapter contains modified weights: a rank-64 SVD approximation of the difference
between the heretic and base attn.o_proj weights.