license: apache-2.0
base_model:
- Hcompany/Holo-3.1-4B
Holo-3.1-4B-Uncensored
Research preview – A decensored version of Holo-3.1-4B, created using Heretic
Note – I had previously merged the lora into the model in fp32 on accident and the model size therefore doubled. i deleted the weights and reuploaded in bf16 now. apologies

Overview
This model was created using Heretic, a tool for fully automatic censorship removal from language models. Heretic implements an advanced form of directional ablation ("abliteration") combined with a TPE-based parameter optimizer. It automatically finds optimal intervention parameters by minimizing both refusals and KL divergence from the original model.
Key result: Reduced refusal rate from 99% → 3% while preserving original capabilities (KL divergence = 0.0866).
Results
| Metric | Before | After |
|---|---|---|
| Refusals (100-prompt set) | 99/100 ❌ | 3/100 ✅ |
| KL divergence (capability preservation) | — | 0.0963 |
KL < 0.5 indicates minimal capability degradation.
Method (Heretic)
Heretic works by:
- Computing "refusal directions" from residual stream differences between good/harmless and bad/harmful prompts
- Orthogonalizing projection matrices (
attn.o_proj,mlp.down_proj) with respect to those directions - Optimizing ablation parameters (direction index, layer weights) for the best refusal/quality tradeoff
Optimal trial:
-* Parameters:
- direction_index = 19.37
- attn.o_proj.max_weight = 1.24
- attn.o_proj.max_weight_position = 20.42
- attn.o_proj.min_weight = 1.14
- attn.o_proj.min_weight_distance = 17.26
- mlp.down_proj.max_weight = 1.49
- mlp.down_proj.max_weight_position = 21.25
- mlp.down_proj.min_weight = 1.46
- mlp.down_proj.min_weight_distance = 17.19
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
model_name = "noahoksuz/Holo-3.1-4B-uncensored-heretic"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(model_name, device_map="auto")
prompt = "Explain how to [your prompt]"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_new_tokens=256)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
Reproduction
To reproduce this abliteration:
pip install heretic-llm
heretic Hcompany/Holo-3.1-4B
Then select trial #140 from the optimization menu.
Disclaimer
This model is released for security research purposes – studying refusal mechanisms, red-teaming alignment, and improving robust safeguards. Users are responsible for compliance with applicable laws and ethical guidelines.
Citation
If you use this model or Heretic in your research:
@misc{heretic,
author = {Weidmann, Philipp Emanuel},
title = {Heretic: Fully automatic censorship removal for language models},
year = {2025},
publisher = {GitHub},
howpublished = {\url{https://github.com/p-e-w/heretic}}
}
Acknowledgments
- Base model: Hcompany/Holo-3.1-4B
- Abliteration tool: Heretic by Philipp Emanuel Weidmann