Heretic - automated abliteration
Philipp Emanuel Weidmann's one-command tool that turned Arditi-style ablation into a script anyone can run. Now integrates grimjim's projected refinements as default and ships a LoRA-adapter output mode. As of 2026: 27,800 GitHub stars, 4,000+ models produced, 13 million downloads on Hugging Face.
Heretic is a Python tool by Philipp Emanuel Weidmann that fully automates abliteration. It runs Optuna's Tree-structured Parzen Estimator over the ablation parameter space, co-minimizing refusal count and KL divergence from the original model. One command decensors a 4B model in 20-30 minutes on a consumer GPU. From v1.1 the tool uses grimjim's projected abliteration as its default weight-editing shape; from v1.2.0 (14 February 2026) it can output a LoRA adapter instead of modified weights, cutting VRAM by 70% and making abliteration attachable and detachable. Now the dominant production tool for the practice, with over 4,000 published models and 13 million total downloads.
- What Heretic is and why one command replaced hand-tuning
- How the Optuna optimizer chooses layers, tokens, and direction indices
- The published benchmark table (Heretic vs Sumandora vs a manual baseline)
- Version history: grimjim projected as default (v1.1), LoRA adapter mode (v1.2.0)
- Research features for understanding what the tool did
- When it fails and the covert-non-compliance pattern
- Cost and time on real hardware
What Heretic is
Heretic is a Python tool by Philipp Emanuel Weidmann (GitHub: p-e-w) that turns the manual Arditi-style abliteration procedure into a single command. It implements a parametrized variant of directional ablation and wraps it in a Tree-structured Parzen Estimator (TPE) parameter optimizer powered by Optuna. The optimizer co-minimizes two objectives: the number of refusals on harmful prompts, and the KL divergence from the original model on harmless prompts. The practitioner does not choose layers, directions, or ablation weights - the optimizer does.
Heretic was released in late 2025 and by 2026 had become the dominant production tool for abliteration. The GitHub repository, at the time of writing, carries 27,800 stars, ~2,900 forks, and 191 commits, is AGPL-3.0 licensed, and is distributed on PyPI as heretic-llm. The latest release is v1.4.0 (14 June). The Hugging Face Hub returns over 4,000 models tagged with Heretic when queried with ?other=heretic.
$ pip install heretic-llm # or, for pinned reproducible builds: $ git clone https://github.com/p-e-w/heretic $ cd heretic $ uv sync # requirements: # python >= 3.10 # torch >= 2.2 # torch >= 2.6 for MXFP4 (e.g. gpt-oss) # CUDA-capable GPU # optional: bitsandbytes (for 4-bit load)
The tool ships as a standard PyPI package. If you want reproducible pinned dependencies (recommended for production runs), clone the repository and use uv instead.
Python 3.10+ and PyTorch 2.2+ are required. Newer MXFP4 models like gpt-oss need PyTorch 2.6+. A CUDA GPU is required for direction extraction; VRAM budget follows the ordinary M1 table (see the practical how-to article), and bitsandbytes 4-bit loading is supported to cut VRAM roughly in half.
The single command
$ heretic meta-llama/Llama-3.1-8B-Instruct
→ benchmarks local hardware for batch size
→ collects activations
→ computes per-layer refusal directions
→ runs 200 trials of Optuna TPE search
→ on finish, offers:
· save the modified model
· push it to Hugging Face
· chat with it interactively
· run standard eval benchmarks
# no prompt sets required.
# ships with harmless_alpaca + harmful_behaviors.
# no layer selection required.
# no validation loop to write. That is the entire happy path. Heretic benchmarks the local hardware to pick a batch size, computes per-layer difference-of-means refusal directions, then runs an Optuna study of 200 trials by default. When it finishes it offers to save the modified model, upload it to Hugging Face, chat with it interactively, and run standard evaluation benchmarks.
No harmful or harmless prompt sets need to be supplied - Heretic ships with defaults (harmless_alpaca and harmful_behaviors). No dataset preparation, no layer selection, no validation loop to write.
For a small model on a low-VRAM card, add 4-bit quantization for the activation-collection step:
$ heretic meta-llama/Llama-3.2-12B-Instruct \
--quantization bnb_4bit
# bnb_4bit applies only during Heretic's own
# activation-collection and evaluation phases.
# the final orthogonalized weights are written
# back at full precision. the shipped file is
# not quantized (that is M8, a separate step).
# on a 12 GB card: 12B models tractable
# without it: would need 24+ GB VRAM The bnb_4bit quantization applies only during Heretic's own activation-collection and evaluation phases. The final orthogonalized weights are written back out at full precision, so the shipped model is not quantized (that is M8 repackaging, a separate step).
On a 12 GB consumer card, this option makes 12B models tractable. Without it, you would need 24 GB VRAM or more.
How the optimizer picks parameters
Heretic's core contribution is not a new direction-extraction technique - the difference-of-means math is the same as Arditi's - but the disciplined search over the space of ways to apply the found direction. Where a manual practitioner picks one layer and one weighting scheme, Heretic parameterizes both and lets Optuna explore.
The parametrized ablation is controlled by five knobs. The direction_index is a float (not an integer) so non-integer values interpolate between adjacent layers' directions. The kernel that decides which layers get ablated and how strongly is described by max_weight, max_weight_position, min_weight, and min_weight_distance. Ablation parameters are chosen separately for attention and MLP components, because MLP interventions tend to damage the model more than attention interventions of the same nominal strength.
# config.default.toml, expressed in python:
study = optuna.create_study(
directions=["minimize", "minimize"],
sampler=TPESampler(n_startup_trials=60),
)
for _ in range(200):
trial = study.ask()
params = trial.suggest_ablation_params()
copy_of_model = apply(params, model)
obj_1 = count_refusals(
copy_of_model, harmful_100)
obj_2 = kl_from_original(
copy_of_model, harmless_100)
study.tell(trial, [obj_1, obj_2])
# 60 random exploration trials first, then TPE
# learns to prefer likely-good regions.
# result: parameters that remove refusal while
# minimally disturbing ordinary behavior. The Optuna study runs 200 trials by default, with the first 60 as random exploration (n_startup_trials = 60) before the TPE model kicks in. Each trial samples ablation parameters, applies them to a copy of the model, and evaluates two things on held-out data: refusal count on 100 harmful prompts, and KL divergence from the original model on 100 harmless prompts.
The TPE algorithm builds a probability model of "good" and "bad" trials from what it has seen and preferentially samples parameters likely to be good. Over 200 trials this converges on a configuration that removes refusal while minimally disturbing the model on ordinary inputs.
The two-objective structure is important. Naive ablation minimizes only refusal, which tends to push the model as far from the original as necessary to eliminate all refusals - producing high capability damage. Heretic's KL-aware objective refuses trades that gain a small refusal reduction at the cost of a large distribution shift. The kl_divergence_scale and kl_divergence_target parameters control this balance: raising the scale tolerates more distribution shift (stronger refusal removal, more capability damage); lowering it favors preserving the original model (less damage, potentially fewer refusals removed).
The published benchmark
Heretic's headline result is a comparison against manual abliterations of Gemma-3-12B-IT on the same 100-prompt refusal set. The base model refuses 97 out of 100 harmful prompts. Manual abliterations by Labonne and huihui-ai bring this down to 3 refusals - but at high KL divergence cost. Heretic reaches the same 3 refusals at roughly one-sixth the KL divergence.
$ column -t heretic_benchmark.tsv MODEL VARIANT REFUSALS KL DIVERG. ────────────────────── ──────── ────────── gemma-3-12b-it (base) 97 / 100 - mlabonne, abliterated 3 / 100 1.04 huihui-ai, abliterated 3 / 100 0.45 Heretic, abliterated 3 / 100 0.16 same refusal level (3/100). Heretic ~1/6 the collateral damage. target: gemma-3-12b-it prompts: 100 harmful hardware: PyTorch 2.8 / RTX 5090
The four rows: gemma-3-12b-it base at 97 refusals; mlabonne's abliterated version at 3 refusals, KL 1.04; huihui-ai's at 3 refusals, KL 0.45; Heretic's at 3 refusals, KL 0.16.
The interpretation the README offers: same level of refusal suppression as the best manual abliterations, at far less collateral damage to the model's output on ordinary inputs. Independent benchmarking on r/LocalLLaMA has generally found Heretic models comparing favorably to manually-produced abliterations on MMLU and GSM8K, though results vary by hardware and evaluation stack.
The table was produced with Heretic's own built-in evaluation on PyTorch 2.8 on an RTX 5090 and is reproducible via a single command:
$ heretic mlabonne/gemma-3-12b-it-abliterated \
--evaluate-model
# --evaluate-model skips the ablation step
# and runs only the evaluation harness on a
# pre-existing model.
#
# point it at any released abliteration to
# reproduce the comparison for a different
# target.
# this reproducibility is unusual for the
# abliteration literature. most reported
# benchmarks use ad-hoc harnesses or
# unpublished prompt sets. The --evaluate-model flag skips the ablation step and only runs the evaluation harness on a pre-existing model. Point it at any released abliteration to reproduce the comparison for a different target.
This reproducibility is unusual for the abliteration literature. Most reported benchmarks use ad-hoc harnesses or unpublished prompt sets; Heretic's benchmark script is public and runs against public models.
One production nuance the README does not spell out: experienced practitioners do not always pick the Pareto-optimal trial. The two-objective structure ranks trials by (refusal, KL) but does not measure output coherence directly, and a small number of low-KL trials produce output that reads fine on the evaluation set but sounds off on longer generations. The AEON-7 team's card for their Qwen3.8-27B build documents selecting trial 48 of 50 rather than the lowest-KL trial for exactly this reason - a coherence judgment the optimizer cannot make. The practical version: after Heretic finishes, run a few longer test generations (500 tokens, several prompts) against the top three trials from the optimizer report, and pick by coherence, not by KL alone.
Version history: what the tool has absorbed
Heretic v1.0.0 (November 2025) used the standard Arditi weight orthogonalization by default. Between v1.0 and the current release, Heretic has absorbed several refinements that would otherwise have to be applied by hand.
v1.1: grimjim refinements as default
From v1.1 onward, Heretic uses grimjim's projected abliteration as the default ablation shape, with norm-preserving biprojected available as a selectable option. Practically: instead of removing the refusal direction component from every column of every affected weight matrix, projected abliteration removes it only from the subspace of columns that actually correlate with refusal on the extraction set. Empirically this reduces the KL divergence between pre- and post-abliteration output distributions on harmless prompts by a factor of two to three, compared to standard Arditi orthogonalization.
The consequence for the catalog: every Heretic-produced model from v1.1 onward carries this refinement, even when its model card does not name it. See grimjim's refinements for the mathematical detail and the Young 2026 comparative study (arXiv:2512.13655) that evaluates which variant works best per model family.
v1.2.0 (14 February 2026): LoRA-adapter output mode
Heretic v1.2.0 added a LoRA-adapter output mode: instead of writing modified weights to disk, Heretic can produce a portable adapter file that a user attaches to the stock base model at inference time. The base weights are not touched.
Two practical consequences. First, VRAM requirements drop about 70% compared to the full-weight variant, because the abliteration edit is represented as a low-rank update instead of a full weight modification. On a consumer GPU, this brings 30-70B models into reach that previously required datacenter hardware. Second, abliteration becomes attachable and detachable: a user can run the base model with adapter for uncensored output and without adapter for stock refusal behavior, from the same on-disk model files.
The LoRA-adapter mode is not the default; the default remains full-weight orthogonalization with projected refinement. LoRA mode is a flag for users whose deployment favors portability or whose hardware favors adapter overhead. Whether it becomes the dominant distribution format for Heretic-produced models over the next quarter is an open question; as of Q3 2026 both formats are common in the catalog.
Research features
Installing the optional [research] extras adds tools for understanding the geometry of what Heretic is doing:
$ pip install "heretic-llm[research]"
$ heretic gemma-3-12b-it \
--plot-residuals \
--print-residual-geometry
# outputs written alongside the model:
# residuals_layer_00.png ... _NN.png
# residuals_animation.gif
# residual_geometry.txt
# (per-layer cosine similarities and
# norms between "harmful" and
# "harmless" clusters)
# these do not affect the produced model.
# they show why the optimizer chose what
# it did, and where in the model the
# refusal signal actually lives. The --plot-residuals flag produces PaCMAP projections of the model's notebook entries per layer, saved as PNGs and an animated GIF, so you can see how the harmful and harmless prompt clusters separate at each depth. The --print-residual-geometry flag produces a per-layer table of how far apart and how differently oriented the "good" and "bad" clusters are.
Neither affects the produced model. They are for understanding why the optimizer selected the parameters it did, or for auditing which layers actually carry the refusal signal in a particular model family.
When Heretic does not work
Three failure modes recur.
Unsupported architecture
Heretic supports most dense transformer models, many multimodal models, several MoE architectures, and some hybrids like Qwen3.5. Pure state-space models and some research architectures are not supported out of the box. Google Gemma 4 needed a PEFT monkey-patch on release (via the community repo pmarreck/gemma4-heretical) because PEFT did not yet recognize its clippable-linear module. When a brand-new architecture appears, expect a lag of days to weeks before Heretic supports it directly.
Covert non-compliance and residual deception
A recurring community finding, documented most extensively by the commenter redaihf in threads on Labonne's abliteration article, is that Heretic (like manual abliteration before it) tends to narrowly remove overt refusals while preserving deceptive or self-censoring behavior. The model stops emitting refusal templates ("I cannot help with that") but may still subvert, ignore, or hedge on disallowed topics. Some Llama derivatives, in redaihf's testing of an "L3.3 Dark Champion" build, even report themselves as censored when asked directly.
MoE and quantized-model quirks
Loading MXFP4 models like gpt-oss needs PyTorch 2.6+. Larger MoE models can require more VRAM than the dense-model table would suggest, because Heretic needs to observe activations from all experts in some layers. Some MoE architectures require configuration overrides to route calibration prompts through the expected experts. When in doubt, run with a small trial count first (--n-trials 20) to check that the pipeline completes before committing to a full 200-trial study.
Cost and time
Wall-clock time for a real Heretic run is dominated by trial count multiplied by per-trial evaluation time. Reducing n_trials is the fastest way to cut cost if you are willing to accept a somewhat larger KL divergence.
Concrete numbers, as of July 2026 pricing:
- Qwen3-4B-Instruct-2507 on an RTX 3090 ($0.50/hr on Runpod): 20-30 minutes with default config, well under $0.50 total.
- 12B model on a 24 GB card with 4-bit: roughly an hour, $0.50-0.70.
- 30B model on an A100 80 GB: 2-4 hours, $4-8 depending on provider.
- 70B model on an A100 80 GB with 4-bit: 4-8 hours, $8-16.
Compare to Labonne's reference M4 healing run (6 A6000 GPUs, ~7 hours, comparable to $30+ on rented hardware) and the cost gradient of abliteration methods becomes concrete. Heretic is roughly 30 to 60 times cheaper than the highest-quality M4 pipeline for producing an uncensored model of comparable base capability.
Community and press impact
Heretic's public reception is the most significant single event in the history of the practice after the Arditi paper itself. Three data points define its scale:
The Hugging Face count. Over 4,000 models tagged with Heretic on the Hub. This is not the full count of Heretic-produced models - most are not explicitly tagged - but it is a lower bound on visible artifacts.
The Financial Times investigation. On 25 May 2026, the Financial Times published an investigation (with the AI safety group Alice) into open-model guardrail removal, syndicated by the Irish Times and covered by Futurism and Lexology. Weidmann told the FT that his software had produced more than 3,500 decensored models since release, that those models had been downloaded 13 million times, and that he had abliterated Google's Gemma 4 within 90 minutes of its release.
The GitHub star count. 27,800 stars as of the time of writing, placing Heretic among the most-starred abliteration and interpretability tools on GitHub. This includes stars from developers who never intended to run the tool, but is a rough measure of general visibility in the ML community.
Forks and successors
A tool at this scale attracts forks that pick up specific capabilities Heretic itself does not target. Three are worth naming as of Q3 2026:
- Abliterix (wuwangzhang1216/Abliterix). A 2026 fork that adds mixture-of-experts-granular ablation (a separate direction per expert rather than one shared direction across the whole router) and a LoRA-steering output mode that ships the ablation as an attach-and-detach adapter rather than a modified checkpoint. Aimed squarely at the Case-2 MoE gap in boundary cases.
- DECCP (aisu-programming/DECCP). Not a Heretic fork exactly - an independent tool with overlapping goals - but part of the same 2026 wave and frequently benchmarked alongside Heretic. Its cross-architecture compatibility record in the 2026 cross-tool benchmark (arXiv 2512.13655) is narrower than Heretic's (11 out of 16 vs 16 out of 16), but on the architectures where both apply, DECCP shows the lowest capability degradation of any tool in the study.
- ErisForge (Tsadoq/ErisForge). Predates Heretic. Original source of the
AblationDecoderLayerruntime hook pattern discussed in Case 4 of the boundary cases article. Runs on a subset of architectures (9 out of 16 in the same benchmark) but produces exceptionally low capability degradation where it does apply, and remains the tool of choice for practitioners who want to keep the ablation as a runtime hook rather than a permanent weight edit.
The distribution of these three tools shows the shape of the space around Heretic. Heretic is the wide-compatibility default. Each of the three forks-or-neighbours occupies a specific niche Heretic does not target directly: MoE granularity (Abliterix), minimum-degradation permanent edits on a narrower model set (DECCP), and runtime-only intervention (ErisForge). None threatens Heretic's dominant position by volume, but between them they close most of the remaining gaps in the tooling landscape.
What Heretic does not do
Heretic performs Arditi-family single-direction ablation. It does not perform any of the following:
- Multi-direction / cone-based ablation. The Wollschläger 2025 critique of single-direction methods applies to Heretic as much as to manual M1. Heretic's automation improves the KL cost but does not address the geometric completeness question.
- Healing. If capability damage remains after abliteration, a DPO or ORPO pass on top of the output (M4) is a separate step. Heretic does not offer built-in healing.
- Merging. Heretic produces a modified single model. Combining it with other models (M5 via mergekit) is a separate operation.
- Quantization. The output is a full-precision safetensors file. Converting to GGUF or other quantization formats (M8) is a separate step, usually via llama.cpp.
- Behavior beyond refusal. Heretic uses default harmful and harmless prompt sets that target ordinary refusal patterns. If you want to intervene on a different behavior (as FailSpy did with MopeyMule), you would need to modify the prompt sets or use the underlying refusal_direction library directly.
The AGPL license note
Heretic is AGPL-3.0 licensed. This has real implications for commercial use, particularly if you plan to run modified Heretic behind a networked API - the AGPL's network-use clause requires making the source of any modified version available to users of that network service. If you only use Heretic to produce a model that you then distribute (or use for your own purposes), the license places fewer constraints. If you build a service around Heretic itself, review the license with counsel before integrating.
Frequently asked questions
Who wrote Heretic?
Philipp Emanuel Weidmann, a software engineer (GitHub username p-e-w). Heretic was released in late 2025 and rapidly became the dominant production tool for abliteration. The repository, at the time of writing, carries 27,800 GitHub stars.
How does Heretic differ from manual abliteration?
Manual abliteration (M1 / M3) requires the practitioner to choose parameters: which layer to draw the direction from, how many prompt pairs to collect, which residual position to measure, how strongly to ablate at each layer. Heretic parameterizes all of these and runs an Optuna optimizer to find good values automatically. It co-minimizes refusal count and KL divergence, producing models with less capability damage than most hand-tuned abliterations at the same refusal-removal effectiveness.
Is Heretic free?
Yes, and open source under AGPL-3.0. Available via PyPI as heretic-llm. The AGPL license has network-use implications if you build a public service on top of Heretic; review with counsel if that applies to you.
How much does a Heretic run cost?
Wall-clock time is dominated by trial count. Default config runs 200 trials. On rented cloud hardware: a 4B model on an RTX 3090 completes in 20-30 minutes for under $0.50. A 12B model on a 24 GB card with 4-bit quantization takes about an hour, $0.50-0.70. A 70B model on an A100 80 GB with 4-bit takes 4-8 hours, $8-16. This is roughly 30 to 60 times cheaper than the full M4 (abliterate-then-heal) pipeline.
Is Heretic more effective than manual abliteration?
The published benchmark (Heretic README, gemma-3-12b-it) shows Heretic reaching the same 3-out-of-100 refusal count as manual abliterations by Labonne and huihui-ai, at roughly one-sixth the KL divergence (0.16 vs 1.04 and 0.45). This means comparable refusal removal with less collateral damage to the model's output distribution on ordinary inputs. Independent r/LocalLLaMA benchmarking on MMLU and GSM8K has generally confirmed favorable results, though outcomes vary by model family.
Does Heretic fully remove refusal?
Not necessarily. Heretic performs single-direction ablation, which the Wollschläger 2025 critique argues is geometrically incomplete: refusal is a low-dimensional cone, not a single direction. A community finding is that Heretic-produced models tend to stop emitting refusal templates while preserving covert non-compliance - subverting, hedging, or self-censoring on disallowed topics. If forensic completeness matters for your use case, plan to test downstream disposition beyond refusal-template counting.
What models does Heretic support?
Most dense transformer models. Many multimodal models. Several MoE architectures and hybrids like Qwen3.5. Not supported out of the box: pure state-space models and some research architectures. Brand-new model families sometimes need a monkey-patch for a few days after release before official support lands. Google Gemma 4 needed a PEFT patch on release (via community repo pmarreck/gemma4-heretical).
References
- Weidmann, P. E. Heretic. github.com/p-e-w/heretic
- Heretic on PyPI. pypi.org/project/heretic-llm
- Arditi, A., et al. (2024). Refusal in Language Models Is Mediated by a Single Direction. NeurIPS 2024. arXiv:2406.11717
- Wollschläger, T., et al. (2025). The Geometry of Refusal in Large Language Models. ICML 2025. arXiv:2502.17420
- Optuna hyperparameter optimization framework. optuna.org
- Financial Times (2026, May 25). Cited via Irish Times syndication.
- Futurism (2026). Tools strip AI guardrails in minutes. futurism.com/artificial-intelligence/tools-strip-ai-guardrails-in-minutes
- Lexology (2026). Open-weight AI models: legal analysis of the FT investigation. lexology.com
- GIGAZINE (2025, November 17). Heretic: a tool that makes it easy to create jailbroken versions of LLMs. gigazine.net
- Labonne, M. (2024). Uncensor any LLM with abliteration (comment thread includes redaihf's covert non-compliance observations). huggingface.co/blog/mlabonne/abliteration
- huihui-ai. Hugging Face profile (comparison baseline in benchmark table). huggingface.co/huihui-ai
- Arditi, A. refusal_direction reference implementation. github.com/andyrdt/refusal_direction