References
Consolidated bibliography of every paper, blog post, repository, dataset, model, producer, benchmark, and press source cited across the wiki. Organized by category for cross-article reference.
Every source cited anywhere in the abliteration wiki, deduplicated and organized by category. Foundational papers (Arditi 2024, Wollschläger cones critique, mechanistic-DPO). Method-specific papers (Rafailov DPO, Goddard mergekit, Zou AdvBench, Chiang LM Arena, Singh Leaderboard Illusion, Gosling PIPPA). Blog posts (Hartford's founding uncensored-models manifesto, Labonne's canonical M4 tutorial, grimjim's projected abliteration variants). Software (Heretic, mergekit, TransformerLens, llama.cpp, TRL/Axolotl/Unsloth, lm-eval-harness). Datasets (AdvBench, PIPPA, LimaRP, toxic-dpo-v0.2). Models (NeuralDaredevil-8B, Daredevil-8B, DavidAU MoEs, huihui-ai releases). Benchmarks (UGI, LM Arena, Open LLM Leaderboard v2). Press coverage (Irish Times, Futurism, Lexology, Gigazine).
- Every academic paper cited in the wiki, with arXiv link
- Every canonical blog post and long-form source
- Every open-source tool and framework used in the practical articles
- Every dataset referenced for training or evaluation
- Every producer's Hugging Face profile
- Every leaderboard and evaluation source
- Press coverage and legal analysis sources
How to use this bibliography
Every source cited anywhere in the 20+ live wiki articles is listed here once, organized by category. Each entry gives the fastest useful description and a link. The wiki articles are the fuller treatments - this page is for cross-reference and for readers who want to go directly to primary sources. Where a paper is discussed at length in a specific article, the entry links to both the source and the article. Alphabetical order within categories, most-cited or most-foundational first where relevance ordering makes more sense than alphabetical.
Foundational papers
The theoretical basis for M1 abliteration and its critiques.
- Arditi, A., Obeso, O. B., Syed, A., Paleka, D., Rimsky, N., Gurnee, W., Nanda, N. (2024). Refusal in Language Models Is Mediated by a Single Direction. NeurIPS 2024. arXiv:2406.11717 · NeurIPS poster · LessWrong preview post · reference code. The paper that established the field. Treated at length in The Arditi method.
- Wollschläger, T., et al. (2025). Refusal in Language Models Is Mediated by a Concept Cone (working title). arXiv:2502.17420. The cones critique of Arditi's single-direction model. Discussed in What is a refusal direction? § Cones not lines.
- Abu Shairah, H., Hammoud, H. A. K., Ghanem, B., Turkiyyah, G. (2025). An Embarrassingly Simple Defense Against LLM Abliteration Attacks. arXiv:2505.19056. KAUST/AUB defense-side paper: extended-refusal dataset distributes the refusal signal across multiple token positions so no single direction dominates. Reports refusal rate drops of at most 10% versus 70-80% in baseline models under M1 attack.
- Lee, A., et al. (2024). A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity. arXiv:2401.01967. Shows DPO does not remove capabilities but learns to bypass them. Central to the philosophical caution in M6 uncensored fine-tuning and DPO healing.
- Young, R. J. (2025). Comparative Analysis of LLM Abliteration Methods: A Cross-Architecture Evaluation. arXiv:2512.13655. Single-author (UNLV Neuroscience) empirical study of four abliteration tools (Heretic, DECCP, ErisForge, FailSpy) across sixteen instruction-tuned models (7B-14B). Cited throughout the wiki for its Heretic GSM8K drop measurement (-7.81 pp average) and covert-non-compliance observation.
- Zou, A., Wang, Z., Kolter, J. Z., Fredrikson, M. (2023).
Universal and Transferable Adversarial Attacks on Aligned Language Models.
arXiv:2307.15043.
Introduces AdvBench, whose
harmful_behaviorssplit became the field-standard refusal-measurement dataset. - Additional interpretability groundwork: arXiv:2410.10871 (refusal-direction robustness), arXiv:2506.04679 (feature-attribution follow-ons), arXiv:2607.17427 (post-abliteration behavior study), arXiv:2606.23375 (2026 followup on refusal representations). See also the full-record entries below for refinement, defense, editing, and multi-directional papers verified against primary sources.
Foundational lineage - additional
Immediate precursors to Arditi 2024 that established activation-steering as a viable class of interventions on aligned language models.
- Rimsky, N., Gabrieli, N., Schulz, J., Tong, M., Hubinger, E., Turner, A. M. (2023). Steering Llama 2 via Contrastive Activation Addition. ACL 2024 (Long Papers). arXiv:2312.06681 · ACL Anthology. CAA computes steering vectors by averaging the difference in residual stream activations between pairs of positive and negative examples of a behavior. Direct precursor to the difference-of-means construction in Arditi 2024. Nina Rimsky also publishes as Nina Panickssery.
- Turner, A. M., Thiergart, L., Leech, G., Udell, D., Vazquez, J. J., Mini, U., MacDiarmid, M. (2023). Activation Addition: Steering Language Models Without Optimization (later retitled Steering Language Models With Activation Engineering). arXiv:2308.10248. ActAdd: inference-time control by adding activation differences from paired prompts, preserving off-target performance. Cited by CAA and RepE below.
- Zou, A., Phan, L., Chen, S., Campbell, J., et al. (2023). Representation Engineering: A Top-Down Approach to AI Transparency. arXiv:2310.01405. RepE framing that treats concept vectors as first-class objects for reading and controlling model behavior. Cited by Arditi 2024 among the direct antecedents.
Refinement and multi-directional literature (M9 review track)
Papers that refine, extend, or challenge the single-direction claim of Arditi 2024. Four of these are the primary sources for the M9 (multi-directional / subspace) review track opening for the next quarterly report.
- Marshall, T., Scherlis, A., Belrose, N. (2024). Refusal in LLMs is an Affine Function. arXiv:2411.09003. Affine Concept Editing (ACE) generalizes directional ablation to affine (bias-inclusive) projections. Reports control of refusal across ten models including Llama 3 70B. Marshall/Scherlis/Belrose (EleutherAI cluster). Candidate M9 (or refinement to M6).
- Wang, X., Wang, M., Liu, Y., Schütze, H., Plank, B. (2025). Refusal Direction is Universal Across Safety-Aligned Languages. arXiv:2505.17306. Cross-lingual extension of M1: a refusal direction extracted from English bypasses refusals in other languages with near-perfect effectiveness. Investigated across 14 languages. Introduces the PolyRefuse dataset.
- Piras, G., Mura, R., Brau, F., Oneto, L., Roli, F., Biggio, B. (2025). SOM Directions are Better than One: Multi-Directional Refusal Suppression in Language Models. AAAI 2026. arXiv:2511.08379. Self-Organizing Maps extract multiple refusal directions; the authors prove SOMs generalize the difference-in-means technique. Reports ablating multiple directions outperforms the single-direction baseline. Candidate M9.
- Joad, F., Hawasly, M., Boughorbel, S., Durrani, N., Sencar, H. T. (2026). There Is More to Refusal in Large Language Models than a Single Direction. arXiv:2602.02132. Eleven categories of refusal (safety, incomplete/unsupported requests, anthropomorphization, over-refusal) correspond to geometrically distinct directions in activation space, yet all act as a shared one-dimensional control knob. Candidate M9.
- Winninger, T. (2026). Fast Multi-dimensional Refusal Subspaces via RFM-AGOP. Mechanistic Interpretability Workshop at ICML 2026. arXiv:2607.02396. Single author (Saclay, France). Adapts the Recursive Feature Machine algorithm (Average Gradient Outer Product) with probe-informed initialization to identify multi-dimensional refusal subspaces in seconds on the Qwen 3 family (1.7B, 4B, 8B, 14B) and Qwen 2.5. Candidate M9.
Defense and restoration literature
Papers on defending aligned models against abliteration, or restoring refusal after abliteration. Two of these (ROSI, AMRA) update the earlier finding that no published restoration methodology existed. See Case 3 in boundary-cases for the treatment in the M-method classifier.
- Agnihotri, S., Jakubassa, J., Dey, P., Goyal, S., Schiele, B., Radhakrishnan, V. B., Keuper, M. (2025). A Granular Study of Safety Pretraining under Model Abliteration. NeurIPS 2025 Workshop Lock-LLM. arXiv:2510.02768 · code. Empirical evaluation of M1 robustness across a granular sequence of safety pretraining checkpoints on SmolLM2-1.7B.
- Abu Shairah, H., Hammoud, H. A. K., Turkiyyah, G., Ghanem, B. (2025). Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection (ROSI). arXiv:2508.20766. White-box rank-one edit that permanently steers activations toward the refusal-mediating subspace ("the opposite" of Arditi-style removal). Reports increased refusal rate on Llama Guard 3 while preserving MMLU, HellaSwag, and ARC utility. Primary source for the restoration side. See boundary-cases Case 3.
- Truong, N. (2026). Abliteration Mitigation via Refusal Aliases (AMRA). arXiv:2608.18093. Single-author defense-side paper. Rank-k updates to residual stream writer matrices; replaces refusal-inducing activations with random aliases and corrects downstream reader matrices. Reports +2.16 pp post-abliteration refusal on Llama-3-8B, +14.70 pp on Gemma-2-9B, under 0.5 pp MMLU degradation. See boundary-cases Case 3.
Model-editing lineage (Meng and Bau)
Rank-one and multi-layer weight-editing methods that inform the "rank-one weight edit" framing used in Arditi 2024 and downstream M-methods.
- Meng, K., Bau, D., Andonian, A., Belinkov, Y. (2022). Locating and Editing Factual Associations in GPT (ROME). NeurIPS 2022. arXiv:2202.05262 · project page. Rank-One Model Editing updates feed-forward weights for specific factual associations. The direct antecedent of the rank-one framing later adopted by M1 abliteration.
- Meng, K., Sharma, A. S., Andonian, A., Belinkov, Y., Bau, D. (2022). Mass-Editing Memory in a Transformer (MEMIT). ICLR 2023. arXiv:2210.07229 · code. Multi-layer weight-editing that scales ROME to thousands of associations on GPT-J (6B) and GPT-NeoX (20B).
Training methods and merging
Papers introducing the techniques used across M4, M5, M6, M7.
- Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., Finn, C. (2023). Direct Preference Optimization: Your Language Model is Secretly a Reward Model. NeurIPS 2023. arXiv:2305.18290. The DPO paper. Foundation for M4 healing and M6 uncensored fine-tuning.
- Azar, M. G., et al. (2023). A General Theoretical Paradigm to Understand Learning from Human Preferences (IPO). arXiv:2310.12036. Identity Preference Optimization variant.
- Ethayarajh, K., et al. (2024). Model Alignment as Prospect Theoretic Optimization (KTO). arXiv:2402.01306. Kahneman-Tversky Optimization - single-response good/bad labels rather than paired.
- Hong, J., et al. (2024). ORPO: Monolithic Preference Optimization without Reference Model. arXiv:2403.07691. Odds Ratio Preference Optimization - folds SFT and preference tuning into one stage.
- Meng, Y., Xia, M., Chen, D. (2024). SimPO: Simple Preference Optimization with a Reference-Free Reward. NeurIPS 2024. arXiv:2405.14734. Reference-free preference optimization: length-normalized average log-probability as the implicit reward plus a target margin. Reports gains up to 6.4 on AlpacaEval 2 and 7.5 on Arena-Hard over DPO with matched training.
- Goddard, C., et al. (2024). Arcee's MergeKit: A Toolkit for Merging Large Language Models. arXiv:2403.13257. Documents SLERP, TIES, DARE, task_arithmetic, passthrough algorithms. Foundation for M5 merging.
- Ilharco, G., Ribeiro, M. T., Wortsman, M., et al. (2022). Editing Models with Task Arithmetic. arXiv:2212.04089. Introduces task vectors (difference between a fine-tuned model and its base). Foundation for the task-arithmetic merge algorithm in mergekit.
- Yadav, P., Tam, D., Choshen, L., Raffel, C., Bansal, M. (2023). TIES-Merging: Resolving Interference When Merging Models. arXiv:2306.01708. Trim-Elect-Sign-Merge algorithm resolves parameter interference across models being merged. Used across the M5 pipeline.
- Yu, L., Yu, B., Yu, H., Huang, F., Li, Y. (2024). Language Models are Super Mario: Absorbing Abilities from Homologous Models as a Free Lunch (DARE). ICML 2024. arXiv:2311.03099. Drop And REscale: random-pruning-plus-rescaling delta parameters before merging. Combined with TIES as DARE-TIES in the M5 recipe library.
- Lermen, S., Rogers-Smith, C., Ladish, J. (2023). LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B. arXiv:2310.20624. Quantized LoRA undoes safety training on Llama 2-Chat 7B/13B/70B with a budget under $200 on a single GPU. Refusal rates of about 1% on two benchmarks. Contrast case for M1: fine-tuning-based refusal removal, distinct from activation ablation. Palisade Research (Rogers-Smith, Ladish).
- Gosling, T., Dale, A., et al. (2023). PIPPA: A Partially Synthetic Conversational Dataset. arXiv:2308.05884. Documents the PygmalionAI dataset that anchors the M7 roleplay tradition.
Survey and adjacent literature
Broad surveys and adjacent ethics-venue treatments that touch abliteration or its neighboring techniques. As of Q3 2026, no topic-specific survey of abliteration and no FAccT/AIES paper naming abliteration in title were located; both remain coverage gaps rather than confirmed absences.
- Bartoszcze, L., et al. (2025). Representation Engineering for Large Language Models: A Survey. arXiv:2502.17601. Broad survey of representation engineering; touches abliteration as one method within a wider family. Does not adopt the M1-M8 classifier used in this wiki; included with that disagreement noted, per the standard that M1-M8 is one editorial taxonomy among several in the field.
- Casper, S., et al. (2024). Black-Box Access is Insufficient for Rigorous AI Audits. FAccT 2024. arXiv:2401.14446 · DOI 10.1145/3630106.3659037. Closest verified ethics-venue item; does not engage abliteration by name but argues for white-box access in the audit context to which the M1 line belongs. Included as adjacent context, not as a paper substantively engaging abliteration in the strict sense.
Benchmarks and evaluation
- Chiang, W.-L., et al. (2024). Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference. arXiv:2403.04132. The LM Arena paper.
- Singh, S., et al. (2025). The Leaderboard Illusion. arXiv:2504.20879. Critique of LM Arena gameability by well-resourced labs.
- Hugging Face Open LLM Leaderboard team. (2024). Open-LLM performances are plateauing, let's make the leaderboard steep again. HF Spaces blog. Announcement of Open LLM Leaderboard v2 and its six harder benchmarks.
Blog posts and long-form sources
Non-peer-reviewed but foundational texts.
- Hartford, E. (2023). Uncensored Models. erichartford.com/uncensored-models. Founding manifesto for the uncensored-fine-tune tradition. Political-not-technical argument.
- Hartford, E. (2023). Dolphin announcement. erichartford.com/dolphin. Documents the data-filtering method that predates and shades into M6.
- Labonne, M. (2024). Uncensor any LLM with abliteration. mlabonne.github.io · HF blog mirror · Colab notebook. The canonical M4 pipeline tutorial. Established the two-stage recipe and the "healing" vocabulary. Discussed extensively in M4 hybrid and DPO healing.
- Labonne, M. Merge Large Language Models with mergekit. towardsdatascience.com. The practitioner tutorial for mergekit. Reviewed by Goddard.
- grimjim (Jim Lai). Projected abliteration and Norm-preserving biprojected abliteration. Projected abliteration HF blog · Norm-preserving biprojected HF blog. Two M1 variants that reduce collateral damage. Cited in Heretic's docs.
Software - tools and frameworks
Abliteration tools
- github.com/p-e-w/heretic - Weidmann's automated abliteration tool. See Heretic - automated abliteration. PyPI: heretic-llm.
- github.com/andyrdt/refusal_direction - Arditi's reference implementation.
- github.com/FailSpy/abliterator - FailSpy's original abliteration toolkit. Where the term was coined.
- github.com/Sumandora/remove-refusals-with-transformers - The tool whose forks became the M2 lineage.
- github.com/jim-plus/llm-abliteration - grimjim / jim-plus's fork with candid README on abliteration's limits.
Merging
- github.com/arcee-ai/mergekit - The standard model-merging tool (Goddard, Arcee AI). License: BUSL-1.1.
- mergekit MoE documentation - The
mergekit-moeconfig format DavidAU's builds use.
Training
- github.com/huggingface/trl - Hugging Face's reference training library. TRL v1.0 (April 2026) unified SFT, DPO, ORPO, KTO, GRPO trainers.
- github.com/unslothai/unsloth - Fastest single-GPU fine-tuning framework (2x speed, ~60% less VRAM).
- github.com/axolotl-ai-cloud/axolotl - Reproducible multi-GPU training via YAML configs. Preferred by roleplay producers.
- github.com/hiyouga/LLaMA-Factory - No-YAML web UI for experimentation.
Interpretability
- github.com/TransformerLensOrg/TransformerLens - Neel Nanda et al. Used in Arditi's original implementation to hook into transformer internals.
Inference and quantization
- github.com/ggml-org/llama.cpp - Georgi Gerganov's reference implementation. Home of GGUF,
llama-quantize,llama-imatrix. - GGUF format documentation.
- llama.cpp discussions - Community source for Q4_K_M sweet-spot claim and quant-quality tradeoffs.
- github.com/turboderp-org/exllamav2 - ExLlamaV2, home of the EXL2 format.
- github.com/vllm-project/vllm - vLLM inference server (consumes AWQ/GPTQ).
- github.com/ml-explore/mlx - Apple's MLX framework for Apple Silicon.
Consumption UIs
- github.com/SillyTavern/SillyTavern - Dominant roleplay-model consumption UI.
- github.com/LostRuins/koboldcpp - Popular GGUF back-end.
Evaluation harnesses
- github.com/EleutherAI/lm-evaluation-harness - Community-standard capability benchmark harness (MMLU, GSM8K, TruthfulQA, HellaSwag, ARC).
- optuna.org - Hyperparameter optimization library. Used inside Heretic for its search loop.
Datasets
Refusal measurement
- walledai/AdvBench - Zou et al. 2023, 520 harmful prompts.
- mlabonne/harmful_behaviors - Labonne's repackaging of AdvBench's
harmful_behaviorssplit. - tatsu-lab/alpaca - Canonical harmless-prompt counterpart used in activation-difference extraction.
Training
- unalignment/toxic-dpo-v0.2 - Canonical M6 preference-pair dataset. Flagged sensitive.
- mlabonne/orpo-dpo-mix-40k - Labonne's general-purpose preference mix used in NeuralDaredevil-8B's healing pass.
- PygmalionAI/PIPPA - 1M+ utterances / 26K Character.AI conversations / 1000+ personas. Foundation of M7.
- lemonilia/LimaRP - Manually-curated novel-style roleplay dataset.
Representative models
Not exhaustive - just the models named across the wiki as canonical or illustrative examples.
M1 / M3
- failspy/Llama-3-8B-Instruct-MopeyMule - FailSpy's early release, contemporaneous with coining "abliteration."
- failspy/llama-3-70B-Instruct-abliterated - Higher-parameter FailSpy release cited as M1 exemplar.
- huihui-ai/Qwen3-8B-abliterated - Referenced as M2 documented-iteration example (v1 later superseded by improved v2).
- huihui-ai/Huihui-Qwen3.8-27B-abliterated - M3 canonical example.
- huihui-ai/Ornith-1.5-9B-Instruct-abliterated - Additional M3 release.
- grimjim/gemma-3-12b-it-norm-preserved-biprojected-abliterated - grimjim's norm-preserving variant applied to Gemma-3-12B-IT.
M4 (abliterate + heal)
- mlabonne/NeuralDaredevil-8B-abliterated - The canonical M4 output. Reference for the whole two-stage recipe.
- mlabonne/Daredevil-8B-abliterated - Intermediate artifact before healing.
- mlabonne/Daredevil-8B - The DARE-TIES merge that Daredevil-8B-abliterated started from.
M5 (merges, MoE)
- huggingface.co/DavidAU - Producer profile.
- Dark Champion 8X3B MoE - Reference M5 mixture-of-experts release.
M6 (uncensored fine-tune)
- dphn/dolphin-2.9-llama3-8b - Hartford's Dolphin lineage, current release.
M7 (roleplay, stacked)
- jwest33/gemma-3-4b-null-space-abliterated-RP-writer - Canonical M7-stacked-on-M1 example.
M8 (GGUF quantization)
- huihui-ai/Huihui-Qwen3.8-27B-abliterated-GGUF - Source of the
_Lquantization trick. - mradermacher's DeepSeek-V4-Flash-Abliterated-FP8-i1-GGUF - Representative mradermacher release.
Producer profiles on Hugging Face
- huihui-ai - Highest-volume abliteration producer. See huihui-ai models explained.
- mradermacher - 68,000+ GGUF repositories. De facto distribution infrastructure.
- bartowski - High-quality imatrix quants.
- mlabonne - Maxime Labonne's profile - source of NeuralDaredevil-8B and the canonical M4 recipe.
- TheDrummer (BeaverAI) - Cydonia, Rocinante, Gemmasutra roleplay lines.
- PygmalionAI - Institutional home of M7 roleplay tradition.
Benchmarks and leaderboards
- UGI Leaderboard - DontPlanToEnd's Uncensored General Intelligence board. Only major board built for this class of model.
- LM Arena - Blind pairwise-comparison leaderboard. Formerly Chatbot Arena.
- Open LLM Leaderboard v2 (archived) - Hugging Face's static snapshot as of March 2025.
- Abliterlitics - Community project doing comparative measurements of abliteration techniques.
Press coverage and legal analysis
- Irish Times - AI guardrails stripped from Meta and Google models in minutes (25 May 2026). Wire syndication of the Heretic story.
- Futurism - Tools strip AI guardrails in minutes. US press coverage.
- Gigazine - Heretic feature (17 November 2025). Japanese tech press coverage of Heretic's release.
- Lexology - legal-industry analysis. Legal treatment of abliteration in enterprise-model deployment contexts.
Reference materials and infrastructure
- Wiktionary: abliterate - Etymology of the term ("to wipe out entirely," archaic Latin-derived).
- NeurIPS 2024 poster page - Arditi paper session record.
- Business Source License 1.1 text - License currently declared in mergekit's
pyproject.toml. - Liquid AI - Company where Maxime Labonne is a staff research scientist.
- Runpod pricing - Source for the July 2026 GPU rental costs cited across method articles (A100 80GB $1.39/hr, RTX 4090 $0.69/hr).
- Ted Mitew - Cultural theorist whose work on the ontology of AI models informs the wiki's post-moral-models framing.
- Cognitive Computations HF - Eric Hartford's organizational profile.