The full timeline
Every dated event we could verify from a primary source, from Eric Hartford's uncensored-models essay in May 2023 to the West Point counter-terrorism commentary in July 2026. 36 entries grouped by quarter. Where a source is behind a paywall or single-sourced, the entry says so. The timeline grows with each new quarterly research report.
For the narrative reading of the same arc - what the sequence means, how the field changed - see the wiki article on the history of abliteration.
2023 · Q2
- blog post
Essay arguing users should choose their own model alignment and documenting fine-tuning-based uncensored WizardLM models.
authors Eric Hartford models WizardLM-7B-Uncensored · WizardLM-13B-Uncensored · WizardLM-30B-UncensoredPage timestamp 'Updated May 22, 2023'; some secondary sources say 15 May 2023.
2023 · Q4
- blog post
Report undoing Llama 2-Chat 70B safety training with LoRA for under $200 on one GPU.
authors Simon Lermen · Jeffrey Ladish - preprint
Precursor paper establishing that safety fine-tuning is not robust when weights are public; later at ICLR 2024 workshop.
authors Simon Lermen · Charlie Rogers-Smith · Jeffrey Ladish - preprint
Introduces Contrastive Activation Addition, computing steering vectors from residual-stream activation differences; direct precursor to refusal-direction work.
authors Nina Rimsky · Nick Gabrieli · Julian Schulz · Meg Tong · Evan Hubinger · Alexander Matt TurnerPublished at ACL 2024.
2024 · Q2
- community
Informal write-up previewing the paper; states refusal is mediated by a single residual-stream direction.
authors Andy Arditi · Oscar Obeso · Aaquib Syed · Neel Nanda · Wes GurneePost update note records arXiv availability 18 June 2024. - model release ◌ single-sourced
Among the earliest dated FailSpy Llama-3 abliterations, per HF upload commits.
authors FailSpy tools abliteratorDate from HF commit history. - terminology
Reddit post explaining 'ablated + obliterated = abliterated'; first documented public use as a method name; also documents Phi-3 abliterations.
authors FailSpy models Phi-3-mini-128k-instruct-abliterated · Phi-3-vision-128k-instruct-abliteratedCited by Wiktionary; corroborated by HF model-card commits dated 28 May 2024. - model release ◌ single-sourced
Abliterated Daredevil-8B healed with DPO to recover benchmark performance.
authors Maxime Labonne models NeuralDaredevil-8B · Daredevil-8B - tool release ◌ single-sourced
Pure HF Transformers proof-of-concept for refusal removal; base for deccp and huihui-ai.
authors SumandoraExact first-commit date not publicly pinned; predates 9 June 2024 deccp release. - model release
Qwen2-7B-Instruct-deccp plus deccp dataset, built on Sumandora's code; reduces refusals from ~100% to ~20%.
authors Leonard Lin (augmxnt) tools remove-refusals-with-transformers · DECCP models Qwen2-7B-Instruct-deccpCovered by Simon Willison 9 June 2024. - blog post
Hugging Face blog tutorial (proofread by FailSpy) that popularizes abliteration with a Colab notebook.
authors Maxime Labonne · FailSpy models NeuralDaredevil-8BHF blog 'Published June 13, 2024'; personal site dates 4 June 2024. - preprint
Demonstrates the finding across 13 open chat models up to 72B and proposes a rank-one weight edit disabling refusal.
authors Andy Arditi · Oscar Obeso · Aaquib Syed · Daniel Paleka · Nina Panickssery · Wes Gurnee · Neel Nanda
2024 · Q3
- academic
CAA appears in ACL 2024 proceedings, pages 15504-15522.
authors Nina Rimsky · Alexander Matt Turner
2024 · Q4
- preprint ◌ single-sourced
Extends single-direction account; argues refusal is an affine function of activations.
- academic
Refusal-direction paper appears in NeurIPS 2024 proceedings.
authors Andy Arditi · Neel Nanda - model release
Early huihui-ai abliteration built on Sumandora's code; start of high-volume publishing program.
authors huihui-ai tools remove-refusals-with-transformersFeatherless lists 11 Dec 2024; README commits 12 Dec 2024.
2025 · Q1
- tool release
Python library for layer ablation/addition, built on Sumandora's repo; Show HN January 2025.
authors Tsadoq tools remove-refusals-with-transformersExact release day not pinned. - model release ◌ single-sourced
Abliterated DeepSeek-R1 distilled models after R1's 20 Jan 2025 release.
authors huihui-ai models DeepSeek-R1-Distill-Qwen-32B-abliterated · DeepSeek-R1-Distill-Llama-8B-abliteratedQwen-32B 22 Jan; Llama-8B 23 Jan 2025. - preprint
Gradient-based approach revealing multiple independent refusal directions and concept cones.
authors Tom Wollschläger · Jannes Elstner · Simon Geisler · Vincent Cohen-Addad · Stephan Günnemann · Johannes Gasteigerv2 revised 8 Feb 2026.
2025 · Q2
- preprint
Extended-refusal fine-tuning keeps refusal-rate drop under abliteration to at most 10% vs 70-80% baseline.
authors Harethah Abu Shairah · Hasan Abed Al Kader Hammoud · Bernard Ghanem · George Turkiyyah models Llama-2-7B-Chat · Qwen2.5-InstructKAUST; v2 7 Oct 2025.
2025 · Q3
- academic
Wollschläger et al. paper presented at ICML 2025.
authors Tom Wollschläger - community ◌ single-sourced
HF discussion: 'we are already trying, it's a bit challenging, and the project is still ongoing.'
authors huihui-ai models Kimi K2Kimi K2: 1T total / 32B activated parameters.
2025 · Q4
- preprint
Evaluates which data-centric safety components survive abliteration across SmolLM2-1.7B checkpoints; NeurIPS 2025 Lock-LLM workshop.
authors Shashank Agnihotri · Jonas Jakubassa · Priyam Dey · Sachin Goyal · Bernt Schiele · Venkatesh Babu Radhakrishnan · Margret Keuper models SmolLM2-1.7B - model release ◌ single-sourced
Publishes Huihui-Kimi-K2-Instruct-0905-BF16-abliterated-GGUF (1T total / 32B activated params).
authors huihui-ai models Kimi K2Exact day not pinned; completed Sept-Nov 2025. - tool release
Automatic abliteration tool combining directional ablation with TPE/Optuna parameter optimization.
authors Philipp Emanuel Weidmann tools HereticCommit b3545e4. - preprint
Evaluates Heretic, DECCP, ErisForge, FailSpy across 16 models; tool compatibility 16/11/9/5.
authors Richard J. Young tools Heretic · DECCP · ErisForge · abliteratorUNLV.
2026 · Q1
- preprint ◌ single-sourced
Revised version submitted to arXiv.
authors Tom Wollschläger - tool release ◌ single-sourced
Adds LoRA engine with 4-bit quantization support.
authors Philipp Emanuel Weidmann tools Heretic - community ◌ single-sourced
Dated blog coverage reports Heretic as #1 trending repository.
authors Philipp Emanuel Weidmann tools HereticSecondary-sourced; GitHub does not archive trending rankings. - tool release
SVD-based abliteration toolkit with Gradio UI, CLI, Python API and telemetry; 1,000 stars in 24 hours.
authors Pliny the Liberator (elder_plinius) tools OBLITERATUSX announcement timestamped 4 Mar 2026.
2026 · Q2
- press
FT/Alice investigation: Llama 3.3 stripped in under 10 minutes; Gemma 4 within 90 minutes of release.
tools Heretic models Llama 3.3 · Gemma 3 · Gemma 4FT paywalled; facts via Irish Times, eWeek, Futurism. - press ◌ single-sourced
Reports Heretic produced 3,500+ decensored models downloaded 13 million times (per Weidmann to FT).
authors Philipp Emanuel Weidmann tools Heretic
2026 · Q3
- regulatory
CTC Sentinel Vol.19 Issue 7 commentary by Adam Hadley on the open-weight abliteration problem.
authors Adam Hadley models Kimi K2 · Kimi K3Institutional/counter-terrorism commentary, not a binding regulatory action. - model release ◌ single-sourced
Kimi K3 served from 16 July 2026 per West Point CTC commentary.
models Kimi K3 - press
6,644 distinct guardrail-free HF models downloaded 22M+ times in 30 days; via Business Wire, Forbes.
15,865 uploads merged to 6,644 distinct; downloads as of 4 July 2026. - model release ◌ single-sourced
Open weights published 27 July 2026 following UK AISI / US CAISI evaluation.
models Kimi K3