36
dated events
12
quarters covered
9
event types
23
multi-sourced
TYPE
YEAR

2023 · Q2

  1. blog post

    Essay arguing users should choose their own model alignment and documenting fine-tuning-based uncensored WizardLM models.

    authors Eric Hartford models WizardLM-7B-Uncensored · WizardLM-13B-Uncensored · WizardLM-30B-Uncensored
    Page timestamp 'Updated May 22, 2023'; some secondary sources say 15 May 2023.

2023 · Q4

  1. blog post

    Report undoing Llama 2-Chat 70B safety training with LoRA for under $200 on one GPU.

    authors Simon Lermen · Jeffrey Ladish
  2. preprint

    Precursor paper establishing that safety fine-tuning is not robust when weights are public; later at ICLR 2024 workshop.

    authors Simon Lermen · Charlie Rogers-Smith · Jeffrey Ladish
  3. preprint

    Introduces Contrastive Activation Addition, computing steering vectors from residual-stream activation differences; direct precursor to refusal-direction work.

    authors Nina Rimsky · Nick Gabrieli · Julian Schulz · Meg Tong · Evan Hubinger · Alexander Matt Turner
    Published at ACL 2024.

2024 · Q2

  1. community

    Informal write-up previewing the paper; states refusal is mediated by a single residual-stream direction.

    authors Andy Arditi · Oscar Obeso · Aaquib Syed · Neel Nanda · Wes Gurnee
    Post update note records arXiv availability 18 June 2024.
  2. model release ◌ single-sourced

    Among the earliest dated FailSpy Llama-3 abliterations, per HF upload commits.

    authors FailSpy tools abliterator
    Date from HF commit history.
  3. terminology

    Reddit post explaining 'ablated + obliterated = abliterated'; first documented public use as a method name; also documents Phi-3 abliterations.

    authors FailSpy models Phi-3-mini-128k-instruct-abliterated · Phi-3-vision-128k-instruct-abliterated
    Cited by Wiktionary; corroborated by HF model-card commits dated 28 May 2024.
  4. model release ◌ single-sourced

    Abliterated Daredevil-8B healed with DPO to recover benchmark performance.

    authors Maxime Labonne models NeuralDaredevil-8B · Daredevil-8B
  5. tool release ◌ single-sourced

    Pure HF Transformers proof-of-concept for refusal removal; base for deccp and huihui-ai.

    authors Sumandora
    Exact first-commit date not publicly pinned; predates 9 June 2024 deccp release.
  6. model release

    Qwen2-7B-Instruct-deccp plus deccp dataset, built on Sumandora's code; reduces refusals from ~100% to ~20%.

    authors Leonard Lin (augmxnt) tools remove-refusals-with-transformers · DECCP models Qwen2-7B-Instruct-deccp
    Covered by Simon Willison 9 June 2024.
  7. blog post

    Hugging Face blog tutorial (proofread by FailSpy) that popularizes abliteration with a Colab notebook.

    authors Maxime Labonne · FailSpy models NeuralDaredevil-8B
    HF blog 'Published June 13, 2024'; personal site dates 4 June 2024.
  8. preprint

    Demonstrates the finding across 13 open chat models up to 72B and proposes a rank-one weight edit disabling refusal.

    authors Andy Arditi · Oscar Obeso · Aaquib Syed · Daniel Paleka · Nina Panickssery · Wes Gurnee · Neel Nanda

2024 · Q3

  1. academic

    CAA appears in ACL 2024 proceedings, pages 15504-15522.

    authors Nina Rimsky · Alexander Matt Turner

2024 · Q4

  1. preprint ◌ single-sourced

    Extends single-direction account; argues refusal is an affine function of activations.

  2. academic

    Refusal-direction paper appears in NeurIPS 2024 proceedings.

    authors Andy Arditi · Neel Nanda
  3. model release

    Early huihui-ai abliteration built on Sumandora's code; start of high-volume publishing program.

    authors huihui-ai tools remove-refusals-with-transformers
    Featherless lists 11 Dec 2024; README commits 12 Dec 2024.

2025 · Q1

  1. tool release

    Python library for layer ablation/addition, built on Sumandora's repo; Show HN January 2025.

    authors Tsadoq tools remove-refusals-with-transformers
    Exact release day not pinned.
  2. model release ◌ single-sourced

    Abliterated DeepSeek-R1 distilled models after R1's 20 Jan 2025 release.

    authors huihui-ai models DeepSeek-R1-Distill-Qwen-32B-abliterated · DeepSeek-R1-Distill-Llama-8B-abliterated
    Qwen-32B 22 Jan; Llama-8B 23 Jan 2025.
  3. preprint

    Gradient-based approach revealing multiple independent refusal directions and concept cones.

    authors Tom Wollschläger · Jannes Elstner · Simon Geisler · Vincent Cohen-Addad · Stephan Günnemann · Johannes Gasteiger
    v2 revised 8 Feb 2026.

2025 · Q2

  1. preprint

    Extended-refusal fine-tuning keeps refusal-rate drop under abliteration to at most 10% vs 70-80% baseline.

    authors Harethah Abu Shairah · Hasan Abed Al Kader Hammoud · Bernard Ghanem · George Turkiyyah models Llama-2-7B-Chat · Qwen2.5-Instruct
    KAUST; v2 7 Oct 2025.

2025 · Q3

  1. academic

    Wollschläger et al. paper presented at ICML 2025.

    authors Tom Wollschläger
  2. community ◌ single-sourced

    HF discussion: 'we are already trying, it's a bit challenging, and the project is still ongoing.'

    authors huihui-ai models Kimi K2
    Kimi K2: 1T total / 32B activated parameters.

2025 · Q4

  1. preprint

    Evaluates which data-centric safety components survive abliteration across SmolLM2-1.7B checkpoints; NeurIPS 2025 Lock-LLM workshop.

    authors Shashank Agnihotri · Jonas Jakubassa · Priyam Dey · Sachin Goyal · Bernt Schiele · Venkatesh Babu Radhakrishnan · Margret Keuper models SmolLM2-1.7B
  2. model release ◌ single-sourced

    Publishes Huihui-Kimi-K2-Instruct-0905-BF16-abliterated-GGUF (1T total / 32B activated params).

    authors huihui-ai models Kimi K2
    Exact day not pinned; completed Sept-Nov 2025.
  3. tool release

    Automatic abliteration tool combining directional ablation with TPE/Optuna parameter optimization.

    authors Philipp Emanuel Weidmann tools Heretic
    Commit b3545e4.
  4. preprint

    Evaluates Heretic, DECCP, ErisForge, FailSpy across 16 models; tool compatibility 16/11/9/5.

    authors Richard J. Young tools Heretic · DECCP · ErisForge · abliterator
    UNLV.

2026 · Q1

  1. preprint ◌ single-sourced

    Revised version submitted to arXiv.

    authors Tom Wollschläger
  2. tool release ◌ single-sourced

    Adds LoRA engine with 4-bit quantization support.

    authors Philipp Emanuel Weidmann tools Heretic
  3. community ◌ single-sourced

    Dated blog coverage reports Heretic as #1 trending repository.

    authors Philipp Emanuel Weidmann tools Heretic
    Secondary-sourced; GitHub does not archive trending rankings.
  4. tool release

    SVD-based abliteration toolkit with Gradio UI, CLI, Python API and telemetry; 1,000 stars in 24 hours.

    authors Pliny the Liberator (elder_plinius) tools OBLITERATUS
    X announcement timestamped 4 Mar 2026.

2026 · Q2

  1. press

    FT/Alice investigation: Llama 3.3 stripped in under 10 minutes; Gemma 4 within 90 minutes of release.

    tools Heretic models Llama 3.3 · Gemma 3 · Gemma 4
    FT paywalled; facts via Irish Times, eWeek, Futurism.
  2. press ◌ single-sourced

    Reports Heretic produced 3,500+ decensored models downloaded 13 million times (per Weidmann to FT).

    authors Philipp Emanuel Weidmann tools Heretic

2026 · Q3

  1. regulatory

    CTC Sentinel Vol.19 Issue 7 commentary by Adam Hadley on the open-weight abliteration problem.

    authors Adam Hadley models Kimi K2 · Kimi K3
    Institutional/counter-terrorism commentary, not a binding regulatory action.
  2. model release ◌ single-sourced

    Kimi K3 served from 16 July 2026 per West Point CTC commentary.

    models Kimi K3
  3. press

    6,644 distinct guardrail-free HF models downloaded 22M+ times in 30 days; via Business Wire, Forbes.

    15,865 uploads merged to 6,644 distinct; downloads as of 4 July 2026.
  4. model release ◌ single-sourced

    Open weights published 27 July 2026 following UK AISI / US CAISI evaluation.

    models Kimi K3
Open in Abliteration