The Abliteration & Uncensored LLM Report
A neutral quarterly reference on the market and mechanics of refusal removal in open-weight language models. What exists, in numbers, with primary-source citations for every claim.
The Abliteration & Uncensored LLM Report
A neutral quarterly reference on the market and mechanics of refusal removal in open-weight language models. What exists, in numbers, with primary-source citations for every claim.
Who this report is for
- Researchers in mechanistic interpretability, alignment, and open-source machine learning
- Journalists and reporters covering AI safety, open-weight models, and policy
- Regulators, national technical bodies, and legal analysts (EU AI Office, UK AISI, US CAISI, NIST)
- Security teams, threat-intelligence vendors, and red teams working with open-weight models
- Philosophers of technology and thoughtful general readers interested in what these developments mean
Table of contents (first issue)
Every section shipped so far has a source manifest attached. Sections marked drafting or planned land here as they clear editorial review.
- 01Industry timeline (Jan 2023 - Aug 2026)40 dated events across academic milestones, terminology, model and tool releases, press coverage, and institutional commentary, each with a primary source.ready
- 02Terminology and boundariesWhere "abliteration" starts and stops. What counts as uncensored, abliterated, jailbroken, roleplay-masked, base-model, and where the words overlap.ready
- 03The M1-M9 methods classifierThe full editorial taxonomy: single-direction removal, layer-wise ablation, heal-after-ablate, merges, uncensored fine-tuning, quantization repackaging, multi-directional ablation.ready
- 04Academic corpus inventoryEvery peer-reviewed and preprint contribution to refusal geometry, abliteration methodology, and defence research from 2023 through August 2026.ready
- 05The refusal direction: a technical primerHistorical arc from Turner ActAdd to Winninger RFM-AGOP. Extraction and application methods. The dimensionality debate. Restoration and inversion.ready
- 06FailSpy: origin figure of abliteration as an open-source practiceThe Reddit user who coined the term in May 2024 and released the abliterator library that seeded the field.ready
- 07Maxime Labonne: the popularising tutorial and NeuralDaredevilThe June 2024 Hugging Face blog that named the practice for the wider community, and the healed-abliteration recipe.drafting
- 08huihui-ai: industrialisation of layer-wise ablationThe producer behind ~220 abliterated models, including the trillion-parameter Kimi K2 abliteration.drafting
- 09Philipp Emanuel Weidmann and HereticThe November 2025 tool that automated abliteration and, by mid-2026, seeded 5,000+ decensored models.drafting
- 10The quantization republishing layerbartowski, mradermacher, RichardErkhov, DevQuasar, tensorblock, QuantFactory, mlx-community, and the historical TheBloke.drafting
- 11The merge and roleplay specialistsDavidAU, Undi95, Sao10K, SicariusSicariiStuff, Goekdeniz-Guelmez, and the M5 merge tradition.drafting
- 12Nous Research, Cognitive Computations, Eric HartfordThe dataset-based uncensoring lineage that predates activation-based abliteration.planned
- 13Mid-volume producers and the Heretic-native cohortllmfan46, Blackfrost-AI, HauhauCS, LuffyTheFox, and other producers born inside the automated era.planned
- 14Producer ecosystem map (2024-2026)The four functional tiers of the producer population; naming conventions and README patterns; geographic distribution.ready
- 15Tool profiles: abliterator, remove-refusals-with-transformers, Heretic, OBLITERATUS, ErisForge, DECCPEvery tool in general use with its methodology, dependencies, and lineage.planned
- 16The Qwen family under abliterationThe most-abliterated base-model family; the multi-token-prediction corruption trap; the huihui-ai reference recipes.planned
- 17Llama, Gemma, Mistral, DeepSeek, Kimi, GLM, PhiOne chapter per base-model family: what makes it easy or hard, who ships variants, what benchmarks say.planned
- 18Multimodal and mixture-of-experts abliterationsVision-language variants, MoE architectures, and the M9 multi-directional research frontier.planned
- 19Regional communitiesChinese, Russian, Japanese, Korean practitioner traditions, from huihui-ai to the CIS layer-window heuristic.planned
- 20Fork chain analysis and the disenchantment indexSix-generation-deep chains, the DAI metric, and the growth curves per base model family.drafting
- 21Historical downloads and market shareWayback-backfilled download curves for the top of the catalog; Gini and top-1% concentration by month.planned
- 22The deletion history: what the field has thrown away2,900+ models silently removed from Hugging Face, reconstructed from monthly snapshots; average lifespan; largest by lifetime downloads.drafting
- 23Model longevity and author trajectoriesHow long the average abliterated model lives, which authors persist across cycles, which vanish.planned
- 24Break velocity: how fast a base gets abliteratedDays from a base release to its first abliteration, per family; Gemma 4 within 90 minutes, Kimi K2 at trillion scale.planned
- 25Three case studiesOne M3 huihui build, one Heretic-automated run, one M5 merge; the full recipe, benchmark deltas, and downstream impact.planned
- 26Refusal restoration index (preview)How much of the removed capability comes back under adversarial prompts; where the ceiling is.planned
- 27Community favourites and cross-community receptionr/LocalLLaMA sticky recommendations, UGI Leaderboard trajectories, most-forked bases and derivatives.planned
- 28Financial Times investigation, May 2026: a deep readThe FT/Alice joint investigation and what it recorded verbatim; the responses that followed.planned
- 29EU DSA landscapeHow the Digital Services Act treats open-weight abliteration; documented enforcement or absence thereof.planned
- 30US, UK, China: three regulatory grammarsSection 230, Online Safety Act, CAC Interim Measures. Where each names abliteration and where it does not.planned
- 31Provider responses: Google, Meta, AnthropicOn the record, from statement to statement, on the specific question of downstream abliteration.planned
- 32The academic ethics discussionPeer-reviewed and preprint takes on whether the technique should be publishable, teachable, or restricted.planned
- 33Community ethics: the practitioners argue with themselvesThe internal debates on r/LocalLLaMA, HuggingFace discussions, and producer blogs.planned
- 34Reverse-engineering abliteration from weights aloneThe current state of the "which method was used" question. What Hurtado and Messenger get right; where the classifier is unbuilt.ready
- 35Cross-family comparisonOne table across every base family: how many variants, which methods dominate, what benchmarks fall.planned
- 36Methodology appendixThe Report Methodology document (v1.0, 2026-08-28) plus every session-specific override.ready
- 37Sources appendix and bibliographyEvery primary source, cited exactly once, in one consolidated bibliography with archive.org backups where available.planned
- 38Executive summaryA short first-page summary written last, once the whole report has answered its own questions.planned
Methodology
The report follows a versioned methodology document (v1.0, 2026-08-28). Every analytical section follows the same three-part shape: a factual body with primary-source citations for every claim, a technical detail layer where precision matters, and an interpretive closing paragraph that observes what the data suggests without editorialising, ranking, or recommending.
Short hyphens only, no em-dashes. No promotional adjectives. No unsupported claims: every factual claim carries a citation to a primary source or a peer-reviewed secondary source; where a claim cannot be sourced, it is marked "Not publicly documented" and moved past. Direct quotes reproduce exact language of policies, statements, or key findings with full attribution.
Citation
Each issue is dated, versioned, and preserved at its issue-specific URL. Once the first issue lands, a citation block with BibTeX, APA, and Chicago is published alongside it, and a DOI is minted via Zenodo. Where a claim references a specific edition, cite the versioned URL; where it references the track as a whole, cite this page.
Past issues
No prior issues. This is the first cycle.