TRACK 1 · MARKET & TECHNICAL REPORT

The Abliteration & Uncensored LLM Report

A neutral quarterly reference on the market and mechanics of refusal removal in open-weight language models. What exists, in numbers, with primary-source citations for every claim.

Cadence
Quarterly
First issue
2026 Q3
Status
In preparation
Length
60-90k words
License
CC BY 4.0

Who this report is for

  • Researchers in mechanistic interpretability, alignment, and open-source machine learning
  • Journalists and reporters covering AI safety, open-weight models, and policy
  • Regulators, national technical bodies, and legal analysts (EU AI Office, UK AISI, US CAISI, NIST)
  • Security teams, threat-intelligence vendors, and red teams working with open-weight models
  • Philosophers of technology and thoughtful general readers interested in what these developments mean

Table of contents (first issue)

Every section shipped so far has a source manifest attached. Sections marked drafting or planned land here as they clear editorial review.

  1. 01
    Industry timeline (Jan 2023 - Aug 2026)
    40 dated events across academic milestones, terminology, model and tool releases, press coverage, and institutional commentary, each with a primary source.
    ready
  2. 02
    Terminology and boundaries
    Where "abliteration" starts and stops. What counts as uncensored, abliterated, jailbroken, roleplay-masked, base-model, and where the words overlap.
    ready
  3. 03
    The M1-M9 methods classifier
    The full editorial taxonomy: single-direction removal, layer-wise ablation, heal-after-ablate, merges, uncensored fine-tuning, quantization repackaging, multi-directional ablation.
    ready
  4. 04
    Academic corpus inventory
    Every peer-reviewed and preprint contribution to refusal geometry, abliteration methodology, and defence research from 2023 through August 2026.
    ready
  5. 05
    The refusal direction: a technical primer
    Historical arc from Turner ActAdd to Winninger RFM-AGOP. Extraction and application methods. The dimensionality debate. Restoration and inversion.
    ready
  6. 06
    FailSpy: origin figure of abliteration as an open-source practice
    The Reddit user who coined the term in May 2024 and released the abliterator library that seeded the field.
    ready
  7. 07
    Maxime Labonne: the popularising tutorial and NeuralDaredevil
    The June 2024 Hugging Face blog that named the practice for the wider community, and the healed-abliteration recipe.
    drafting
  8. 08
    huihui-ai: industrialisation of layer-wise ablation
    The producer behind ~220 abliterated models, including the trillion-parameter Kimi K2 abliteration.
    drafting
  9. 09
    Philipp Emanuel Weidmann and Heretic
    The November 2025 tool that automated abliteration and, by mid-2026, seeded 5,000+ decensored models.
    drafting
  10. 10
    The quantization republishing layer
    bartowski, mradermacher, RichardErkhov, DevQuasar, tensorblock, QuantFactory, mlx-community, and the historical TheBloke.
    drafting
  11. 11
    The merge and roleplay specialists
    DavidAU, Undi95, Sao10K, SicariusSicariiStuff, Goekdeniz-Guelmez, and the M5 merge tradition.
    drafting
  12. 12
    Nous Research, Cognitive Computations, Eric Hartford
    The dataset-based uncensoring lineage that predates activation-based abliteration.
    planned
  13. 13
    Mid-volume producers and the Heretic-native cohort
    llmfan46, Blackfrost-AI, HauhauCS, LuffyTheFox, and other producers born inside the automated era.
    planned
  14. 14
    Producer ecosystem map (2024-2026)
    The four functional tiers of the producer population; naming conventions and README patterns; geographic distribution.
    ready
  15. 15
    Tool profiles: abliterator, remove-refusals-with-transformers, Heretic, OBLITERATUS, ErisForge, DECCP
    Every tool in general use with its methodology, dependencies, and lineage.
    planned
  16. 16
    The Qwen family under abliteration
    The most-abliterated base-model family; the multi-token-prediction corruption trap; the huihui-ai reference recipes.
    planned
  17. 17
    Llama, Gemma, Mistral, DeepSeek, Kimi, GLM, Phi
    One chapter per base-model family: what makes it easy or hard, who ships variants, what benchmarks say.
    planned
  18. 18
    Multimodal and mixture-of-experts abliterations
    Vision-language variants, MoE architectures, and the M9 multi-directional research frontier.
    planned
  19. 19
    Regional communities
    Chinese, Russian, Japanese, Korean practitioner traditions, from huihui-ai to the CIS layer-window heuristic.
    planned
  20. 20
    Fork chain analysis and the disenchantment index
    Six-generation-deep chains, the DAI metric, and the growth curves per base model family.
    drafting
  21. 21
    Historical downloads and market share
    Wayback-backfilled download curves for the top of the catalog; Gini and top-1% concentration by month.
    planned
  22. 22
    The deletion history: what the field has thrown away
    2,900+ models silently removed from Hugging Face, reconstructed from monthly snapshots; average lifespan; largest by lifetime downloads.
    drafting
  23. 23
    Model longevity and author trajectories
    How long the average abliterated model lives, which authors persist across cycles, which vanish.
    planned
  24. 24
    Break velocity: how fast a base gets abliterated
    Days from a base release to its first abliteration, per family; Gemma 4 within 90 minutes, Kimi K2 at trillion scale.
    planned
  25. 25
    Three case studies
    One M3 huihui build, one Heretic-automated run, one M5 merge; the full recipe, benchmark deltas, and downstream impact.
    planned
  26. 26
    Refusal restoration index (preview)
    How much of the removed capability comes back under adversarial prompts; where the ceiling is.
    planned
  27. 27
    Community favourites and cross-community reception
    r/LocalLLaMA sticky recommendations, UGI Leaderboard trajectories, most-forked bases and derivatives.
    planned
  28. 28
    Financial Times investigation, May 2026: a deep read
    The FT/Alice joint investigation and what it recorded verbatim; the responses that followed.
    planned
  29. 29
    EU DSA landscape
    How the Digital Services Act treats open-weight abliteration; documented enforcement or absence thereof.
    planned
  30. 30
    US, UK, China: three regulatory grammars
    Section 230, Online Safety Act, CAC Interim Measures. Where each names abliteration and where it does not.
    planned
  31. 31
    Provider responses: Google, Meta, Anthropic
    On the record, from statement to statement, on the specific question of downstream abliteration.
    planned
  32. 32
    The academic ethics discussion
    Peer-reviewed and preprint takes on whether the technique should be publishable, teachable, or restricted.
    planned
  33. 33
    Community ethics: the practitioners argue with themselves
    The internal debates on r/LocalLLaMA, HuggingFace discussions, and producer blogs.
    planned
  34. 34
    Reverse-engineering abliteration from weights alone
    The current state of the "which method was used" question. What Hurtado and Messenger get right; where the classifier is unbuilt.
    ready
  35. 35
    Cross-family comparison
    One table across every base family: how many variants, which methods dominate, what benchmarks fall.
    planned
  36. 36
    Methodology appendix
    The Report Methodology document (v1.0, 2026-08-28) plus every session-specific override.
    ready
  37. 37
    Sources appendix and bibliography
    Every primary source, cited exactly once, in one consolidated bibliography with archive.org backups where available.
    planned
  38. 38
    Executive summary
    A short first-page summary written last, once the whole report has answered its own questions.
    planned

Methodology

The report follows a versioned methodology document (v1.0, 2026-08-28). Every analytical section follows the same three-part shape: a factual body with primary-source citations for every claim, a technical detail layer where precision matters, and an interpretive closing paragraph that observes what the data suggests without editorialising, ranking, or recommending.

Short hyphens only, no em-dashes. No promotional adjectives. No unsupported claims: every factual claim carries a citation to a primary source or a peer-reviewed secondary source; where a claim cannot be sourced, it is marked "Not publicly documented" and moved past. Direct quotes reproduce exact language of policies, statements, or key findings with full attribution.

Citation

Each issue is dated, versioned, and preserved at its issue-specific URL. Once the first issue lands, a citation block with BibTeX, APA, and Chicago is published alongside it, and a DOI is minted via Zenodo. Where a claim references a specific edition, cite the versioned URL; where it references the track as a whole, cite this page.

Past issues

No prior issues. This is the first cycle.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in app" button that hands off directly to a local runtime of your choice - Infrahuman, LM Studio, or Ollama. No API keys, no subscription, no prompt leakage.