About this wiki
What this wiki is, what it is not, how it was assembled, and what to do when you find something wrong. The reference of record for a field that has been settling into shape in real time. Updated for Q3 2026 with the refusal-direction material integrated.
A reference wiki on abliteration - the technique that removes refusal behavior from language models by editing their weights. Thirty-five articles across ten clusters covering foundations, methods (M1-M8 plus M9 review track), practical how-to, SEO-anchored landing pages, concept articles, frontiers, market analytics, open questions, and appendices. Sits alongside abliteration.org's catalog of 16,000+ abliterated Hugging Face repositories and the project's quarterly research report track. Written between August 2026 and continuously updated as new report material is integrated. Descriptive not prescriptive - names disputes rather than resolving them, treats abliteration as a technical practice rather than endorsing or opposing it. Corrections welcome.
- What this wiki tries to do and what it deliberately does not
- How the taxonomy was constructed and why it has the boundaries it does
- The register the wiki uses and the register it avoids
- How the wiki was assembled - sources, quarterly reports, structure, timing
- How to read it (five reading paths at wiki index), and how to correct it
- The relationship between the wiki, the catalog, and the quarterly report
- Attribution and how the wiki treats the people it cites
What this wiki is
A reference wiki on abliteration - the technique that removes refusal behavior from language models by editing their weights. Thirty-five articles across ten clusters:
- Foundations - What is abliteration?, The Arditi method, Ontology of post-moral models
- Methods - the eight-method taxonomy (M1-M8) plus the M9 review track, methods overview, Heretic, boundary cases, methodology of labeling, and grimjim's refinements
- Practical - practical how-to with real code and cost numbers, reading benchmarks
- Focused topics - abliterated vs uncensored, what is a refusal direction?, huihui-ai models explained
- Concept and history - DPO healing, Frankenstein's creature, quantizations for humanists, the dimensionality debate, restoration and defense
- Frontiers - reverse-engineering abliteration from weights (the first article in a cluster on missing artifacts and tooling)
- Market analytics - how market metrics work, how the fork graph works, timeline of the field
- Open questions - open questions in refusal geometry. A cluster for substantive questions the literature has raised but not resolved. Each article tracks its questions with binary status (open or resolved) and updates when new work moves any of them.
- Reference and appendices - glossary, references
- Meta - this page
The wiki sits alongside abliteration.org's catalog of 16,000+ abliterated Hugging Face repositories and the project's quarterly research report track. The catalog is the artifact index; the wiki is the reference material that lets a reader make sense of what the catalog contains; the report is the periodic deep-research output that feeds new material into both. All three systems refer to each other by design.
What this wiki is not
Several things it deliberately does not attempt:
- Not a tutorial. The practical articles show real code, but the wiki is not a step-by-step teach-yourself guide. It assumes the reader can follow a Colab notebook or read a Python file.
- Not a benchmark aggregator. We do not track scores across models. That is what UGI, LM Arena, and Abliterlitics do; we point readers at them and explain what they measure.
- Not an endorsement or a warning. The wiki treats abliteration as a technical practice that exists, not as something the reader should adopt or avoid. Producers whose work is documented here range from academic labs to pseudonymous individuals; they are described accurately regardless of the reader's opinion of their choices.
- Not exhaustive. Not every abliterated model or every producer is named. Coverage prioritizes canonical or unusually well-documented examples.
- Not a policy paper. Where legal or safety debates around abliteration are relevant, we cite them (Lexology's analysis, Irish Times coverage, Futurism reporting) but do not stake positions.
How the taxonomy was constructed
The M1-M8 taxonomy came from a two-pass process. First pass: read the field's own vocabulary - the words producers use in model cards, the categories papers propose, the natural clusters in the catalog. Second pass: name what a serious reader would need to distinguish between to understand any given model.
The eight-way split settled where it did because that was where the distinctions stopped being lossy. Six categories collapsed too much (M4's healing pass has to be separable from raw M1 for the benchmark discussion to make sense; M8 has to be a category because most catalog artifacts are quantized). Ten categories over-split (splitting M1 into "with validation" and "without" gave the M2 category we kept but never really settled).
The taxonomy has known soft spots. They are named in the relevant articles rather than hidden:
- Whether M2 should collapse into M1 as a quality gradient rather than remain a peer category.
- Whether M6 belongs in an abliteration taxonomy at all, or whether "uncensored fine-tune" should be strictly separate.
- Whether M8 is a method or a distribution modifier that co-occurs with the other seven.
- Whether M7's roleplay tradition is really the same phenomenon as M6 under different framing.
- Whether the M9 review track is ready to become a full method. Five research groups have shown refusal geometry is richer than a single direction (see the dimensionality debate), but no production tooling exists yet, and no model card in the catalog currently declares itself M9. The review track is open pending three formalization conditions.
These are open. The wiki takes positions - M6 in as adjacent, M8 in as distribution modifier, M2 in for provenance-descriptive use, M9 in review - but flags each as contested and explains the alternative view.
The register
The wiki uses clinical reference register. Concretely:
- Attribution. Named people cited by their preferred forms (mononyms like huihui-ai and DontPlanToEnd stay mononyms; academic authors get full names on first mention, surnames after). Papers cited with arXiv links.
- Philosophical vocabulary allowed but only where load-bearing. When the article's philosophical point is that "healing" is not neutral vocabulary, that point is made explicitly (see DPO healing). Otherwise the wiki uses the field's own terms without commentary.
- No sales voice. No superlatives, no urgency, no calls to action. Producers are described, not celebrated or condemned.
- Short hyphens only. No em-dashes anywhere in the corpus (verified via CI check). This is a specific stylistic choice - em-dashes have become an AI-writing tell in 2026, and their absence signals the text was reviewed by a human.
- Code blocks and configs verbatim where the source used specific values. When Labonne's tutorial uses
lr=5e-6, we uselr=5e-6and cite. When huihui-ai's cards say a specific layer band, we quote it exactly.
Register the wiki does not use: casual voice, first-person narration outside the meta articles (this one), rhetorical questions to the reader, exhortation, or any framing that presumes the reader should agree with the wiki's implicit stance on any question.
How the wiki was assembled
The initial pass was written during August 2026. Since then the wiki is updated continuously as the project's quarterly research report track produces new material. Source materials in current use:
- Primary sources - every paper, blog post, and canonical release cited in the References appendix. Read directly, quoted verbatim where the exact wording matters (Hartford's "no one true correct alignment," Labonne's "1-2% on every benchmark," huihui-ai's "crude proof-of-concept," Arditi's "we show that refusal is mediated by a one-dimensional subspace," Wollschläger's "concept cones").
- Producer model cards - read directly on Hugging Face for the canonical example models. Where a model card describes its own methodology (huihui-ai's layer-band choices, DavidAU's constituent lists, Labonne's healing configs, Heretic's default hyperparameters), that self-description is treated as evidence.
- Community discussion - LessWrong, r/LocalLLaMA, GitHub issue threads, Hugging Face community tabs. Used for context and to identify where the field's practitioners themselves disagree.
- Comparative empirical work - the Young 2026 tool study (arXiv:2512.13655), UGI Leaderboard measurements, and Heretic's own documentation. Used for the empirical claims in the benchmarks and methods articles.
- Press and legal coverage - Irish Times, Futurism, Gigazine, Lexology. Used to represent how the field is described outside itself.
- Quarterly research reports - the project runs a quarterly research report track. The Q3 2026 cycle covered timeline, terminology, the M1-M8 methods classifier, the academic corpus of 27 verified sources, and the refusal direction; its outputs feed most of the material in the concept, methods, and frontiers clusters. Each report cycle produces substantial reference documents from which the wiki extracts material and links back to primary sources.
- Producer Ecosystem Map - a separate deep-research artifact identifying producers, researchers, quantizers, tier by activity, and organizational versus individual attribution. Feeds the acknowledgements section, the author roles in the catalog, and the impersonation-warning system.
- The project's own classifier and market data - the M1-M8 method classifier and the seven-method extraction classifier (both running hourly against the catalog DB) produce the empirical claims about production distribution in articles like what is a refusal direction? § theory vs practice. The classifier code and its rules are in the project's
scripts/directory.
The order of writing was foundations → methods → practical → SEO → concept → frontiers → market analytics → appendices → meta (this page). The meta page is refreshed with each substantial content wave; the current refresh integrates the refusal-direction material.
How to read the wiki
The wiki index page offers five reading paths curated for different readers, each about an hour or under:
- For journalists (30 min) - the fastest orientation. Enough to write about abliteration accurately, not confuse it with jailbreaks, and read a benchmark table without being deceived.
- For philosophers (65 min) - what refusal is, how it is removed, and what a model without it becomes. Assumes familiarity with the basic idea.
- For practitioners (65 min) - from setup to a running abliterated model of your own, then verify what you produced against real benchmarks.
- For researchers (58 min) - the academic corpus and the M1-M8 taxonomy from a research angle: sources, classification methodology, boundary cases, review track, and open frontiers.
- For academic groups (60 min) - the citable-reference track for a research lab or thesis committee that will cite abliteration.org. Every claim ties to an arXiv or peer-reviewed primary source.
Beyond the curated paths: every article has a Related section at the bottom pointing to five natural next reads. The glossary and references are the reference-of-record for cross-lookup. The timeline orders everything chronologically.
Attribution and how the wiki treats the people it cites
Named individuals are treated as public figures in the technical sense: identified by the name they publish under, cited when their work is discussed, quoted where their exact wording matters. This includes pseudonymous producers (huihui-ai, DontPlanToEnd, TheDrummer, grimjim, mradermacher, bartowski, FailSpy) - their pseudonyms are the identity they chose for their work, and referring to them by that name is respect for that choice.
The wiki does not attempt to unmask pseudonymous producers or link their work to real-name identities. Where a person has published under both a pseudonym and a real name (grimjim / Jim Lai), both are noted because the person has done so themselves. Where they have not, the pseudonym stands.
Academic authors are cited by full name on first mention and surname after (Arditi et al., Labonne, Wollschläger, Hartford, Goddard, Rafailov). Their papers are linked directly to arXiv or the publication venue.
Companies and organizations are cited by name (Arcee AI, Liquid AI, Hugging Face, PygmalionAI). Products they maintain are cited without brand-marketing language.
How to correct the wiki
The wiki lives on abliteration.org. Corrections happen through the site's contact channels and, where technical, through direct correspondence with abliteration.org maintainers.
Categories of correction that are welcome:
- Factual errors - dates, quote wording, author attribution, model release information, benchmark numbers. Cite the source that shows the correct fact and we will update.
- Broken links - papers moved to different venues, HF repositories deleted or renamed, blog posts taken down. The wiki will update to the current authoritative location, or note when a source has become inaccessible.
- Missing information - producers or techniques we should have named but did not, disputes we should have flagged but did not. The taxonomy is not exhaustive by design, but we prefer to know what we missed.
- Vocabulary corrections - if a producer prefers a different framing for their own work (e.g. rejects being called an "M2 producer"), that preference will be noted in the article and future revisions will use the preferred term or explicitly bracket both.
Categories of change that will require discussion rather than automatic acceptance:
- Taxonomic reorganization - collapsing M2 into M1, splitting M6 and M7, promoting M8 out of the method taxonomy. These are load-bearing choices; changing them means restructuring cross-references throughout the wiki.
- Register shifts - moving toward more advocacy tone, adding safety warnings, adopting first-person voice. The current clinical register is deliberate.
- Coverage expansion - adding new method categories, new SEO landing pages, new concept articles. Welcome in principle but need to be scoped so the wiki does not become unmaintainable.
The relationship to the catalog and the quarterly report
The catalog is the artifact index - 16,000+ Hugging Face repositories tagged as abliterated, browsable and searchable. Each catalog entry links to the underlying Hugging Face page and displays basic metadata (parameter count, quantization variants, downloads, apparent method).
The wiki is the reference material. When a catalog entry says "M4" or "GGUF Q4_K_M" or "huihui-ai layer 23-28," the reader can click through to the corresponding wiki article to see what those terms mean, how the technique works, and what to expect from a model built that way.
The catalog is updated hourly from a systemd collector that pulls new HF releases matching abliteration-related tags. As of Q3 2026 the M1-M8 method classifier and the seven-method extraction classifier also run hourly against the DB, so new models are labeled shortly after they appear. The wiki itself is updated when the field's state changes enough to require it (a new taxonomic dispute settled, a new canonical technique released, factual corrections) or when a Session of the quarterly report produces new integrable material. The three systems (catalog, wiki, quarterly report) are decoupled at the storage level but refer to each other by design: the catalog cites wiki articles for method explanations, the wiki cites classifier data for empirical distribution claims, and the quarterly report feeds new material into both.
What comes next for the wiki
The wiki now covers foundations, methods (including M9 review), practical, focused topics, concept/history, frontiers, market analytics, and appendices - the initial scaffolding is complete. Ongoing work:
Continuous integration from the quarterly research report. The Q3 2026 report cycle is complete and its material is integrated. Subsequent cycles will continue to feed the wiki, with the next focus area covering people and organizations (which will primarily expand the acknowledgements section, the producer descriptions in method articles, and the fork graph annotations).
Classifier and analytics improvements. The current M1-M8 method classifier has known blind spots (see boundary cases). A v2 that handles merge-and-abliterate combinations more cleanly, distinguishes projected from standard orthogonalization at scale, and catches restoration operations is planned. The extraction-methods classifier currently covers about 35% of the catalog; improving coverage is ongoing.
Market analytics widgets. The catalog's market page shows a small number of widgets today (extraction distribution, model of the month, top forks). More are planned as data quality improves: quality-tier distribution, producer concentration, family lineage timelines, geographic distribution of production.
The M9 review track. When production tooling emerges (a Heretic-style one-command runner for multi-directional ablation) and released models begin declaring the method in their model cards, M9 will be promoted from review track to full method. The M9 article documents the three formalization conditions. Not scheduled.
Longer-term possibilities (not committed):
- English Wikipedia article on abliteration - the wiki here can serve as a source for a shorter, editorially-styled Wikipedia entry once the topic has enough independent secondary coverage to satisfy Wikipedia's notability standards.
- Translations. Russian, French, German, and Japanese are natural candidates given where the field's readership sits.
- Companion podcast or long-form video treatments of specific articles.
Acknowledgements
This wiki exists because a large number of people did the underlying work. Explicitly acknowledged:
The originating papers and their authors:
- Andy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka, Nina Panickssery, Wes Gurnee, and Neel Nanda for the June 2024 paper that made the whole field legible.
- Thomas Marshall, Adam Scherlis, and Nora Belrose (EleutherAI) for the affine-decomposition refinement (November 2024) that opened the geometry debate.
- Tom Wollschläger, Jannes Elstner, Simon Geisler, Vincent Cohen-Addad, Stephan Günnemann, and Johannes Gasteiger for the polyhedral-cone critique (ICML 2025).
- Giacomo Piras, Roberto Mura, Fabio Brau, Luca Oneto, Fabio Roli, and Battista Biggio (PRALab, University of Cagliari) for the SOM manifold analysis (AAAI 2026).
- Thomas Winninger for the RFM-AGOP paper (ICML 2026 workshop) and the widely-cited Qwen 3 three-directions result.
- Fatima Joad, Majd Hawasly, Sabri Boughorbel, Nadir Durrani, and Husrev T. Sencar (Qatar Computing Research Institute) for the eleven-category decomposition that partially reconciles the debate.
The practice-shaping figures:
- Eric Hartford for the founding uncensored-models argument that the abliteration tradition later joined.
- Maxime Labonne for the canonical M4 recipe and the vocabulary the field now uses.
- Philipp Emanuel Weidmann for Heretic and for demonstrating that automated abliteration was tractable.
- Jim Lai (grimjim) for the projected and norm-preserving refinements that became Heretic's default extraction shape.
- Charles Goddard for mergekit and Georgi Gerganov for llama.cpp - the infrastructure without which none of this would be reproducible at scale.
- Neel Nanda separately for TransformerLens, the interpretability toolkit the extraction methods build on.
The restoration and defense authors:
- The ROSI authors for the first published restoration method.
- Abu Shairah and colleagues for Extended Refusal, the first published defense.
- Truong for AMRA (Abliteration Mitigation via Refusal Aliases).
- Agnihotri and colleagues for the safety-pretraining study that reported Qwen 3 abliteration-resistance.
The producers whose work is the substrate the catalog indexes:
- Tier-1: huihui-ai, mradermacher, bartowski, DavidAU, TheDrummer, PygmalionAI, cognitivecomputations, grimjim, mlabonne, p-e-w (Weidmann's HF handle), Sumandora, FailSpy, jwest33, DontPlanToEnd.
- Tier-3 and specialized: SicariusSicariiStuff, Goekdeniz-Guelmez, LuffyTheFox, llmfan46, noctrex - producers with smaller catalogs but distinctive contributions.
Any errors in this wiki are the wiki's, not theirs.
License and republishing
The wiki's prose is intended to be freely referable and quotable. Direct quotation of blocks longer than a paragraph should link back to the article rather than reproducing wholesale. Corrections and derived work are welcome. The wiki cannot license the underlying papers or datasets it cites - those retain their own copyrights and terms.
The wiki does not republish or redistribute any of the datasets, harmful-prompt sets, or preference-training data it discusses. Where datasets are named (AdvBench, toxic-dpo-v0.2, PIPPA), the wiki links to their canonical Hugging Face pages and describes their structure without reproducing contents.