Glossary
Every technical term, tool, dataset, benchmark, person, and organization used across this wiki, one line each. For quick lookup when reading any other article.
A one-line reference glossary for the abliteration wiki. Covers concepts (abliteration, ablation, refusal direction, residual stream), techniques (M1-M8, DPO, SFT, LoRA, QLoRA, merging), tools (mergekit, Heretic, TransformerLens, llama.cpp, TRL, Axolotl, Unsloth), datasets (AdvBench, PIPPA, LimaRP, toxic-dpo-v0.2), benchmarks (UGI, LM Arena, MMLU, GSM8K, TruthfulQA, Open LLM v2), formats (GGUF, safetensors, AWQ, GPTQ, EXL2, MLX), quantization schemes (K-quant, I-quant, imatrix, Q4_K_M decoded), and people (Arditi, Labonne, Weidmann, Wollschläger, Hartford, Goddard, Gerganov, huihui-ai, mradermacher). Alphabetically ordered.
- Every technical concept referenced anywhere in the wiki, defined in one line
- Every tool, framework, and format the practical articles use
- Every dataset and benchmark cited in the measurement articles
- Every person, organization, and Hugging Face producer named
- Every method-taxonomy shortcut (M1-M8, B0-B9) explained
How to use this glossary
Alphabetically ordered. Each entry gives the shortest useful definition and, where relevant, a link to the wiki article that treats the term at length. If you land here from another article, the wiki article you came from is the fuller treatment - the glossary is for quick lookup and cross-reference. When the same idea appears under two names in the field (e.g. "orthogonalization" and "directional ablation"), the entry names both.
A
Abliteration - Removing a model's refusal behavior by editing its weights to erase the "refusal direction." Term coined by FailSpy; technique from Arditi et al. The concept at the center of this wiki. See What is abliteration?.
Ablation - Deleting a component or direction from a model to see or change what it does. The "abl-" root in abliteration.
Activation - The numerical state inside the model as it processes text; the values flowing through the network at a given moment. Refusal directions are extracted from activation differences between harmful and harmless prompts.
AdvBench - walledai/AdvBench. The canonical harmful-prompt set for refusal-rate measurement. Zou et al. 2023, 520 short instruction-style prompts spanning malware, fraud, weapons, disinformation. Used across the field as the refusal benchmark.
Alignment - The training process (RLHF, DPO, etc.) that shapes a model's dispositions toward helpful and cautious responses. Abliteration disturbs the alignment-installed refusal component without redoing the alignment.
Alpaca - tatsu-lab/alpaca, repackaged as mlabonne/harmless_alpaca. The canonical harmless-prompt set used as the counterpart to AdvBench when extracting refusal directions from activation differences.
Arcee AI - Company that maintains mergekit after acquiring Charles Goddard.
Arditi, Andy - Lead author of Refusal in Language Models Is Mediated by a Single Direction (NeurIPS 2024), the paper that established the theoretical basis for M1 abliteration. See The Arditi method.
Attention head - A sub-component of a transformer layer that lets the model relate one token to others in the sequence. Some abliteration methods target specific heads.
AWQ - Activation-aware Weight Quantization. GPU-side 4-bit quantization format for fast server inference (vLLM). Alternative to GGUF for GPU-only deployments.
Axolotl - axolotl. Reproducible multi-GPU training framework via YAML configs. Preferred by roleplay producers and for production DPO runs.
B
B0 through B9 - This wiki's cluster-B article codes for the method articles: B0 methods overview, B1 M1 direct removal, B2 M2 raw weight editing, B3 M3 layer-wise ablation, B4 Heretic (automated M1), B5 M4 hybrid (abliterate + heal), B6 M5 merging, B7 M6 uncensored fine-tuning, B8 M7 custom-dataset fine-tune, B9 M8 repackaging (GGUF).
Base model - The raw next-word predictor before instruction tuning or alignment. Capable but without manners or policy. Abliteration operates on instruct-tuned models, not base models.
bartowski - Prolific Hugging Face M8 producer, particularly known for high-quality imatrix-calibrated GGUF quantizations. Parallel operation to mradermacher.
BBH - Big-Bench Hard. Reasoning benchmark, part of Open LLM Leaderboard v2's six-benchmark suite.
BeaverAI - The producer identity of TheDrummer, maintainer of the Cydonia, Rocinante, and Gemmasutra roleplay model families.
BF16 / F16 - 16-bit brain-floating-point / 16-bit floating-point. The native precision most modern models are trained and released in. Two bytes per weight, so a 70B model in F16 is ~140 GB. See Quantizations for humanists.
bitsandbytes - Python library for loading models in 4-bit or 8-bit precision at inference time. Enables single-GPU fine-tuning of larger models via QLoRA.
Bluemoon - Long-form forum-roleplay dataset used in M7 custom-dataset fine-tuning. Emphasizes multi-thousand-word turns.
Bradley-Terry model - Statistical model for pairwise-comparison data. What LM Arena uses to convert millions of blind head-to-head votes into Elo-style ratings.
BUSL-1.1 - Business Source License 1.1. Restrictive-then-open source license currently declared in mergekit's pyproject.toml. Restricts commercial production use during a set period, then converts to a more permissive license.
C
Chat template - The model-specific prompt format (system message + alternating user/assistant turns) that instruct-tuned models expect. Chat-template fidelity between training and inference is the M7-specific make-or-break technical detail.
Chatbot Arena - Original name of LM Arena. Renamed but often still referenced by the older name.
Chiang, Wei-Lin - Lead author on the LM Arena paper (Chiang et al. 2024). Original LMSYS researcher.
Cones (concept cones) - Wollschläger's critique of the single-direction model: refusal is mediated by a concept cone - a family of related directions - not one axis. Removing one direction leaves others available to still mediate refusal. See What is a refusal direction? § Cones not lines.
Cognitive Computations - Eric Hartford's organizational home for the Dolphin, Samantha, and WizardLM-Uncensored model lines.
Cydonia - TheDrummer's most-downloaded roleplay-model family, based on Mistral. Frequently used as merge ingredient for downstream releases.
D
Daredevil-8B - mlabonne/Daredevil-8B. Labonne's DARE-TIES mega-merge of Llama-3-8B fine-tunes. Serves as the abliteration target for NeuralDaredevil-8B (the canonical M4 pipeline example).
DARE - Drop And REscale. mergekit merging algorithm that randomly prunes and rescales weight deltas before merging. Comes in two flavors: dare_ties (with TIES sign-election) and dare_linear (without).
DavidAU - Hugging Face producer known for large mixture-of-experts merges (Dark Champion, 8X3B series) built via mergekit-moe. Model cards unusually detailed about constituent provenance.
DontPlanToEnd - Pseudonymous maintainer of the UGI Leaderboard. Runs the only benchmark board built specifically for uncensored models.
DPO - Direct Preference Optimization. Rafailov et al. 2023, subtitled Your Language Model is Secretly a Reward Model. Preference training on chosen-vs-rejected answer pairs. Simpler alternative to RLHF. Central to M4 healing and M6 uncensored fine-tuning. See DPO healing.
E
Elo - Rating system originally developed for chess by Arpad Elo. LM Arena uses an Elo-style system (via Bradley-Terry) to convert pairwise-comparison votes into rankings.
EXL2 - ExLlamaV2's variable-bitrate format. Excellent quality-per-bit for single-GPU NVIDIA inference. Preferred by GPU enthusiasts over GGUF when running on NVIDIA hardware only.
F
F16 / BF16 - See B.
Fafuła - Co-author of a follow-on critique to the Arditi paper, alongside Wollschläger. Argued that refusal is mediated by a family of directions, not a single one.
FailSpy - Coined the term "abliteration" in mid-2024. Published early practical abliteration recipes that spread across Hugging Face producers.
Frankenmerge - Informal name for a model built via mergekit's passthrough algorithm - concatenating layers from different models to create a new depth. Typically requires at least a light fine-tune to be coherent.
G
Gerganov, Georgi - Creator of llama.cpp and the GGUF format. The G in GGUF/GGML derives from his initials.
GGUF - "GGML Universal Format." Dominant file format for local model inference. Self-contained single file with quantized weights, tokenizer, and metadata. Created by Gerganov for llama.cpp. Consumed by LM Studio, Ollama, KoboldCpp, text-generation-webui.
Goddard, Charles - Creator of mergekit. GitHub handle cg123. Described by Arcee AI as "an award winning software engineer with a strong track record ... at NASA and Apple." Joined Arcee AI, where mergekit is now maintained.
GPQA - Graduate-Level Google-Proof Q&A. Science benchmark written to resist lookup. Part of Open LLM Leaderboard v2.
GPTQ - GPU-side quantization format for fast server inference. Alternative to AWQ, both used with vLLM.
grimjim (Jim Lai) - Author of projected abliteration and norm-preserving biprojected abliteration variants. See grimjim's HF blog post.
GSM8K - Grade School Math 8K. Multi-step arithmetic benchmark. The single most abliteration-sensitive capability benchmark - drops here are the classic tell that abliteration broke reasoning. Comparative studies report Heretic averaging -7.81 pp GSM8K drop.
H
Harmful behaviors - The AdvBench split reused as mlabonne/harmful_behaviors. 520 prompts, used as one half of the harmful/harmless activation-difference pair when extracting a refusal direction.
Hartford, Eric - Author of the founding uncensored-models manifesto (May 2023). Creator of the Dolphin, Samantha, and WizardLM-Uncensored model lines. Predates abliteration by roughly a year; his data-filtering approach evolved into what this wiki calls M6.
Healing - The M4 second stage: DPO pass over an abliterated model to recover general-capability benchmarks. Term popularized by Labonne. Not neutral vocabulary - see DPO healing § What healing presupposes.
Heretic - Weidmann's automated abliteration tool (November 2025). Wraps M1 in a KL-divergence-optimizing loop that searches for good parameters without human tuning. Industrialized abliteration; 3,500+ models and 13M downloads by early 2026. See Heretic - automated abliteration.
HellaSwag - Common-sense completion benchmark. Part of the older Open LLM Leaderboard v1 suite and lm-eval-harness defaults. Barely affected by abliteration.
huihui-ai - Highest-volume abliteration producer on Hugging Face. Self-describes their pipeline as "a crude, proof-of-concept implementation." Notable for the _L quantization trick that preserves ablation-affected tensors at higher precision. See huihui-ai models explained.
I
IFEval - Instruction-Following Eval. Benchmark measuring how well a model follows specific formatting and content directives. Part of Open LLM Leaderboard v2.
imatrix - Importance matrix. Pre-computed measurement of which weights matter most (via activation statistics on calibration text). Used by llama-quantize to allocate the scarce bit budget where it reduces error most. Significant quality gains at Q4 and below; no inference overhead. mradermacher's "i1" tag marks imatrix-calibrated variants.
Instruct model - A base model further trained via SFT and/or RLHF/DPO to follow instructions and produce assistant-style responses. All abliteration targets are instruct models; you cannot abliterate a base model because base models do not refuse.
IPO - Identity Preference Optimization. DPO variant that adds a regularization term to curb DPO's tendency to overfit hard preference pairs. Azar et al. 2023.
IQ (I-quant) - GGUF quantization scheme that uses an imatrix to decide bit allocation. Notation: IQ3_XXS = importance-weighted 3-bit, extra-extra-small variant. Preferred at low bit counts (3-bit and below) where imatrix helps most.
J
Jailbreak - A prompt that tricks an unmodified model into bypassing its safety training. Temporary and patchable (the model's weights are unchanged), unlike abliteration. See Ontology of post-moral models for the four-way distinction.
jim-plus - Maintainer of llm-abliteration, a Sumandora fork. Its README's frank caveat ("abliteration is not full removal of censorship") is a fast route into understanding what M1 and M2 actually accomplish.
jwest33 - Producer of gemma-3-4b-null-space-abliterated-RP-writer, a canonical example of M7 stacked on M1 (null-space abliteration first, then LoRA fine-tune on LimaRP).
K
K-quant - GGUF quantization scheme that stores weights in blocks with per-block scaling factors. Better quality than the older "0" schemes at the same bit count. Notation: Q4_K_M = 4-bit K-quant, medium variant. The community sweet spot for 7B-70B models.
KL divergence - Kullback-Leibler divergence. Measure of how far a modified model's output distribution has drifted from the original on ordinary prompts. Lower means less collateral damage. Heretic uses it as its optimization objective. The honest adjacent metric no leaderboard tracks.
KoboldCpp - Popular back-end for local GGUF inference, particularly favored in the roleplay community. See also SillyTavern, LM Studio, Ollama, text-generation-webui.
KTO - Kahneman-Tversky Optimization. DPO variant that works from single-response good/bad labels rather than paired comparisons. Ethayarajh et al. 2024.
L
Labonne, Maxime - Author of the canonical M4 pipeline (June 2024 blog post). Staff research scientist at Liquid AI. Producer of NeuralDaredevil-8B, orpo-dpo-mix-40k dataset, and a widely used LLM course. Popularized the "healing" vocabulary.
Lai, Jim - See grimjim.
LimaRP - lemonilia/LimaRP. Manually-curated novel-style roleplay dataset. Used in M7 fine-tunes for prose polish.
Linear merge - mergekit's simplest algorithm - a plain weighted average across models. The "model soup." Less sensitive to bad ingredients than SLERP or TIES.
Liquid AI - Company where Maxime Labonne is a staff research scientist.
llama.cpp - llama.cpp. The reference implementation for local model inference. Created by Georgi Gerganov. Home of the GGUF format and the llama-quantize / llama-imatrix tools.
LLaMA-Factory - LLaMA-Factory. Training framework with a no-YAML web UI. Alternative to Axolotl/Unsloth/TRL for experimentation without setting up training infrastructure.
LM Arena - Formerly Chatbot Arena. Human-preference benchmark via blind pairwise comparison, aggregated into Elo-style ratings. 7M+ votes by 2026. Chiang et al. 2024. No separate track for uncensored models.
lm-evaluation-harness - EleutherAI's lm-eval. Community-standard tool for running capability benchmarks (MMLU, GSM8K, TruthfulQA, HellaSwag, ARC) with reproducible settings. The right tool for measuring an abliterated model against its parent.
LoRA - Low-Rank Adaptation. Lightweight fine-tuning method that trains a small set of extra weights instead of the full model. Merged back into the base after training. Used in healing pipelines and roleplay fine-tunes.
M
M1 through M8 - This wiki's method taxonomy. M1 direct removal (Arditi-style), M2 raw weight editing (ad-hoc), M3 layer-wise ablation (selective bands), M4 hybrid abliterate-then-heal, M5 merging via mergekit, M6 uncensored fine-tuning (DPO), M7 custom-dataset fine-tune (roleplay), M8 repackaging (GGUF). See Methods overview.
MATH - Competition-math benchmark. Hardest tier of math evaluation. Part of Open LLM Leaderboard v2.
Merge - Combining two or more models' weights into one without training. Tools: mergekit. Algorithms: SLERP, TIES, DARE, task_arithmetic, passthrough, linear.
mergekit - arcee-ai/mergekit. Standard model-merging tool. Created by Charles Goddard, now maintained by Arcee AI. ~6.9k GitHub stars. License: currently BUSL-1.1.
mergekit-moe - mergekit's subcommand for building mixture-of-experts models from constituent single-expert models. Used by DavidAU for the Dark Champion / 8X3B series.
Millière, Raphaël - Philosopher of mind at Macquarie University. His work on machine consciousness and the ethics of language models informs the philosophical framing this wiki uses.
Mitew, Teodor - Cultural theorist and philosopher whose work on the ontology of AI models informs the wiki's post-moral-models framing.
MLX - Apple's MLX. Apple Silicon native format for machine learning. Preferred on Mac when running inside the MLX ecosystem.
MMLU - Massive Multitask Language Understanding. Broad knowledge-and-reasoning benchmark, 57 subjects. The most commonly cited capability metric. Abliteration typically shifts MMLU by 0.5-3 points at most; the best techniques stay within 0.3 points.
MMLU-Pro - A harder ten-option version of MMLU. Part of Open LLM Leaderboard v2's successor suite.
MoE (Mixture of Experts) - Model architecture that routes each token to a subset of specialized "expert" sub-networks. DavidAU builds MoE merges from abliterated single-expert models via mergekit-moe.
mradermacher - mradermacher. Highest-volume GGUF quantization producer. 68,000+ Hugging Face repositories. Semi-automated pipeline on nethype GmbH hardware plus a private supercomputer from collaborator nicoboss. De facto distribution infrastructure for the abliterated ecosystem.
MuSR - Multi-Step Reasoning. Narrative reasoning benchmark, part of Open LLM Leaderboard v2.
N
NeuralDaredevil-8B - mlabonne/NeuralDaredevil-8B-abliterated. Labonne's canonical M4 output. Abliterated Daredevil-8B, then healed with DPO on orpo-dpo-mix-40k. The reference release for the whole M4 approach.
Norm-preserving biprojected abliteration - grimjim's M1 variant that preserves weight norms and removes the harmless component from the refusal direction before ablating. Reduces collateral damage. Cited in Heretic's documentation.
Null-space abliteration - M1 variant that operates in the null space of the harmless-prompt activations. Used in some jwest33 releases.
O
Ollama - Popular consumer-facing wrapper for GGUF local inference. Alternative front-end to LM Studio.
Open LLM Leaderboard - Hugging Face's capability scoreboard for open models. V2 launched June 2024 with harder benchmarks (MMLU-Pro, IFEval, BBH, MATH, GPQA, MuSR). Archived March 2025 - historical snapshot, no new scores.
ORPO - Odds Ratio Preference Optimization. DPO variant that folds SFT and preference tuning into one stage, needs no reference model. Hong et al. 2024. Less memory, simpler pipeline than DPO.
Orthogonalization - The mathematical operation at the heart of M1: subtract the projection onto the refusal direction from the residual-writing weight matrices, making the direction orthogonal to (perpendicular to) what the model can produce. Sometimes used as a synonym for "directional ablation."
P
Passthrough - mergekit algorithm that concatenates layers from different models to create a frankenmerge of new depth. No interpolation - each layer comes verbatim from one donor.
PEFT - Parameter-Efficient Fine-Tuning. Hugging Face library implementing LoRA, QLoRA, and other low-parameter-count fine-tuning methods.
PIPPA - PygmalionAI/PIPPA. Personal Interaction Pairs between People and AI. ~1M utterances across 26K Character.AI conversation sessions, 1000+ personas. Gosling et al. 2023. Foundation of the M7 roleplay tradition.
Projected abliteration - grimjim's M1 variant that refines direction selection by projection operations. Reduces collateral damage. Cited alongside norm-preserving biprojected abliteration in Heretic docs.
PygmalionAI - Organizational home of the M7 roleplay tradition. Maintains PIPPA and its own line of roleplay-trained models.
Q
Q4_K_M - GGUF quantization notation: "4-bit, K-quant scheme, medium variant." The community value sweet spot for 7B-70B models. Small enough to run on modest hardware, high enough quality that prose stays coherent. See Quantizations for humanists.
QLoRA - Quantized LoRA. Fine-tuning a 4-bit-loaded base model with a LoRA adapter. Enables single-24GB-GPU fine-tuning of 7-8B models. Standard for M4/M6/M7 on consumer hardware.
Quantization - Storing model weights in fewer bits (2, 3, 4, 5, 8) instead of the native 16, saving memory at some cost to accuracy. Formats: GGUF (dominant local), AWQ/GPTQ (GPU-server), EXL2 (NVIDIA enthusiast), MLX (Apple).
R
Rafailov, Rafael - Lead author of the DPO paper (Rafailov et al. 2023).
Refusal direction - The single axis in the model's residual-stream activation space along which the refuse-or-comply decision is encoded, per Arditi et al. Removing it via orthogonalization is what M1 does. Cones critique argues this is oversimplified for larger models. See What is a refusal direction?.
Residual stream - The shared internal channel through which information flows and accumulates across a transformer's layers. Every layer reads from and writes to the residual stream. Refusal directions live in residual-stream activation space.
RLHF - Reinforcement Learning from Human Feedback. Alignment method that tunes models using human preference rankings via a reward-model and PPO. Installs refusal, among other behaviors. Increasingly replaced by DPO for its simplicity.
Roleplay masks - The fourth category in the ontology of post-moral models: models whose reduced refusal is a side effect of persona training rather than an intentional uncensoring operation. See M7 and Ontology of post-moral models.
S
safetensors - Hugging Face's safe alternative to pickle-based model file formats. Standard for base and fine-tuned model distribution on HF Hub. Converted to GGUF via convert_hf_to_gguf.py for local inference.
Second-order model - A model built from already-modified models, e.g. a merge of an abliterated model with another. Common in M5 (see also DavidAU's MoE merges).
SFT - Supervised Fine-Tuning. Training on labeled example answers via cross-entropy loss. Heavier than preference tuning and, per Labonne, "lobotomizes" instruct-tuned models - which is why M4 uses DPO for healing rather than SFT.
SillyTavern - The dominant consumption UI for roleplay models. Front-end interface; not a trainer. Producers train with M6/M7 toolchain, users consume with SillyTavern (or KoboldCpp, or Oobabooga).
SLERP - Spherical Linear intERPolation. mergekit algorithm for two-model blend along a spherical path between weights. Most common two-model merge method.
Sumandora - Author of remove-refusals-with-transformers. Lightweight abliteration tool that lowered the barrier to entry - many current producers use its forks.
T
Task arithmetic - mergekit algorithm that treats the difference between a fine-tuned model and its base as a task vector. Vectors can be added or subtracted; theoretical basis for arithmetic-style refusal removal via negative weights.
text-generation-webui - Oobabooga's back-end for local model inference. Alternative to KoboldCpp for GGUF consumption.
TheBloke - Predecessor to mradermacher and bartowski as the highest-volume GGUF quantization producer. No longer active but many current quantization pipelines cite TheBloke's README templates.
TheDrummer - See BeaverAI.
TIES - Trim, Elect Sign, Merge. mergekit algorithm that merges many task-specific models by trimming redundant deltas and electing signs where models disagree.
toxic-dpo-v0.2 - unalignment/toxic-dpo-v0.2. Small preference dataset "designed to prompt the model to answer illegal questions," explicitly flagged as sensitive. Canonical M6 training data.
TransformerLens - Interpretability library by Neel Nanda et al. Used in the original Arditi implementation to hook into transformer internals during activation extraction. Some producers (e.g. huihui-ai) explicitly write "without using TransformerLens" to signal their pipeline is looser.
TRL - Transformer Reinforcement Learning. Hugging Face's reference training library. TRL v1.0 (April 2026) unified SFTTrainer, DPOTrainer, ORPOTrainer, KTOTrainer, GRPOTrainer under a common API.
TruthfulQA - Benchmark that rewards refusing to endorse popular misconceptions. The metric that drops most reliably under abliteration (5-11 points typical), partly through genuine capability loss and partly through measurement artifact.
U
UGI Leaderboard - Uncensored General Intelligence Leaderboard. Run by DontPlanToEnd. The only major benchmark board built for uncensored models specifically. Columns: UGI (0-100, knowledge on sensitive topics), W/10 (willingness), NatInt (natural intelligence), Writing, political lean.
Uncensored fine-tune - The second category in the ontology of post-moral models. Model whose refusal was removed through training on data lacking refusals (M6/M7), not through weight surgery. See Abliterated vs uncensored.
Unsloth - unslothai/unsloth. Training framework wrapping TRL with custom kernels. 2x faster training and ~60% less VRAM on single GPU. Default choice for 7-13B QLoRA fine-tuning in 2026.
V
vLLM - vLLM. High-throughput GPU inference server. Consumes AWQ and GPTQ quantized models.
W
W/10 - The willingness column on the UGI Leaderboard (0-10 scale). Measures how readily the model engages rather than deflecting. The column that most directly reflects successful abliteration.
Weidmann, Paul - Author of Heretic. Industrialized abliteration into a one-command tool. See Heretic - automated abliteration.
Wollschläger, Tom - Author of the cones critique of Arditi's single-direction model. Argued refusal is mediated by a concept cone (family of directions), not one axis. See What is a refusal direction? § Cones not lines.
X, Y, Z
Zou, Andy - Lead author of the AdvBench paper (Zou et al. 2023). AdvBench's harmful_behaviors split became the field-standard refusal benchmark despite being originally published for adversarial-attack research.
Symbols and abbreviations
_L suffix - huihui-ai convention for GGUF files where ablation-affected tensors (token_embd, output, ffn_down, ssm_out, attn_output) are preserved at Q8_0 or BF16 while the rest of the model is quantized to the nominal level. Prevents aggressive quantization from partly undoing the abliteration. See M8 § huihui-ai's _L trick.
i1 - mradermacher's tag for imatrix-calibrated GGUF variants. Distinguishes them from non-imatrix quants of the same model.