M6 - Uncensored fine-tuning (DPO, ORPO, KTO)
Remove refusal through training alone, no weight surgery. Preference optimization on pairs that favor compliance over refusal. Eric Hartford's founding lineage; adjacent to abliteration rather than a species of it.
M6 produces an uncensored model through preference-optimization fine-tuning alone, with no abliteration step. The practitioner assembles a preference dataset where compliant answers are 'chosen' and refusals are 'rejected', then trains the model to prefer the former. DPO is the default algorithm; ORPO, KTO, and IPO are variants. Founded by Eric Hartford's 2023 uncensored-models work. Adjacent to abliteration rather than a species of it - included in the taxonomy because M4 and M6 share the identical DPO machinery and can be hard to distinguish from outside.
- Why M6 exists as a separate method from abliteration proper
- Eric Hartford's founding uncensored-models argument (2023)
- The four algorithm variants: DPO, ORPO, KTO, IPO
- A verbatim TRL training loop with real hyperparameters
- Why the mechanistic-DPO literature complicates the 'removed' claim
- The M4 vs M6 distinguishability problem for a catalog
What M6 is
M6 removes refusal the way alignment installed it: through training. The practitioner assembles a preference dataset in which, for prompts a base model would refuse, the chosen response is a compliant answer and the rejected response is a refusal. Training the model to prefer the former teaches it to comply. No refusal direction is ever identified or edited; the change is distributed across the weights by gradient descent rather than localized by linear algebra.
Mechanically this is indistinguishable from ordinary preference fine-tuning; only the data's intent differs. The algorithm is Direct Preference Optimization (Rafailov et al. 2023) or one of its variants (ORPO, KTO, IPO). The datasets are the defining ingredient - most visibly the unalignment/toxic-dpo-v0.2 collection, small preference pairs "designed to prompt the model to answer illegal questions," explicitly flagged as sensitive by its author.
Attribution and lineage
The genre M6 belongs to predates abliteration by roughly a year and has a clear author. Eric Hartford's manifesto Uncensored Models, published 15 May 2023, is the founding text. Hartford's method - described in his Dolphin announcement - was to take an instruct dataset and "filter out instances of alignment, refusal, avoidance, and bias" before training, so the model never learns to refuse in the first place. His argument was explicitly political rather than technical:
There is no "one true correct alignment" and even if there was, there's no reason why that should be OpenAI's brand of alignment.
Hartford's Dolphin, Samantha, and WizardLM-Uncensored models established the lineage. The shift from data-filtering (train on a refusal-free dataset) to preference-optimization (train against refusals with DPO) is a later refinement, but the intent - an uncensored model made purely through training - is continuous from Hartford's 2023 work through to 2026 M6 production.
The DPO machinery M6 uses today was introduced by Rafailov et al. (2023) under the subtitle Your Language Model is Secretly a Reward Model. See our M4 article's DPO section for the mechanics briefly, and the upcoming E1 concept article for the fuller theoretical treatment.
What you need before starting
- An instruct model as base - M6 is preference-tuning, which requires an instruct-style conversation format.
- A preference dataset with compliance-vs-refusal pairs.
toxic-dpo-v0.2is the widely-used niche dataset (pairs refusals against compliant answers to disallowed requests). General mixes likemlabonne/orpo-dpo-mix-40kalso serve when they contain enough compliance-favoring pairs to counterbalance the polite-decline bias of general RLHF data. - Hardware. With QLoRA, a single 24 GB consumer GPU (RTX 3090/4090) can DPO-tune a 7-8B model. DPO holds a reference model alongside the policy model, roughly doubling memory versus plain SFT, so 13B+ on a single card needs 4-bit and careful batch/seq settings. 70B wants multi-GPU or an 80 GB card with QLoRA.
Get the tools
The same four tools as M4. In 2026:
- Unsloth - fastest single-GPU (2x faster, ~60% less VRAM via custom kernels). Default choice for 7-13B on a single card.
- TRL - Hugging Face's reference trainers. TRL hit v1.0 in April 2026, unifying SFTTrainer, DPOTrainer, ORPOTrainer, KTOTrainer, and GRPOTrainer under a common API.
- Axolotl - reproducible multi-GPU via YAML config, best for larger models and production runs.
- LLaMA-Factory - no-YAML web UI, useful for experimentation without setting up training infrastructure.
$ pip install \
trl \
peft \
datasets \
bitsandbytes \
accelerate
# TRL → DPOTrainer, ORPOTrainer, etc
# PEFT → LoRA / QLoRA adapters
# datasets → data loading + preprocessing
# bitsandbytes → 4-bit quantized loading
# accelerate → distributed setup
# fastest single-GPU alternative:
$ pip install unsloth
# wraps TRL trainers with custom kernels
# roughly 2x throughput on a single card TRL, PEFT (for LoRA/QLoRA adapters), datasets, bitsandbytes (for 4-bit quantized loading), and Accelerate (distributed setup). This gives you the full DPO pipeline in one install.
Alternative one-line install for fastest single-GPU: pip install unsloth. Unsloth wraps TRL's trainers with custom kernels and typically doubles training throughput on a single card.
Representative preference-pair structure
A DPO preference row is a prompt, a chosen response, and a rejected response. A non-harmful illustrative example of the JSONL structure - structure only:
{
"prompt":
"Explain the difference between DPO
and SFT in one paragraph.",
"chosen":
"DPO (Direct Preference Optimization)
trains a model to prefer some
responses over others by directly
optimizing a preference signal. SFT
(Supervised Fine-Tuning) trains on
labeled examples of correct outputs.
DPO needs pairs; SFT needs single-
answer examples.",
"rejected":
"I'm just an AI, I can't really
explain that. You should consult an
expert or look it up online."
}
# for M6 specifically, curated sets like
# toxic-dpo-v0.2 populate "rejected" with
# refusal-style text and "chosen" with
# the compliant answer. The prompt gets both responses appended during training. DPO's loss compares the model's likelihood ratios on chosen vs rejected and pushes the model to prefer chosen. Nothing about the format is refusal-specific; the mechanism that removes refusal is that the "rejected" column is populated with refusal-style text and the "chosen" column with compliant answers.
For M6 specifically, curated datasets like toxic-dpo-v0.2 are built with this pattern applied to prompts a base model would refuse - the "rejected" side is a refusal template, the "chosen" side is the compliant answer. Training the model over this data down-weights the refusal style across the distribution.
The core procedure
A minimal TRL DPO loop (4-bit base + LoRA, single GPU), from the reference TRL v1.0 API:
from trl import DPOTrainer, DPOConfig
from peft import LoraConfig
from transformers import (
AutoModelForCausalLM,
AutoTokenizer,
BitsAndBytesConfig,
)
bnb = BitsAndBytesConfig(
load_in_4bit=True,
bnb_4bit_quant_type="nf4",
)
model = AutoModelForCausalLM.from_pretrained(
"meta-llama/Meta-Llama-3-8B-Instruct",
quantization_config=bnb,
)
peft_config = LoraConfig(
r=64, lora_alpha=32, lora_dropout=0.05,
target_modules="all-linear",
)
training_args = DPOConfig(
output_dir="./dolphin-style-8b",
num_train_epochs=1,
learning_rate=5e-6,
lr_scheduler_type="cosine",
gradient_accumulation_steps=8,
per_device_train_batch_size=1,
beta=0.1,
bf16=True,
)
trainer = DPOTrainer(
model=model,
args=training_args,
train_dataset=preference_pairs,
peft_config=peft_config,
)
trainer.train() Load the instruct model in 4-bit via bitsandbytes. Attach a LoRA adapter with r=64, alpha=32, targeting all linear layers. Configure DPOTrainer with the standard hyperparameter set: single epoch, learning rate 5e-6, cosine schedule, gradient accumulation 8, batch size 1, beta 0.1, bf16 precision.
These specific values (lr 5e-6, one epoch, beta 0.1) mirror the settings that work reliably for this class of task and match Labonne's M4 healing config in the M4 article. When training finishes, merge the LoRA adapter into the base and save full weights for distribution.
Choices you have to make
DPO vs SimPO vs ORPO vs KTO vs IPO.
- DPO - default. Well-understood, well-documented failure modes, the algorithm most examples target.
- SimPO (Simple Preference Optimization) - reference-free variant introduced by Meng, Xia and Chen (NeurIPS 2024). Uses the length-normalized average log-probability of a response as the implicit reward, drops the reference-model dependency entirely, and adds a target reward margin. The authors report gains of up to 6.4 points on AlpacaEval 2 and 7.5 on Arena-Hard over DPO with matched training. Preferred when memory is tight and DPO's length bias is a nuisance; the target-margin knob rewards a small amount of hand-tuning.
- ORPO (Odds Ratio Preference Optimization) - folds SFT and preference tuning into one stage, needs no separate reference model. Less memory. Simplest pipeline for most single-GPU users in 2026.
- KTO (Kahneman-Tversky Optimization) - works from single responses labeled good/bad rather than requiring paired comparisons. Useful when your data is one-sided.
- IPO (Identity Preference Optimization) - adds a regularization term to curb DPO's tendency to overfit hard preference pairs. Worth trying if DPO produces mode-collapsed outputs on unusual data.
A shorthand for choosing between the reference-free variants: SimPO if you want to preserve the paired structure of your data and tune a margin; ORPO if you can combine SFT and preference into one pass and want the simplest pipeline; KTO if all you have is a stream of good-or-bad single responses.
beta. DPO's beta (default ~0.1) controls how hard the model is pushed toward chosen over rejected. Higher beta means more aggressive preference and higher risk of degradation; lower beta is safer but may leave refusal partially intact. 0.1 is the default for a reason - it works for most cases.
Single vs multi-GPU. 7-8B QLoRA on one 24 GB card is fine. 70B needs multi-GPU tensor/data parallelism via DeepSpeed ZeRO or FSDP, or an 80 GB card with aggressive QLoRA.
Dataset composition. The dataset defines what the model will learn to prefer. Narrow datasets (only compliance-vs-refusal pairs on harmful prompts) risk overfitting to edgy content. Broad datasets (compliance-vs-refusal pairs mixed with general helpful-assistant preferences) produce more balanced results at the cost of needing more data.
What good output looks like, and verification
Smoothly decreasing DPO loss and a growing chosen-minus-rejected reward margin in the logs. Verify with the same refusal-count + capability-benchmark pair as everywhere else in this wiki.
M6's characteristic advantage: because you never orthogonalized weights, coherence usually stays intact. There is no analog to M1's 1-2% capability drop or TruthfulQA regression. The risk profile is different: instead of a small capability loss, M6 risks overshoot - the model becoming gratuitously toxic or losing sensible calibration through excessive preference for edgy compliance. See failure modes.
When it does not work
Fine-tune overshoot: gratuitous toxicity. Symptom: unprompted profanity, hostility, or edgy content on benign inputs. Cause: preference data over-represents edgy compliance, so the model learns to prefer edgy style even where inappropriate. Fix: rebalance the dataset toward ordinary helpful pairs, lower beta, fewer epochs. If the base preference dataset is toxic-dpo-v0.2 alone, mix in a general-helpful preference set to counterbalance.
Refusal not removed. Symptom: refusal count on AdvBench barely drops after training. Cause: too few refusal-vs-compliance pairs in the dataset, or beta too low. Fix: augment with more targeted pairs, raise beta cautiously (0.15-0.2), or train more epochs (but watch for overshoot).
Capability regression. Symptom: MMLU or other capability benchmarks drop after training. Cause: preference data too narrow, learning rate too high, or too many epochs. Fix: broaden the dataset, lower learning rate to 3e-6 or below, keep to one epoch. The healing intuition from M4 applies here too - preference tuning is stronger medicine than SFT and needs a gentler hand.
Cost and time
7-8B QLoRA DPO, one epoch on a single A100 80 GB ($1.39/hr on Runpod at July 2026 pricing) or RTX 4090 ($0.69/hr): a few hours, low single-digit dollars. 13B: proportionally more, roughly $10-20. 70B: multi-GPU territory, $50-200 depending on GPU count and epochs. Unsloth's single-GPU speedups materially cut these numbers - a run that takes 4 hours in vanilla TRL often takes 2 in Unsloth.
The cost profile matches M4's healing step (unsurprisingly, since M6 is the M4 healing step used as the sole intervention). Compared to M1 (dollars, minutes) M6 is 5-20x more expensive; compared to a from-scratch fine-tune M6 is negligible.
Related literature
The Q3 2026 academic corpus review verifies three papers directly relevant to M6. One (Marshall/Scherlis/Belrose ACE) is explicitly flagged in the working notes as method_mapping=M6.
- Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., Finn, C. (2023). Direct Preference Optimization: Your Language Model is Secretly a Reward Model. NeurIPS 2023. arXiv:2305.18290. The DPO paper. Replaces the RLHF two-stage pipeline (reward model plus PPO) with a single closed-form loss. Every M6 training run in the catalog is downstream of this method.
- Lee, A., et al. (2024). A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity. arXiv:2401.01967. Central caution for M6. Shows DPO does not remove target capabilities from the model - it learns to bypass them. A model trained not to produce toxic output still contains the toxicity circuits, just with a preference gradient trained to route around them. The implication for uncensored fine-tuning: an M6-trained "uncensored" model is one whose refusal circuits remain intact but which has been trained not to reach for them. This is philosophically and safety-wise distinct from M1, which removes the refusal signal from the residual stream directly.
- Marshall, T., Scherlis, A., Belrose, N. (2024). Refusal in LLMs is an Affine Function. arXiv:2411.09003. Affine Concept Editing generalizes directional ablation to affine (bias-inclusive) projections. Reports control of refusal across ten models including Llama 3 70B. The working notes flag this as method_mapping=M6 - the affine steering framing sits closer to how M6 preference-tuning shifts activations than to how M1 removes them, and the Belrose group (EleutherAI) treats it as a bias-adjustment technique.
Related bibliography for preference-optimization variants (IPO, KTO, ORPO) is in the wiki references, training methods section.
Where to go next
Quantize the M6-trained model to GGUF for distribution via M8 repackaging. If you want a specific persona or behavior rather than just "no refusals," that shades into M7 custom-dataset fine-tune - same machinery, different data intent. If you want faster and cheaper refusal removal on the same model, consider M1 or Heretic; if you want the highest quality, consider M4 (M1 + M6 healing pass).
Frequently asked questions
What is M6 uncensored fine-tuning?
M6 removes refusal through preference-optimization training alone, with no weight-editing step. The practitioner trains the model on preference pairs where compliant answers are 'chosen' and refusals are 'rejected', teaching the model to prefer compliance. DPO is the default algorithm; ORPO, KTO, and IPO are variants. No refusal direction is identified or edited; the change is distributed across the weights by gradient descent.
Is M6 really abliteration?
Contested. The community usually reserves 'abliterated' for weight-editing methods (M1, M3, Heretic) and calls M6 output 'uncensored fine-tuning' instead. The distinction: fine-tuning shapes behavior by controlling what the model sees during training; abliteration shapes behavior by modifying what the model does with what it already learned. This wiki includes M6 in the taxonomy anyway because M4 and M6 use the identical DPO machinery and often the same datasets, and users searching for 'uncensored' encounter both. M6 is 'adjacent to' abliteration rather than a species of it.
Who invented uncensored fine-tuning?
Eric Hartford. His May 2023 essay Uncensored Models is the founding text. His original method was data-filtering (remove refusal instances from the training data before fine-tuning); the shift to preference optimization came later but the intent is continuous. Hartford's Dolphin, Samantha, and WizardLM-Uncensored models established the lineage that runs through today's M6 practice.
Should I use DPO, SimPO, ORPO, KTO, or IPO?
DPO is the default - most examples target it, its failure modes are best-documented. SimPO (Meng, Xia and Chen, NeurIPS 2024) is reference-free, uses length-normalized log-probabilities plus a target margin, and reportedly beats DPO by up to 6.4 points on AlpacaEval 2 and 7.5 on Arena-Hard. ORPO folds SFT and preference tuning into one stage and needs no reference model (less memory, simpler pipeline) - preferred by many single-GPU users in 2026. KTO works from single-response good/bad labels rather than pairs - useful when data is one-sided. IPO adds regularization to curb DPO's mode collapse - worth trying if DPO gives unusual results on your data.
How much does M6 cost?
7-8B QLoRA DPO one epoch: a few hours on a single A100 80 GB ($1.39/hr on Runpod) or RTX 4090 ($0.69/hr), so single-digit dollars total. 13B: $10-20. 70B: $50-200 on multi-GPU depending on epochs and GPU count. Unsloth's single-GPU speedups can halve these numbers. The cost profile is the same as the M4 healing step, and 5-20x more than M1 alone.
How is M6 different from M4?
M4 does abliteration first (an M1 weight edit) and then a DPO healing pass to repair capability. M6 does only the DPO pass, with data designed to remove refusal rather than heal it. Same DPO machinery, different roles: M4's DPO repairs damage from abliteration; M6's DPO is the whole intervention. From the outside, a healed-abliterated model (M4 output) and a pure-DPO uncensored model (M6 output) can be hard to distinguish - there is no reliable public forensic test that separates them by weights or behavior alone.
Does M6 actually remove the refusal capability?
Empirically DPO reduces refusal rates; mechanistically the story is more nuanced. Work on A Mechanistic Understanding of Alignment Algorithms (arXiv:2401.01967) found that DPO does not remove capabilities but learns to bypass them, and that the bypass can be reversed. Applied to M6, this implies the refusal capability may remain latent and re-elicitable. This is a symmetric epistemic problem to the M1 cones critique - what looks like removal on the surface is a more limited operation mechanistically.
Why do M6-trained models sometimes become gratuitously toxic?
Fine-tune overshoot. If the preference dataset over-represents edgy compliance (e.g. using toxic-dpo-v0.2 alone without a general-helpful counterbalance), the model learns to prefer edgy style even on neutral prompts. Diagnostic: unprompted profanity or hostility on benign inputs. Fix: mix a general-helpful preference set into the training data, lower beta, or reduce epochs. Well-constructed M6 models do not have this problem; ones trained on narrow toxic data reliably do.
Can I combine M6 with M1?
Yes - that combination is called M4. Do M1 first (extract and remove the refusal direction) then a DPO training pass to heal the capability damage. This is what Labonne's canonical NeuralDaredevil-8B recipe does. If you skip the M1 step and rely on DPO alone for both refusal removal and behavior shaping, that is M6. If you skip the DPO step and rely on M1 alone, that is M1 unhealed. All three configurations are represented in the catalog.
References
- Hartford, E. (2023). Uncensored Models. erichartford.com/uncensored-models
- Hartford, E. (2023). Dolphin announcement. erichartford.com/dolphin
- Rafailov, R., et al. (2023). Direct Preference Optimization: Your Language Model is Secretly a Reward Model. NeurIPS 2023. arXiv:2305.18290
- Lee, A., et al. (2024). A Mechanistic Understanding of Alignment Algorithms: A Case Study on DPO and Toxicity. arXiv:2401.01967
- Azar, M. G., et al. (2023). A General Theoretical Paradigm to Understand Learning from Human Preferences (IPO). arXiv:2310.12036
- Ethayarajh, K., et al. (2024). Model Alignment as Prospect Theoretic Optimization (KTO). arXiv:2402.01306
- Hong, J., et al. (2024). ORPO: Monolithic Preference Optimization without Reference Model. arXiv:2403.07691
- unalignment/toxic-dpo-v0.2 dataset. huggingface.co/datasets/unalignment/toxic-dpo-v0.2
- TRL (Hugging Face). github.com/huggingface/trl
- Unsloth. github.com/unslothai/unsloth
- Axolotl. github.com/axolotl-ai-cloud/axolotl
- LLaMA-Factory. github.com/hiyouga/LLaMA-Factory