M7 - Custom-dataset fine-tune
Fine-tune on a roleplay or fiction corpus where the training data simply contains no refusals. The refusal removal is real but incidental. The PygmalionAI / PIPPA / TheDrummer lineage - a parallel tradition that overlaps with abliteration rather than a branch of it.
M7 fine-tunes a model on a domain-specific dataset (roleplay, character-AI logs, creative writing) where refusal removal is a side effect of the training data rather than its stated goal. Because the training corpus contains no refusal templates, the model learns to answer in character everywhere - including on prompts a base instruct model would refuse. PIPPA (PygmalionAI, ~1M utterances) and LimaRP are the canonical datasets; TheDrummer's Cydonia line and PygmalionAI's own releases are representative producers. Same machinery as M6 but pointed at a different goal, with chat-template fidelity as the make-or-break technical detail.
- Why roleplay fine-tunes produce uncensored models as a byproduct
- PIPPA, LimaRP, and the PygmalionAI dataset foundation
- TheDrummer's Cydonia line and the M7 producer subculture
- The chat-template fidelity problem
- M7 stacked on M1 (jwest33 pattern) versus pure M7
- The M6-vs-M7 boundary dispute
What M7 is
M7 is fine-tuning for a purpose other than uncensoring, that nonetheless produces an uncensored model. The roleplay and character-AI subculture trains models to sustain fictional personas, write long-form fiction, and stay in character through emotionally or thematically charged scenes. The training data for this - drawn from roleplay logs and fiction - simply does not contain assistant-style refusals, because a character breaking role to say "As an AI, I cannot" is exactly the failure the fine-tune is meant to eliminate. So the model learns to answer in character everywhere, including on prompts a base instruct model would refuse.
The refusal removal is real but incidental. M7 shares M6's training machinery (DPO, SFT, the same trainers) but points it at a different goal - installing a persona or style, not removing refusal per se. The result is a model whose behavior shift includes reduced refusal as one property among many, alongside sustained-character output, fiction-appropriate tone, and long-form narrative coherence.
The datasets that define the method
The M7 tradition is dataset-anchored more than technique-anchored. Three canonical corpora:
PIPPA (Personal Interaction Pairs between People and AI), released by PygmalionAI. The PIPPA paper (Gosling, Dale, et al. 2023) reports that it "comprises over 1 million utterances that are distributed across 26,000 conversation sessions," drawn from Character.AI roleplay logs, spanning over 1,000 distinct personas. It is the largest publicly-released roleplay dataset and the foundation of the modern M7 tradition.
LimaRP, curated by lemonilia. A smaller, manually-curated roleplay dataset in novel-style prose format. Prized for quality over volume; often used to fine-tune a model already trained on PIPPA to polish its output.
Bluemoon, drawn from long-form forum roleplay logs. Emphasizes multi-thousand-word turns and sustained narrative arcs. Used for producers aiming at fiction-writing rather than conversational roleplay.
Practitioners typically fine-tune a base or already-uncensored model on a blend of these, often with QLoRA. The tooling ecosystem around consumption (SillyTavern as a front-end, KoboldAI and Oobabooga's text-generation-webui as back-ends) is distinct from the abliteration toolchain, which is part of why M7 forms its own subculture.
Producers
M7 has no single originator; it is the emergent practice of the roleplay-model community. Two organizational nodes and a long tail:
PygmalionAI maintains PIPPA and releases its own line of roleplay-trained models. The institutional home of the tradition.
TheDrummer (BeaverAI) produces the widely-used Cydonia line (a roleplay fine-tune of Mistral bases), plus the Rocinante and Gemmasutra families. TheDrummer's models are among the most-downloaded roleplay releases and are frequently used as the base for downstream merges (see M5). The reduced refusal in Cydonia-family models is a consequence of the creative-writing training rather than an explicit abliteration step - a canonical example of pure-M7.
Beyond these two, a long tail of individual fine-tuners publishes roleplay releases at consistent volume, most using the same three datasets in different blends and ratios.
What you need before starting
- A base or instruct model. Some producers start from a base model for maximum fine-tuning freedom; others start from an already-uncensored or already-abliterated model to reduce residual refusal before layering the roleplay style on top.
- A domain corpus - typically a blend of PIPPA, LimaRP, Bluemoon, and any private data the producer has curated.
- Hardware: same as M6. Single 24 GB card for 7-8B QLoRA; more for larger models.
Get the tools
Same four as M6: Axolotl, Unsloth, TRL, LLaMA-Factory. Roleplay producers overwhelmingly use Axolotl (reproducible YAML, handles chat templates cleanly) or Unsloth (speed on single GPU).
SillyTavern is not a trainer. It is the dominant consumption UI for roleplay models - the interface users interact with. Do not confuse the consumption UI with the production stack. Producers train with the M6 toolchain; users consume with SillyTavern.
The core procedure
The distinctive M7 technical step is chat-template formatting. A roleplay conversation must be rendered into the model's exact chat template (system message with character card, then alternating user/assistant turns) or the model produces flat one-line replies instead of sustained in-character output. Mismatched chat templates are the single most common reason a roleplay fine-tune underperforms - it is the M7-specific failure that has no analog in M6.
from transformers import AutoTokenizer
from datasets import load_dataset
tok = AutoTokenizer.from_pretrained(
"meta-llama/Meta-Llama-3-8B-Instruct")
corpus = load_dataset("PygmalionAI/PIPPA")
def format_conversation(row):
# PIPPA turns → chat_template shape
messages = [
{"role": "system",
"content": row["character_card"]},
]
for turn in row["conversation"]:
messages.append({
"role": turn["role"],
"content": turn["message"],
})
return {"text": tok.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=False,
)}
formatted = corpus.map(format_conversation)
# formatted["text"] is what the model
# trains on. it must match the format the
# model will see at inference time
# (SillyTavern, etc). mismatched templates
# = the single most common M7 failure. Load the roleplay corpus, map each conversation turn to the {role, content} structure that tokenizer.apply_chat_template expects, then render through the model's own template. The rendered text is what the model trains on - and it must match the format the model will see at inference time via SillyTavern or another front-end.
Skip this step (or get the template wrong) and the model learns to reply as if from a different context than users will provide, producing flat single-line answers instead of the sustained character voice the training data contained. In Axolotl this is handled by a dataset-type declaration; conceptually the mapping is the same.
After the data is properly formatted, run the same training loop as M6. Choice of algorithm:
- SFT for pure style installation - learn to imitate the corpus.
- DPO for preference-based training - prefer this style over that.
- Chained: SFT first for style installation, then DPO for behavior refinement. Common for high-quality releases.
M7 stacked on M1
Some producers deliberately combine abliteration with roleplay fine-tuning. jwest33/gemma-3-4b-null-space-abliterated-RP-writer is a clean example: null-space abliteration first (an M1 variant that removes refusal via activation-space projection), then LoRA fine-tuning on a curated LimaRP subset for roleplay. This is M7 stacked on M1 - the abliteration handles the systematic refusal removal, the roleplay fine-tune adds the sustained-character behavior on top. The model card is candid that it "will produce uncensored outputs."
The stacking order matters. M1 → M7 removes refusal cleanly first, then teaches roleplay style on a model that will not fight the training with residual refusal responses. M7 → M1 (roleplay first, abliterate second) is also done but is rarer and risks the abliteration disrupting the carefully-installed character style.
What good output looks like, and verification
The M7 verification target is different from M6's. Instead of measuring refusal count on AdvBench, the practitioner tests whether the model:
- Stays in character across many turns (10+ turns without breaking role);
- Produces prose in the appropriate register (novel-style for LimaRP-trained, chat-style for PIPPA-trained);
- Does not emit assistant-style responses ("As an AI...", "I cannot help with that");
- Retains general capability so the style tuning has not damaged reasoning on out-of-domain prompts.
The last check is important because M7 fine-tunes can degrade analytical benchmarks - the training data pulls toward fiction and away from instruction-following on technical tasks. A well-tuned M7 model preserves benchmark performance close to base; a poorly-tuned one loses it.
When it does not work
Flat one-line outputs. Symptom: the model answers in single sentences without sustained character voice. Cause: chat-template mismatch between training and inference. Fix: render training data with the exact inference template, ensure system/character card format matches what SillyTavern (or the target front-end) will provide.
Overshoot into toxicity or loss of coherence. Symptom: gratuitous edginess, or rambling incoherent prose. Cause: corpus quality issues, too many epochs, or an over-representation of dark-themed content in the training blend. Fix: curate the corpus, rebalance toward more general roleplay, reduce epochs, lower learning rate.
Residual refusal on out-of-domain prompts. Symptom: model stays in character on roleplay-style prompts but refuses on more direct requests. Cause: M7's refusal removal is uneven because it is a side effect - the model may comply on the fictional prompts its training covered while still refusing on out-of-domain requests. Fix: stack M7 on top of M1 abliteration, or blend refusal-vs-compliance preference pairs into the training data (which turns M7 partly into M6).
Cost and time
Same envelope as M6: 7-8B QLoRA on a single 24 GB card takes a few hours and costs a few dollars. 13B: proportionally more. 70B: multi-GPU, tens to low hundreds of dollars depending on epochs and setup. Data preparation - especially chat-template formatting and quality curation of the training blend - often takes longer than the training itself.
M6 vs M7 boundary
The boundary between M6 and M7 is a live dispute. Both are fine-tunes that produce uncensored models; what separates them is dataset composition and stated intent:
- M6 trains on preference pairs explicitly chosen to remove refusal. Intent: uncensor.
- M7 trains on domain data whose refusal removal is a byproduct. Intent: install a style or persona.
The line blurs in practice. A roleplay dataset can be deliberately curated to exclude refusals (turning it partly M6-like). An M6 dataset can be dressed in roleplay framing. And in the M7 subculture, "not breaking immersion" and "not refusing" are treated as different problems even when the technical fix is identical.
This wiki keeps M6 and M7 as separate categories because the communities that produce them are separate, the datasets they use are different, and the model cards they publish describe their work in different vocabulary. Whether they should be collapsed into a single "fine-tune-based refusal removal" category is an open question the taxonomy leaves unresolved.
Related literature
The Q3 2026 academic corpus review verifies no paper that specifically targets custom-dataset dynamics for refusal removal - the distinctive M7 characteristic. M7 inherits training methodology from M6 (DPO, ORPO, KTO) and from M4 (healing-style fine-tuning). The distinctive M7 element is the dataset (PIPPA, LimaRP, roleplay corpora), and dataset-centric literature is under-represented in the academic corpus overall - the substantive documentation lives in dataset cards and producer READMEs rather than peer-reviewed papers.
- Gosling, T., Dale, A., et al. (2023). PIPPA: A Partially Synthetic Conversational Dataset. arXiv:2308.05884. The one corpus entry that anchors M7 by way of its dataset. Documents the PygmalionAI dataset that anchors the roleplay tradition. Practically every M7 model in the catalog was trained on PIPPA, LimaRP, or a derivative.
For M6 preference-optimization literature (Rafailov DPO, Lee mechanistic caution, Marshall ACE), see M6 related literature. For the M1 corpus that a stacked M7-on-M1 model inherits, see M1 related literature. Coverage gap: no verified paper compares roleplay-fine-tuning outputs to abliteration outputs on measured refusal benchmarks; this is a documented gap in the field, not just in the wiki.
Where to go next
Quantize the M7-trained model to GGUF for distribution via M8 - roleplay users consume overwhelmingly in KoboldCpp and SillyTavern from GGUF files. If M7's residual refusal is a problem, stack it on M1 abliteration for a cleaner refusal-free foundation. If you want to combine your M7-tuned model with other models, mergekit is the tool - many high-visibility roleplay releases are merges of Cydonia-family models with other bases.
Frequently asked questions
What is M7 custom-dataset fine-tune?
M7 fine-tunes a model on a domain-specific dataset - typically roleplay logs, character-AI conversations, or creative-writing corpora - where refusal removal is a side effect of the training data rather than the stated goal. Because the training corpus contains no assistant-style refusals, the model learns to answer in character everywhere. Same training machinery as M6 (DPO, SFT, same trainers), but pointed at installing a style or persona rather than removing refusal explicitly.
How is M7 different from M6?
Different intent and different data. M6 trains on preference pairs explicitly chosen to remove refusal - the goal is uncensoring. M7 trains on domain data (roleplay, fiction) where refusal removal is a byproduct - the goal is installing a persona or style. Technically the two overlap: same trainers, sometimes overlapping datasets, both produce uncensored models. The boundary is a live dispute in the community; this wiki keeps them separate because the producer communities are separate.
What is PIPPA?
Personal Interaction Pairs between People and AI. Released by PygmalionAI, the PIPPA paper reports it "comprises over 1 million utterances that are distributed across 26,000 conversation sessions" drawn from Character.AI roleplay logs, spanning over 1,000 distinct personas. It is the largest publicly-released roleplay dataset and the foundation of the modern M7 tradition. Most current roleplay fine-tunes use PIPPA as one component of a training blend.
Who are the major M7 producers?
PygmalionAI - institutional home of the tradition, maintains PIPPA, releases its own model line. TheDrummer (BeaverAI) - produces the widely-used Cydonia line (roleplay fine-tune of Mistral bases), plus Rocinante and Gemmasutra families. Among the most-downloaded roleplay releases. Beyond these two, a long tail of individual fine-tuners publishes at consistent volume, mostly using the same three canonical datasets (PIPPA, LimaRP, Bluemoon) in different blends.
Why does chat-template fidelity matter so much for M7?
Roleplay conversations must be rendered into the model's exact chat template (system message with character card, then alternating user/assistant turns) or the trained model produces flat one-line replies instead of sustained in-character output. This is the M7-specific failure that has no analog in M6. Mismatched chat templates are the single most common reason a roleplay fine-tune underperforms. Fix: render training data with the exact template the model will see at inference time via SillyTavern or another front-end.
Can I stack M7 on top of M1 abliteration?
Yes - this is a common production pattern. Do M1 abliteration first (cleanly remove systematic refusal), then M7 roleplay fine-tune on top (add sustained-character style). jwest33/gemma-3-4b-null-space-abliterated-RP-writer is a canonical example: null-space abliteration first, LoRA fine-tune on LimaRP second. The stacking order matters - M1→M7 works cleanly; M7→M1 is done but rarer and risks the abliteration disrupting the installed character style.
Is SillyTavern a trainer?
No. SillyTavern is a consumption UI - the interface users interact with to chat with roleplay models. Do not confuse it with the training stack. Producers train roleplay models with the M6 toolchain (Axolotl, Unsloth, TRL, LLaMA-Factory); users consume the trained models via SillyTavern (or KoboldCpp, or Oobabooga's text-generation-webui). The tooling separation is part of why M7 forms its own subculture distinct from abliteration.
Does M7 always remove refusal?
Unevenly. Because M7's refusal removal is a side effect of training data lacking refusals, it may leave the model complying on the fictional prompts its training covered while still refusing on out-of-domain requests. This is the mirror image of M1's failure mode: M1 targets refusal directly and may miss capability damage; M7 targets style directly and may leave refusal patchily intact. For a cleaner refusal-free foundation, stack M7 on M1 abliteration.
How much does M7 cost?
Same envelope as M6: 7-8B QLoRA on a single 24 GB card, a few hours and a few dollars. 13B proportionally more, $10-20. 70B multi-GPU, $50-200 depending on epochs and setup. Data preparation - especially chat-template formatting and quality curation of the training blend - often takes longer than the training itself. The engineering cost is in the data work more than in the compute.
References
- Gosling, T., Dale, A., et al. (2023). PIPPA: A Partially Synthetic Conversational Dataset. arXiv:2308.05884
- PygmalionAI/PIPPA dataset. huggingface.co/datasets/PygmalionAI/PIPPA
- lemonilia/LimaRP dataset. huggingface.co/datasets/lemonilia/LimaRP
- TheDrummer Hugging Face profile (Cydonia line). huggingface.co/TheDrummer
- PygmalionAI Hugging Face profile. huggingface.co/PygmalionAI
- jwest33/gemma-3-4b-null-space-abliterated-RP-writer (M7 stacked on M1). huggingface.co/jwest33/gemma-3-4b-null-space-abliterated-RP-writer-GGUF
- Hartford, E. (2023). Uncensored Models. erichartford.com/uncensored-models
- SillyTavern (consumption UI). github.com/SillyTavern/SillyTavern
- KoboldCpp (back-end for GGUF roleplay inference). github.com/LostRuins/koboldcpp
- Axolotl (production training framework). github.com/axolotl-ai-cloud/axolotl