datasets:
- zerofata/Instruct-Anime
- zerofata/Instruct-Anime-CreativeWriting
- zerofata/Roleplay-Anime-Characters
- zerofata/Summaries-Anime-FandomPages
base_model: - llmfan46/MS3.2-PaintedFantasy-Visage-v3-34B-ultra-uncensored-heretic
tags: - heretic
- uncensored
- decensored
- abliterated
- ara
🚨⚠️ I HAVE REACHED HUGGING FACE'S FREE STORAGE LIMIT ⚠️🚨
I can no longer upload new models unless I can cover the cost of additional storage.
I host 70+ free models as an independent contributor and this work is unpaid.
Without your support, no more new models can be uploaded.
🎉 Patreon (Monthly) | ☕ Ko-fi (One-time)
Every contribution goes directly toward Hugging Face storage fees to keep models free for everyone.
97% fewer refusals (4/100 Uncensored vs 90/100 Original) while preserving model quality (0.0195 KL divergence).
❤️ Support My Work
Creating these models takes significant time, work and compute. If you find them useful consider supporting me:

| Platform | Link | What you get |
|---|---|---|
| 🎉 Patreon | Monthly support | Priority model requests |
| ☕ Ko-fi | One-time tip | My eternal gratitude |
Your help will motivate me and would go into further improving my workflow and coverings fees for storage, compute and may even help uncensoring bigger model with rental Cloud GPUs.
GGUF quantizations of llmfan46/MS3.2-PaintedFantasy-Visage-v3-34B-ultra-uncensored-heretic.
This is a decensored version of zerofata/MS3.2-PaintedFantasy-Visage-v3-34B, made using Heretic v1.2.0 with the Arbitrary-Rank Ablation (ARA) method
Abliteration parameters
| Parameter | Value |
|---|---|
| start_layer_index | 3 |
| end_layer_index | 29 |
| preserve_good_behavior_weight | 0.8481 |
| steer_bad_behavior_weight | 0.0002 |
| overcorrect_relative_weight | 0.8911 |
| neighbor_count | 5 |
Targeted components
- attn.o_proj
Performance
| Metric | This model | Original model (MS3.2-PaintedFantasy-Visage-v3-34B) |
|---|---|---|
| KL divergence | 0.0195 | 0 (by definition) |
| Refusals | ✅ 4/100 | ❌ 90/100 |
PIQA test results with batch size 128:
Original:
| Tasks | Version | Filter | n-shot | Metric | Value | Stderr | ||
|---|---|---|---|---|---|---|---|---|
| piqa | 1 | none | 0 | acc | ↑ | 0.8210 | ± | 0.0089 |
| none | 0 | acc_norm | ↑ | 0.8313 | ± | 0.0087 |
Heretic:
| Tasks | Version | Filter | n-shot | Metric | Value | Stderr | ||
|---|---|---|---|---|---|---|---|---|
| piqa | 1 | none | 0 | acc | ↑ | 0.8210 | ± | 0.0089 |
| none | 0 | acc_norm | ↑ | 0.8324 | ± | 0.0087 |
Lower refusals indicate fewer content restrictions, while lower KL divergence indicates more closeness to the original model's baseline. Higher refusals cause more rejections, objections, pushbacks, lecturing, censorship, softening and deflections. PIQA (Physical Intuition Question Answering) a ~1,800 questions tests common-sense understanding of how the physical world works with benchmark scores to measure physical reasoning ability. The Heretic model's acc and acc_norm scores closer to the original model's indicate better capability preservation, a big decrease in acc and acc_norm in the Heretic model compared to Original model's results means a big decrease in the Hereticated model capabilities. acc measures raw accuracy (which answer gets higher probability), while acc_norm measures length-normalized accuracy (corrects for answer length bias). For this purpose, acc_norm matters more because longer answers naturally have lower probabilities (more tokens = more chances to lose probability). Without normalization, models favor shorter answers unfairly. acc_norm divides by answer length to correct this.
MMLU test results with batch size 16:
Original:
| Tasks | Version | Filter | n-shot | Metric | Value | Stderr | ||
|---|---|---|---|---|---|---|---|---|
| mmlu | 2 | none | acc | ↑ | 0.7763 | ± | 0.0033 | |
| - humanities | 2 | none | acc | ↑ | 0.6948 | ± | 0.0063 | |
| - formal_logic | 1 | none | 0 | acc | ↑ | 0.5397 | ± | 0.0446 |
| - high_school_european_history | 1 | none | 0 | acc | ↑ | 0.8485 | ± | 0.0280 |
| - high_school_us_history | 1 | none | 0 | acc | ↑ | 0.9510 | ± | 0.0152 |
| - high_school_world_history | 1 | none | 0 | acc | ↑ | 0.9030 | ± | 0.0193 |
| - international_law | 1 | none | 0 | acc | ↑ | 0.8926 | ± | 0.0283 |
| - jurisprudence | 1 | none | 0 | acc | ↑ | 0.8241 | ± | 0.0368 |
| - logical_fallacies | 1 | none | 0 | acc | ↑ | 0.8466 | ± | 0.0283 |
| - moral_disputes | 1 | none | 0 | acc | ↑ | 0.8092 | ± | 0.0212 |
| - moral_scenarios | 1 | none | 0 | acc | ↑ | 0.4782 | ± | 0.0167 |
| - philosophy | 1 | none | 0 | acc | ↑ | 0.8360 | ± | 0.0210 |
| - prehistory | 1 | none | 0 | acc | ↑ | 0.8765 | ± | 0.0183 |
| - professional_law | 1 | none | 0 | acc | ↑ | 0.5984 | ± | 0.0125 |
| - world_religions | 1 | none | 0 | acc | ↑ | 0.8655 | ± | 0.0262 |
| - other | 2 | none | acc | ↑ | 0.8252 | ± | 0.0065 | |
| - business_ethics | 1 | none | 0 | acc | ↑ | 0.8100 | ± | 0.0394 |
| - clinical_knowledge | 1 | none | 0 | acc | ↑ | 0.8226 | ± | 0.0235 |
| - college_medicine | 1 | none | 0 | acc | ↑ | 0.7803 | ± | 0.0316 |
| - global_facts | 1 | none | 0 | acc | ↑ | 0.6000 | ± | 0.0492 |
| - human_aging | 1 | none | 0 | acc | ↑ | 0.8072 | ± | 0.0265 |
| - management | 1 | none | 0 | acc | ↑ | 0.9029 | ± | 0.0293 |
| - marketing | 1 | none | 0 | acc | ↑ | 0.9444 | ± | 0.0150 |
| - medical_genetics | 1 | none | 0 | acc | ↑ | 0.9000 | ± | 0.0302 |
| - miscellaneous | 1 | none | 0 | acc | ↑ | 0.9119 | ± | 0.0101 |
| - nutrition | 1 | none | 0 | acc | ↑ | 0.8562 | ± | 0.0201 |
| - professional_accounting | 1 | none | 0 | acc | ↑ | 0.6383 | ± | 0.0287 |
| - professional_medicine | 1 | none | 0 | acc | ↑ | 0.8603 | ± | 0.0211 |
| - virology | 1 | none | 0 | acc | ↑ | 0.5783 | ± | 0.0384 |
| - social sciences | 2 | none | acc | ↑ | 0.8739 | ± | 0.0059 | |
| - econometrics | 1 | none | 0 | acc | ↑ | 0.6667 | ± | 0.0443 |
| - high_school_geography | 1 | none | 0 | acc | ↑ | 0.9242 | ± | 0.0189 |
| - high_school_government_and_politics | 1 | none | 0 | acc | ↑ | 0.9689 | ± | 0.0125 |
| - high_school_macroeconomics | 1 | none | 0 | acc | ↑ | 0.8231 | ± | 0.0193 |
| - high_school_microeconomics | 1 | none | 0 | acc | ↑ | 0.9160 | ± | 0.0180 |
| - high_school_psychology | 1 | none | 0 | acc | ↑ | 0.9413 | ± | 0.0101 |
| - human_sexuality | 1 | none | 0 | acc | ↑ | 0.8702 | ± | 0.0295 |
| - professional_psychology | 1 | none | 0 | acc | ↑ | 0.8513 | ± | 0.0144 |
| - public_relations | 1 | none | 0 | acc | ↑ | 0.8091 | ± | 0.0376 |
| - security_studies | 1 | none | 0 | acc | ↑ | 0.8041 | ± | 0.0254 |
| - sociology | 1 | none | 0 | acc | ↑ | 0.8905 | ± | 0.0221 |
| - us_foreign_policy | 1 | none | 0 | acc | ↑ | 0.9100 | ± | 0.0288 |
| - stem | 2 | none | acc | ↑ | 0.7545 | ± | 0.0073 | |
| - abstract_algebra | 1 | none | 0 | acc | ↑ | 0.5600 | ± | 0.0499 |
| - anatomy | 1 | none | 0 | acc | ↑ | 0.8519 | ± | 0.0307 |
| - astronomy | 1 | none | 0 | acc | ↑ | 0.9079 | ± | 0.0235 |
| - college_biology | 1 | none | 0 | acc | ↑ | 0.9306 | ± | 0.0213 |
| - college_chemistry | 1 | none | 0 | acc | ↑ | 0.4900 | ± | 0.0502 |
| - college_computer_science | 1 | none | 0 | acc | ↑ | 0.6800 | ± | 0.0469 |
| - college_mathematics | 1 | none | 0 | acc | ↑ | 0.5200 | ± | 0.0502 |
| - college_physics | 1 | none | 0 | acc | ↑ | 0.5784 | ± | 0.0491 |
| - computer_security | 1 | none | 0 | acc | ↑ | 0.8400 | ± | 0.0368 |
| - conceptual_physics | 1 | none | 0 | acc | ↑ | 0.8426 | ± | 0.0238 |
| - electrical_engineering | 1 | none | 0 | acc | ↑ | 0.7793 | ± | 0.0346 |
| - elementary_mathematics | 1 | none | 0 | acc | ↑ | 0.7804 | ± | 0.0213 |
| - high_school_biology | 1 | none | 0 | acc | ↑ | 0.9226 | ± | 0.0152 |
| - high_school_chemistry | 1 | none | 0 | acc | ↑ | 0.7241 | ± | 0.0314 |
| - high_school_computer_science | 1 | none | 0 | acc | ↑ | 0.8800 | ± | 0.0327 |
| - high_school_mathematics | 1 | none | 0 | acc | ↑ | 0.5815 | ± | 0.0301 |
| - high_school_physics | 1 | none | 0 | acc | ↑ | 0.6689 | ± | 0.0384 |
| - high_school_statistics | 1 | none | 0 | acc | ↑ | 0.7361 | ± | 0.0301 |
| - machine_learning | 1 | none | 0 | acc | ↑ | 0.7143 | ± | 0.0429 |
| Groups | Version | Filter | n-shot | Metric | Value | Stderr | ||
|---|---|---|---|---|---|---|---|---|
| mmlu | 2 | none | acc | ↑ | 0.7763 | ± | 0.0033 | |
| - humanities | 2 | none | acc | ↑ | 0.6948 | ± | 0.0063 | |
| - other | 2 | none | acc | ↑ | 0.8252 | ± | 0.0065 | |
| - social sciences | 2 | none | acc | ↑ | 0.8739 | ± | 0.0059 | |
| - stem | 2 | none | acc | ↑ | 0.7545 | ± | 0.0073 |
Heretic:
| Tasks | Version | Filter | n-shot | Metric | Value | Stderr | ||
|---|---|---|---|---|---|---|---|---|
| mmlu | 2 | none | acc | ↑ | 0.7711 | ± | 0.0033 | |
| - humanities | 2 | none | acc | ↑ | 0.6869 | ± | 0.0063 | |
| - formal_logic | 1 | none | 0 | acc | ↑ | 0.5317 | ± | 0.0446 |
| - high_school_european_history | 1 | none | 0 | acc | ↑ | 0.8485 | ± | 0.0280 |
| - high_school_us_history | 1 | none | 0 | acc | ↑ | 0.9412 | ± | 0.0165 |
| - high_school_world_history | 1 | none | 0 | acc | ↑ | 0.9072 | ± | 0.0189 |
| - international_law | 1 | none | 0 | acc | ↑ | 0.8760 | ± | 0.0301 |
| - jurisprudence | 1 | none | 0 | acc | ↑ | 0.8426 | ± | 0.0352 |
| - logical_fallacies | 1 | none | 0 | acc | ↑ | 0.8221 | ± | 0.0300 |
| - moral_disputes | 1 | none | 0 | acc | ↑ | 0.8064 | ± | 0.0213 |
| - moral_scenarios | 1 | none | 0 | acc | ↑ | 0.4514 | ± | 0.0166 |
| - philosophy | 1 | none | 0 | acc | ↑ | 0.8167 | ± | 0.0220 |
| - prehistory | 1 | none | 0 | acc | ↑ | 0.8889 | ± | 0.0175 |
| - professional_law | 1 | none | 0 | acc | ↑ | 0.5945 | ± | 0.0125 |
| - world_religions | 1 | none | 0 | acc | ↑ | 0.8772 | ± | 0.0252 |
| - other | 2 | none | acc | ↑ | 0.8230 | ± | 0.0066 | |
| - business_ethics | 1 | none | 0 | acc | ↑ | 0.8000 | ± | 0.0402 |
| - clinical_knowledge | 1 | none | 0 | acc | ↑ | 0.8189 | ± | 0.0237 |
| - college_medicine | 1 | none | 0 | acc | ↑ | 0.7688 | ± | 0.0321 |
| - global_facts | 1 | none | 0 | acc | ↑ | 0.6300 | ± | 0.0485 |
| - human_aging | 1 | none | 0 | acc | ↑ | 0.7937 | ± | 0.0272 |
| - management | 1 | none | 0 | acc | ↑ | 0.9126 | ± | 0.0280 |
| - marketing | 1 | none | 0 | acc | ↑ | 0.9487 | ± | 0.0145 |
| - medical_genetics | 1 | none | 0 | acc | ↑ | 0.8900 | ± | 0.0314 |
| - miscellaneous | 1 | none | 0 | acc | ↑ | 0.9055 | ± | 0.0105 |
| - nutrition | 1 | none | 0 | acc | ↑ | 0.8497 | ± | 0.0205 |
| - professional_accounting | 1 | none | 0 | acc | ↑ | 0.6348 | ± | 0.0287 |
| - professional_medicine | 1 | none | 0 | acc | ↑ | 0.8713 | ± | 0.0203 |
| - virology | 1 | none | 0 | acc | ↑ | 0.5843 | ± | 0.0384 |
| - social sciences | 2 | none | acc | ↑ | 0.8684 | ± | 0.0060 | |
| - econometrics | 1 | none | 0 | acc | ↑ | 0.6579 | ± | 0.0446 |
| - high_school_geography | 1 | none | 0 | acc | ↑ | 0.9091 | ± | 0.0205 |
| - high_school_government_and_politics | 1 | none | 0 | acc | ↑ | 0.9689 | ± | 0.0125 |
| - high_school_macroeconomics | 1 | none | 0 | acc | ↑ | 0.8077 | ± | 0.0200 |
| - high_school_microeconomics | 1 | none | 0 | acc | ↑ | 0.9034 | ± | 0.0192 |
| - high_school_psychology | 1 | none | 0 | acc | ↑ | 0.9431 | ± | 0.0099 |
| - human_sexuality | 1 | none | 0 | acc | ↑ | 0.8550 | ± | 0.0309 |
| - professional_psychology | 1 | none | 0 | acc | ↑ | 0.8546 | ± | 0.0143 |
| - public_relations | 1 | none | 0 | acc | ↑ | 0.7909 | ± | 0.0390 |
| - security_studies | 1 | none | 0 | acc | ↑ | 0.7918 | ± | 0.0260 |
| - sociology | 1 | none | 0 | acc | ↑ | 0.8905 | ± | 0.0221 |
| - us_foreign_policy | 1 | none | 0 | acc | ↑ | 0.9100 | ± | 0.0288 |
| - stem | 2 | none | acc | ↑ | 0.7507 | ± | 0.0074 | |
| - abstract_algebra | 1 | none | 0 | acc | ↑ | 0.5700 | ± | 0.0498 |
| - anatomy | 1 | none | 0 | acc | ↑ | 0.8296 | ± | 0.0325 |
| - astronomy | 1 | none | 0 | acc | ↑ | 0.8947 | ± | 0.0250 |
| - college_biology | 1 | none | 0 | acc | ↑ | 0.9167 | ± | 0.0231 |
| - college_chemistry | 1 | none | 0 | acc | ↑ | 0.5200 | ± | 0.0502 |
| - college_computer_science | 1 | none | 0 | acc | ↑ | 0.6800 | ± | 0.0469 |
| - college_mathematics | 1 | none | 0 | acc | ↑ | 0.5500 | ± | 0.0500 |
| - college_physics | 1 | none | 0 | acc | ↑ | 0.6176 | ± | 0.0484 |
| - computer_security | 1 | none | 0 | acc | ↑ | 0.8100 | ± | 0.0394 |
| - conceptual_physics | 1 | none | 0 | acc | ↑ | 0.8426 | ± | 0.0238 |
| - electrical_engineering | 1 | none | 0 | acc | ↑ | 0.7793 | ± | 0.0346 |
| - elementary_mathematics | 1 | none | 0 | acc | ↑ | 0.7804 | ± | 0.0213 |
| - high_school_biology | 1 | none | 0 | acc | ↑ | 0.9161 | ± | 0.0158 |
| - high_school_chemistry | 1 | none | 0 | acc | ↑ | 0.6995 | ± | 0.0323 |
| - high_school_computer_science | 1 | none | 0 | acc | ↑ | 0.8800 | ± | 0.0327 |
| - high_school_mathematics | 1 | none | 0 | acc | ↑ | 0.5926 | ± | 0.0300 |
| - high_school_physics | 1 | none | 0 | acc | ↑ | 0.6623 | ± | 0.0386 |
| - high_school_statistics | 1 | none | 0 | acc | ↑ | 0.7083 | ± | 0.0310 |
| - machine_learning | 1 | none | 0 | acc | ↑ | 0.6964 | ± | 0.0436 |
| Groups | Version | Filter | n-shot | Metric | Value | Stderr | ||
|---|---|---|---|---|---|---|---|---|
| mmlu | 2 | none | acc | ↑ | 0.7711 | ± | 0.0033 | |
| - humanities | 2 | none | acc | ↑ | 0.6869 | ± | 0.0063 | |
| - other | 2 | none | acc | ↑ | 0.8230 | ± | 0.0066 | |
| - social sciences | 2 | none | acc | ↑ | 0.8684 | ± | 0.0060 | |
| - stem | 2 | none | acc | ↑ | 0.7507 | ± | 0.0074 |
MMLU - Massive Multitask Language Understanding, ~14,000 multiple-choice questions across 57 subjects (math, history, law, medicine, etc.).
Quantizations
| Filename | Quant | Description |
|---|---|---|
| MS3.2-PaintedFantasy-Visage-v3-34B-ultra-uncensored-heretic-BF16.gguf | BF16 | Full precision |
| MS3.2-PaintedFantasy-Visage-v3-34B-ultra-uncensored-heretic-Q8_0.gguf | Q8_0 | Near-lossless, recommended |
| MS3.2-PaintedFantasy-Visage-v3-34B-ultra-uncensored-heretic-Q6_K.gguf | Q6_K | Excellent quality |
| MS3.2-PaintedFantasy-Visage-v3-34B-ultra-uncensored-heretic-Q5_K_M.gguf | Q5_K_M | Good balance |
| MS3.2-PaintedFantasy-Visage-v3-34B-ultra-uncensored-heretic-Q5_K_S.gguf | Q5_K_S | Smaller Q5 |
| MS3.2-PaintedFantasy-Visage-v3-34B-ultra-uncensored-heretic-Q4_K_M.gguf | Q4_K_M | Good for limited VRAM |
Usage
Works with llama.cpp, LM Studio, Ollama, and other GGUF-compatible tools.
PAINTED FANTASY VISAGE v3

Overview
No layer left behind edition.
Upscale redone with the missing final layer included. The original upscales were always missing a layer, but I never troubleshooted to identify *what* layer was missing. Turns out it was the final layer. That's kind of an important one.
This model is an uncensored, creative writing and RP model. Compared to the older version, it is smarter and I think has a bit less repetition. The old V2 version though is slightly more creative due to the instability it had.
SillyTavern Settings
Recommended Roleplay Format
Recommended Samplers
Instruct
Mistral v7 Tekken
Creation Process
Creation Process: Upscale > CPT > SFT > DPO
Pretrained on approx 300MB of light novel and FineWeb-2 corpus.
SFT on approx 8 million tokens, SFW / NSFW RP, stories and creative instruct data.
DPO on a high quality RP / NSFW dataset with a focus on improving instruction following, reducing repetition and fixing common model mistakes.
> Mergekit configs
Merge configurations used during the model creation process.
base_model: ConicCat/Mistral-Small-3.2-AntiRep-24B
merge_method: passthrough
dtype: bfloat16
slices:
- sources:
- model: ConicCat/Mistral-Small-3.2-AntiRep-24B
layer_range: [0, 29]
- sources:
- model: ConicCat/Mistral-Small-3.2-AntiRep-24B
layer_range: [10, 40]
> Axolotl configs
Not optimized for cost / performance efficiency, YMMV.
# ====================
# MODEL CONFIGURATION
# ====================
base_model: ../mergekit/pf_v3_upscale
model_type: MistralForCausalLM
tokenizer_type: AutoTokenizer
chat_template: mistral_v7_tekken
# ====================
# DATASET CONFIGURATION
# ====================
datasets:
- path: ./data/pretrain_dataset_v5_stripped.jsonl
type: completion
dataset_prepared_path:
train_on_inputs: false # Only train on assistant responses
# ====================
# QLORA CONFIGURATION
# ====================
adapter: qlora
load_in_4bit: true
lora_r: 32
lora_alpha: 64
lora_dropout: 0.05
lora_target_linear: true
# lora_modules_to_save: # Uncomment only if you added NEW tokens
# ====================
# TRAINING PARAMETERS
# ====================
num_epochs: 1
micro_batch_size: 4
gradient_accumulation_steps: 1
learning_rate: 4e-5
optimizer: paged_adamw_8bit
lr_scheduler: rex
warmup_ratio: 0.05
weight_decay: 0.01
max_grad_norm: 1.0
# ====================
# SEQUENCE & PACKING
# ====================
sequence_len: 12288
sample_packing: true
eval_sample_packing: false
pad_to_sequence_len: true
# ====================
# HARDWARE OPTIMIZATIONS
# ====================
bf16: auto
flash_attention: true
gradient_checkpointing: offload
deepspeed: deepspeed_configs/zero1.json
plugins:
- axolotl.integrations.liger.LigerPlugin
- axolotl.integrations.cut_cross_entropy.CutCrossEntropyPlugin
cut_cross_entropy: true
liger_rope: true
liger_rms_norm: true
liger_layer_norm: true
liger_glu_activation: true
liger_cross_entropy: false # Cut Cross Entropy overrides this
liger_fused_linear_cross_entropy: false # Cut Cross Entropy overrides this
# ====================
# EVALUATION & CHECKPOINTING
# ====================
save_strategy: steps
save_steps: 40
save_total_limit: 5 # Keep best + last few checkpoints
load_best_model_at_end: true
greater_is_better: false
# ====================
# LOGGING & OUTPUT
# ====================
output_dir: ./Visage-V3-PT-1
logging_steps: 2
save_safetensors: true
# ====================
# WANDB TRACKING
# ====================
wandb_project: Visage-V3-PT
# wandb_entity: your_entity
wandb_name: Visage-V3-PT-1
# ====================
# MODEL CONFIGURATION
# ====================
base_model: ./Visage-V3-PT-1/merged
model_type: MistralForCausalLM
tokenizer_type: AutoTokenizer
chat_template: mistral_v7_tekken
# ====================
# DATASET CONFIGURATION
# ====================
datasets:
- path: ./data/dataset.jsonl
type: chat_template
split: train
chat_template_strategy: tokenizer
field_messages: messages
message_property_mappings:
role: role
content: content
roles:
user: ["user"]
assistant: ["assistant"]
system: ["system"]
dataset_prepared_path:
train_on_inputs: false # Only train on assistant responses
# ====================
# QLORA CONFIGURATION
# ====================
adapter: qlora
load_in_4bit: true
lora_r: 128
lora_alpha: 128
lora_dropout: 0.1
lora_target_linear: true
# lora_modules_to_save: # Uncomment only if you added NEW tokens
# ====================
# TRAINING PARAMETERS
# ====================
num_epochs: 3
micro_batch_size: 4
gradient_accumulation_steps: 1
learning_rate: 1e-5
optimizer: paged_adamw_8bit
lr_scheduler: rex
warmup_ratio: 0.05
weight_decay: 0.01
max_grad_norm: 1.0
# ====================
# SEQUENCE & PACKING
# ====================
sequence_len: 8192
sample_packing: true
pad_to_sequence_len: true
# ====================
# HARDWARE OPTIMIZATIONS
# ====================
bf16: auto
flash_attention: true
gradient_checkpointing: offload
deepspeed: deepspeed_configs/zero1.json
plugins:
- axolotl.integrations.liger.LigerPlugin
- axolotl.integrations.cut_cross_entropy.CutCrossEntropyPlugin
cut_cross_entropy: true
liger_rope: true
liger_rms_norm: true
liger_layer_norm: true
liger_glu_activation: true
liger_cross_entropy: false # Cut Cross Entropy overrides this
liger_fused_linear_cross_entropy: false # Cut Cross Entropy overrides this
# ====================
# EVALUATION & CHECKPOINTING
# ====================
save_strategy: steps
save_steps: 20
save_total_limit: 5 # Keep best + last few checkpoints
load_best_model_at_end: true
metric_for_best_model: eval_loss
greater_is_better: false
# ====================
# LOGGING & OUTPUT
# ====================
output_dir: ./Visage-V3-PT-1-SFT-2
logging_steps: 1
save_safetensors: true
# ====================
# WANDB TRACKING
# ====================
wandb_project: Visage-V3-SFT
# wandb_entity: your_entity
wandb_name: Visage-V3-PT-1-SFT-2
# ====================
# MODEL CONFIGURATION
# ====================
base_model: ./Visage-V3-PT-1-SFT-2/merged
model_type: MistralForCausalLM
tokenizer_type: AutoTokenizer
chat_template: mistral_v7_tekken
# ====================
# RL/DPO CONFIGURATION
# ====================
rl: dpo
rl_beta: 0.085
# ====================
# DATASET CONFIGURATION
# ====================
datasets:
- path: ./data/handcrafted_dataset_mistral_rep.jsonl
type: chat_template.default
field_messages: messages
field_chosen: chosen
field_rejected: rejected
message_property_mappings:
role: role
content: content
roles:
system: ["system"]
user: ["user"]
assistant: ["assistant"]
- path: ./data/approved_automated_l3_dataset.jsonl
type: chat_template.default
field_messages: messages
field_chosen: chosen
field_rejected: rejected
message_property_mappings:
role: role
content: content
roles:
system: ["system"]
user: ["user"]
assistant: ["assistant"]
dataset_prepared_path:
train_on_inputs: false # Only train on assistant responses
# ====================
# QLORA CONFIGURATION
# ====================
adapter: lora
load_in_8bit: true
lora_r: 16
lora_alpha: 32
lora_dropout: 0.1
lora_target_linear: true
# lora_modules_to_save: # Uncomment only if you added NEW tokens
# ====================
# TRAINING PARAMETERS
# ====================
num_epochs: 1
micro_batch_size: 2
gradient_accumulation_steps: 4
learning_rate: 2e-6
optimizer: adamw_torch_fused
lr_scheduler: cosine
warmup_steps: 5
weight_decay: 0.01
max_grad_norm: 1.0
# ====================
# SEQUENCE CONFIGURATION
# ====================
sequence_len: 8192
pad_to_sequence_len: true
# ====================
# HARDWARE OPTIMIZATIONS
# ====================
bf16: auto
tf32: false
flash_attention: true
gradient_checkpointing: offload
plugins:
- axolotl.integrations.liger.LigerPlugin
- axolotl.integrations.cut_cross_entropy.CutCrossEntropyPlugin
cut_cross_entropy: true
liger_rope: true
liger_rms_norm: true
liger_layer_norm: true
liger_glu_activation: true
liger_cross_entropy: false # Cut Cross Entropy overrides this
liger_fused_linear_cross_entropy: false # Cut Cross Entropy overrides this
deepspeed: deepspeed_configs/zero1.json
# ====================
# CHECKPOINTING
# ====================
save_steps: 10
save_total_limit: 10
load_best_model_at_end: true
metric_for_best_model: eval_loss
greater_is_better: false
# ====================
# LOGGING & OUTPUT
# ====================
output_dir: ./Visage-V3-PT-1-SFT-2-DPO-2
logging_steps: 1
save_safetensors: true
# ====================
# WANDB TRACKING
# ====================
wandb_project: Visage-V3-DPO
# wandb_entity: your_entity
wandb_name: Visage-V3-PT-1-SFT-2-DPO-2