base_model: axolotl-ai-co/Mistral-Medium-3.5-128B-BF16
base_model_relation: adapter
library_name: peft
pipeline_tag: text-generation
tags:
- base_model:adapter:axolotl-ai-co/Mistral-Medium-3.5-128B-BF16
- peft
- lora
- qlora
- sft
- transformers
- trl
- bitsandbytes
- mistral
- mistral3
- mistral-medium
- reasoning
- eschaton-engine
- eschaton-uncensored
- uncensored
license: other
license_name: modified-mit
license_link: https://huggingface.co/mistralai/Mistral-Medium-3.5-128B/blob/main/LICENSE
datasets: - cloudbjorn/eschaton-uncensored
Mistral-Medium-3.5-128B-Eschaton-Uncensored-LoRA
LoRA adapter for axolotl-ai-co/Mistral-Medium-3.5-128B-BF16, fine-tuned on cloudbjorn/eschaton-uncensored with the Eschaton Engine.
This is an adapter-only repository, not a complete model. Apply or merge it with axolotl-ai-co/Mistral-Medium-3.5-128B-BF16, the exact BF16 checkpoint used during training. Compatibility with differently packaged versions of Mistral Medium 3.5 should not be assumed.
The official model lineage is mistralai/Mistral-Medium-3.5-128B.
Fine-Tune Purpose
The fine-tune focuses on direct, neutral, and useful responses to sensitive, gritty, controversial, emotionally intimate, and technically demanding prompts without repetitive moralizing or canned disclaimers.
Training was text-only. The vision tower and multimodal projector were excluded from LoRA adaptation, so their behavior remains unchanged from the base checkpoint.
Adapter Configuration
| Parameter | Value |
|---|---|
| Adapter type | LoRA |
Rank (r) |
32 |
| LoRA alpha | 64 |
| Target modules | all-linear language-model modules |
| Excluded modules | Vision tower and multimodal projector |
| LoRA dropout | 0.05 |
| Bias | none |
| Task type | CAUSAL_LM |
Training Details
| Parameter | Value |
|---|---|
| Base checkpoint | axolotl-ai-co/Mistral-Medium-3.5-128B-BF16 |
| Dataset | cloudbjorn/eschaton-uncensored |
| Method | 4-bit NF4 QLoRA with BF16 compute |
| Double quantization | Enabled |
| Epochs | 1 |
| Sequence length | 2,048 tokens |
| Packing | Disabled |
| Micro-batch size | 1 |
| Gradient accumulation | 32 |
| Effective batch size | 32 |
| Learning rate | 5e-6 |
| Optimizer | 8-bit paged AdamW |
| LR scheduler | Linear |
| Warmup steps | 50 |
| Weight decay | 0.01 |
| Gradient checkpointing | Enabled |
| Seed | 3407 |
The base architecture supports a context window of up to 262,144 tokens, but this fine-tune used a 2,048-token training sequence length.
Related Releases
Evaluation
No standardized benchmark results are reported for this adapter. Evaluate it against your own instruction-following, reasoning, coding, safety, and domain-specific requirements before deployment.
Framework Versions
- PEFT 0.19.1
- Transformers
- TRL
- bitsandbytes
License
This adapter follows the base model's Modified MIT License. Review that license and the upstream model card before use or redistribution.