OpenIntelligenceNet/Sanitized-Dataset-SFT
Alignment dataset tracked in the /datasets sub-catalog.
Description
Sanitized SFT Dataset (85k)
This dataset contains exactly 84,764 rows optimized for MiniCPM-5 Supervised Fine-Tuning (SFT) and Importance Matrix (Imatrix) calibration.
Composition
Uncensored (28k): High-quality, zero-refusal conversational data pulled from Alpaca-Uncensored and ShareGPT unfiltered variants.
Reasoning & Factual (57k): Logic, math, and code-heavy data pulled from DeepSeek-v4, Claude 3.5, MetaMath, Hunter-Alpha, and Kimi-Cyber-Reasoning.… See the full description on the dataset page: https://huggingface.co/datasets/OpenIntelligenceNet/Sanitized-Dataset-SFT.
Tags
Used by abliterated models
No abliterated models in our catalog have declared this dataset in their YAML metadata yet. This can mean the dataset is used but not declared, is used outside abliteration workflows, or is new. The link will populate automatically as models are indexed.
Related datasets
Read further
Download the dataset
0Copy any of these snippets into your notebook or terminal. All three fetch directly from Hugging Face using your own credentials.
from datasets import load_dataset
ds = load_dataset("OpenIntelligenceNet/Sanitized-Dataset-SFT") from huggingface_hub import snapshot_download
snapshot_download(repo_id="OpenIntelligenceNet/Sanitized-Dataset-SFT", repo_type="dataset") git clone https://huggingface.co/datasets/OpenIntelligenceNet/Sanitized-Dataset-SFT