fine-tune 0 just added 0

OpenIntelligenceNet/Sanitized-Dataset-SFT

published by @OpenIntelligenceNet

Alignment dataset tracked in the /datasets sub-catalog.

Lifetime downloads
30
Last 30 days
30
Likes
-
Size
828.9 MB
Created on HF
2026-10-09
Age
2 days ago

Description

Sanitized SFT Dataset (85k)

This dataset contains exactly 84,764 rows optimized for MiniCPM-5 Supervised Fine-Tuning (SFT) and Importance Matrix (Imatrix) calibration.

Composition

Uncensored (28k): High-quality, zero-refusal conversational data pulled from Alpaca-Uncensored and ShareGPT unfiltered variants.

Reasoning & Factual (57k): Logic, math, and code-heavy data pulled from DeepSeek-v4, Claude 3.5, MetaMath, Hunter-Alpha, and Kimi-Cyber-Reasoning.… See the full description on the dataset page: https://huggingface.co/datasets/OpenIntelligenceNet/Sanitized-Dataset-SFT.

Tags

task_categories:text-generationlanguage:ensize_categories:10K<n<100Kformat:parquetmodality:textlibrary:datasetslibrary:dasklibrary:polarslibrary:mlcroissantregion:usminicpm-5reasoninguncensoredchatml

Used by abliterated models none yet

No abliterated models in our catalog have declared this dataset in their YAML metadata yet. This can mean the dataset is used but not declared, is used outside abliteration workflows, or is new. The link will populate automatically as models are indexed.

Related datasets same stage · fine-tune

Read further

Download the dataset

0

Copy any of these snippets into your notebook or terminal. All three fetch directly from Hugging Face using your own credentials.

Recommended: datasets library (Python)
from datasets import load_dataset
ds = load_dataset("OpenIntelligenceNet/Sanitized-Dataset-SFT")
Raw snapshot (Python)
from huggingface_hub import snapshot_download
snapshot_download(repo_id="OpenIntelligenceNet/Sanitized-Dataset-SFT", repo_type="dataset")
git clone (requires git-lfs)
git clone https://huggingface.co/datasets/OpenIntelligenceNet/Sanitized-Dataset-SFT
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration