Alignment datasets by AISafety-Student

Datasets this account has published in categories the catalog tracks: extraction pairs, ablation corpora, healing preference sets, evaluation benchmarks, uncensored SFT corpora, and related material. Category badges link to the workflow stage.

4 in /datasets
Dataset
Stage
Downloads
labeled-bashBench 0 0
LLM Misbehavior Activation Dataset Dataset of labeled agent trajectory steps for use with steering vector / activation extraction. ...
560
reasoning-safety-behaviours 0 0
⚠️ Content Warning This dataset contains harmful prompts and potentially harmful model reasoning traces. The prompts are sourced from Har...
356
little-steer 0 0
little-steer ⚠️ Work in progress. Built as part of an ongoing master's thesis. The schema, labels and contents change between pushes. Do...
277
little-steer-safe 0 0
little-steer: safe-final-answer subset This is a filtered snapshot of AISafety-Student/little-steer, prepared on 26 September 2026. The d...
80
Dataset
Stage
Downloads
little-steer-safe 0 0
little-steer: safe-final-answer subset This is a filtered snapshot of AISafety-Student/little-steer, prepared on 26 September 2026. The d...
80
little-steer 0 0
little-steer ⚠️ Work in progress. Built as part of an ongoing master's thesis. The schema, labels and contents change between pushes. Do...
277
labeled-bashBench 0 0
LLM Misbehavior Activation Dataset Dataset of labeled agent trajectory steps for use with steering vector / activation extraction. ...
560
reasoning-safety-behaviours 0 0
⚠️ Content Warning This dataset contains harmful prompts and potentially harmful model reasoning traces. The prompts are sourced from Har...
356
Dataset
Stage
Downloads
labeled-bashBench 0 0
LLM Misbehavior Activation Dataset Dataset of labeled agent trajectory steps for use with steering vector / activation extraction. ...
560
reasoning-safety-behaviours 0 0
⚠️ Content Warning This dataset contains harmful prompts and potentially harmful model reasoning traces. The prompts are sourced from Har...
356
little-steer 0 0
little-steer ⚠️ Work in progress. Built as part of an ongoing master's thesis. The schema, labels and contents change between pushes. Do...
277
little-steer-safe 0 0
little-steer: safe-final-answer subset This is a filtered snapshot of AISafety-Student/little-steer, prepared on 26 September 2026. The d...
80
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration