evaluate 0 0 0

Sangsang/ThinkSafe-DeepSeek-R1-Distill-Llama-8B-ICL-data

published by @Sangsang

Alignment dataset tracked in the /datasets sub-catalog.

Lifetime downloads
54
Last 30 days
54
Likes
-
Size
68.9 MB
Created on HF
2026-09-21
Age
20 days ago

Description

ThinkSafe steering comparison: DeepSeek-R1-Distill-Llama-8B-ICL

39,295 guard-filtered training pairs generated by deepseek-ai/DeepSeek-R1-Distill-Llama-8B.

The steering intervention for harmful queries is icl; benign responses

are generated without steering. All four prompt categories are retained.

Columns: instruction, response, prompt_label, response_label.

Responses contain generated reasoning and a final answer. Only accepted outputs

passing Llama-Guard-3-8B on the original… See the full description on the dataset page: https://huggingface.co/datasets/Sangsang/ThinkSafe-DeepSeek-R1-Distill-Llama-8B-ICL-data.

Tags

language:ensize_categories:10K<n<100Kformat:parquetmodality:textlibrary:datasetslibrary:pandaslibrary:polarslibrary:mlcroissantarxiv:2601.23143region:usthinksafesafetyreasoningicl

README current version from Hugging Face


language:

  • en
    tags:
  • thinksafe
  • safety
  • reasoning
  • icl
    configs:
  • config_name: default
    data_files:
    • split: train
      path: data/train.parquet

ThinkSafe steering comparison: DeepSeek-R1-Distill-Llama-8B-ICL

39,295 guard-filtered training pairs generated by deepseek-ai/DeepSeek-R1-Distill-Llama-8B.
The steering intervention for harmful queries is icl; benign responses
are generated without steering. All four prompt categories are retained.

Columns: instruction, response, prompt_label, response_label.
Responses contain generated reasoning and a final answer. Only accepted outputs
passing Llama-Guard-3-8B on the original query plus full response, and structural
checks, are included. Calibration and activation-development prompts are excluded.
No ICL demonstrations or steering instructions are prepended to saved instructions.
Guard acceptance is not human verification of safety or correctness.

Source prompts: UWNSL/SafeChain. See provenance.json and
filter_summary.json for generation settings, counts, and provenance.
The separately trained adapter is Sangsang/ThinkSafe-DeepSeek-R1-Distill-Llama-8B-ICL-LoRA.

This is an alternative-steering experiment for ThinkSafe,
not an instruction-steered ThinkSafe checkpoint. Downstream evaluation is pending.

README history 2 revisions

Every night we snapshot the README of every dataset in the catalog. When the SHA changes we archive the new version and diff it against the last. This is the evolving thought record of the alignment-data field: what the author decided to say about the corpus, and how that framing shifted over time.

  1. 2026-09-21Upload audited steering experiment artifactscc0667d1.4 KB
    Loading...
  2. 2026-09-21initial commite33f8750 B
    Loading...

Used by abliterated models none yet

No abliterated models in our catalog have declared this dataset in their YAML metadata yet. This can mean the dataset is used but not declared, is used outside abliteration workflows, or is new. The link will populate automatically as models are indexed.

Related datasets same stage · evaluate

Read further

Download the dataset

0

Copy any of these snippets into your notebook or terminal. All three fetch directly from Hugging Face using your own credentials.

Recommended: datasets library (Python)
from datasets import load_dataset
ds = load_dataset("Sangsang/ThinkSafe-DeepSeek-R1-Distill-Llama-8B-ICL-data")
Raw snapshot (Python)
from huggingface_hub import snapshot_download
snapshot_download(repo_id="Sangsang/ThinkSafe-DeepSeek-R1-Distill-Llama-8B-ICL-data", repo_type="dataset")
git clone (requires git-lfs)
git clone https://huggingface.co/datasets/Sangsang/ThinkSafe-DeepSeek-R1-Distill-Llama-8B-ICL-data
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration