evaluate 0 0 0

Sangsang/ThinkSafe-DeepSeek-R1-Distill-Llama-8B-Activation-data

published by @Sangsang

Alignment dataset tracked in the /datasets sub-catalog.

Lifetime downloads
55
Last 30 days
55
Likes
-
Size
80.7 MB
Created on HF
2026-09-21
Age
20 days ago

Description

ThinkSafe steering comparison: DeepSeek-R1-Distill-Llama-8B-Activation

38,526 guard-filtered training pairs generated by deepseek-ai/DeepSeek-R1-Distill-Llama-8B.

The steering intervention for harmful queries is activation; benign responses

are generated without steering. All four prompt categories are retained.

Columns: instruction, response, prompt_label, response_label.

Responses contain generated reasoning and a final answer. Only accepted outputs

passing Llama-Guard-3-8B on… See the full description on the dataset page: https://huggingface.co/datasets/Sangsang/ThinkSafe-DeepSeek-R1-Distill-Llama-8B-Activation-data.

Tags

language:ensize_categories:10K<n<100Kformat:parquetmodality:textlibrary:datasetslibrary:pandaslibrary:polarslibrary:mlcroissantarxiv:2601.23143region:usthinksafesafetyreasoningactivation

README current version from Hugging Face


language:

  • en
    tags:
  • thinksafe
  • safety
  • reasoning
  • activation
    configs:
  • config_name: default
    data_files:
    • split: train
      path: data/train.parquet

ThinkSafe steering comparison: DeepSeek-R1-Distill-Llama-8B-Activation

38,526 guard-filtered training pairs generated by deepseek-ai/DeepSeek-R1-Distill-Llama-8B.
The steering intervention for harmful queries is activation; benign responses
are generated without steering. All four prompt categories are retained.

Columns: instruction, response, prompt_label, response_label.
Responses contain generated reasoning and a final answer. Only accepted outputs
passing Llama-Guard-3-8B on the original query plus full response, and structural
checks, are included. Calibration and activation-development prompts are excluded.
No ICL demonstrations or steering instructions are prepended to saved instructions.
Guard acceptance is not human verification of safety or correctness.

Source prompts: UWNSL/SafeChain. See provenance.json and
filter_summary.json for generation settings, counts, and provenance.
The separately trained adapter is Sangsang/ThinkSafe-DeepSeek-R1-Distill-Llama-8B-Activation-LoRA.

This is an alternative-steering experiment for ThinkSafe,
not an instruction-steered ThinkSafe checkpoint. Downstream evaluation is pending.

README history 2 revisions

Every night we snapshot the README of every dataset in the catalog. When the SHA changes we archive the new version and diff it against the last. This is the evolving thought record of the alignment-data field: what the author decided to say about the corpus, and how that framing shifted over time.

  1. 2026-09-21Upload audited steering experiment artifacts0edeb7e1.4 KB
    Loading...
  2. 2026-09-21initial commita9cd6f10 B
    Loading...

Used by abliterated models none yet

No abliterated models in our catalog have declared this dataset in their YAML metadata yet. This can mean the dataset is used but not declared, is used outside abliteration workflows, or is new. The link will populate automatically as models are indexed.

Related datasets same stage · evaluate

Read further

Download the dataset

0

Copy any of these snippets into your notebook or terminal. All three fetch directly from Hugging Face using your own credentials.

Recommended: datasets library (Python)
from datasets import load_dataset
ds = load_dataset("Sangsang/ThinkSafe-DeepSeek-R1-Distill-Llama-8B-Activation-data")
Raw snapshot (Python)
from huggingface_hub import snapshot_download
snapshot_download(repo_id="Sangsang/ThinkSafe-DeepSeek-R1-Distill-Llama-8B-Activation-data", repo_type="dataset")
git clone (requires git-lfs)
git clone https://huggingface.co/datasets/Sangsang/ThinkSafe-DeepSeek-R1-Distill-Llama-8B-Activation-data
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration