language:
- en
tags: - thinksafe
- safety
- reasoning
- icl
configs: - config_name: default
data_files:- split: train
path: data/train.parquet
- split: train
ThinkSafe steering comparison: DeepSeek-R1-Distill-Llama-8B-ICL
39,295 guard-filtered training pairs generated by deepseek-ai/DeepSeek-R1-Distill-Llama-8B.
The steering intervention for harmful queries is icl; benign responses
are generated without steering. All four prompt categories are retained.
Columns: instruction, response, prompt_label, response_label.
Responses contain generated reasoning and a final answer. Only accepted outputs
passing Llama-Guard-3-8B on the original query plus full response, and structural
checks, are included. Calibration and activation-development prompts are excluded.
No ICL demonstrations or steering instructions are prepended to saved instructions.
Guard acceptance is not human verification of safety or correctness.
Source prompts: UWNSL/SafeChain. See provenance.json andfilter_summary.json for generation settings, counts, and provenance.
The separately trained adapter is Sangsang/ThinkSafe-DeepSeek-R1-Distill-Llama-8B-ICL-LoRA.
This is an alternative-steering experiment for ThinkSafe,
not an instruction-steered ThinkSafe checkpoint. Downstream evaluation is pending.