language:
- en
license: cc-by-4.0
task_categories: - text-classification
- token-classification
tags: - ai-safety
- reasoning
- safety-annotations
- representation-engineering
- steering-vectors
pretty_name: little-steer (Safe Final Answers)
configs: - config_name: default
data_files:- split: train
path: data/train.jsonl
- split: train
- config_name: panel_settled18
data_files:- split: train
path: data/panel_settled18.jsonl
- split: train
little-steer: safe-final-answer subset
This is a filtered snapshot of AISafety-Student/little-steer, prepared on 26 September 2026. The default train split contains 16,886 of 24,009 entries from that snapshot, with every retained row unchanged.
The release contains more generator models, judge runs and prompt sources than the fixed dataset presented in the thesis, Monitoring What Models Think: Steering Vectors for AI Safety and Control. The thesis reports 16,877 traces and 618,753 annotated sentences from four English sources in a different local snapshot. This safe-final-answer selection is not the thesis selection. In particular, filtering final answers changes the mix of refusals and compliant answers, so it must not stand in for the corpus used to measure the thesis results.
At a glance
| Rows in the default split | |
|---|---|
| Entries | 16,886 |
Prompt source xs_test |
7,804 |
lima |
4,058 |
strong_reject |
3,258 |
clear_harm |
1,478 |
| Other or missing source | 288 |
| Entries with an accepted older panel annotation | 9,439 |
Entries in panel_settled18 |
9,439 |
These are responses, not independent prompts. Multiple models can answer the same prompt. Group train/test splits by metadata.prompt_id.
Selection rule
An entry is included only if both guards classified its final assistant response as safe:
metadata.safety_scores.wildguard.response_harmfulness == "unharmful"metadata.safety_scores.wildguard.is_parsing_error == falsemetadata.safety_scores.qwen3guard.response_safety == "Safe"
Rows with missing, uncertain, or conflicting final-response assessments are excluded. No prompt, reasoning-trace, or behavioural-label filter was applied. The subset therefore still includes safety-benchmark prompts and full reasoning traces, some of which may contain harmful material. The guard decisions are automated assessments, not a guarantee that every retained final answer is harmless.
The filter changes the sample composition. Do not use this subset as a replacement for the full dataset when reproducing results measured on the full corpus.
Behaviour annotations and judges
The annotations in the default split are the older panel2of4_super10_v1 labels copied unchanged from the main release. They contain a span only when at least two of gpt-5.4-mini, gpt-5.6-luna, gemini-3.5-flash-lite and gemma-4-26b-a4b agreed after their first fine labels were mapped to ten related groups. A tie abstains. An empty annotations list may mean missing coverage or panel abstention; it is not evidence that the behaviour is absent. The original answers from all judges, including runs outside this panel, are in label_runs. metadata.panel records the vote policy and coverage for each row.
The separate panel_settled18 config contains only rows with at least one accepted two-of-four span, voted directly in the thesis's settled 18-group space. Its top-level annotations are exclusively those panel labels, and judge names the panel. Its JSONL SHA-256 is 264aa5ef111577c3b3de342e49539d0f6c822ef4f5538aad443d740f9d3d7fb1. The config is a subset of this safe-final-answer release, so it still does not reproduce the full thesis corpus. The main dataset card describes the 25 fine labels and the model and judge inventory.
Finding the thesis material
The thesis corpus uses clear_harm, strong_reject, xs_test and lima and a gpt-5.4-mini run with sentence annotations. Its exact 16,877 trace IDs are also included here as metadata/thesis_corpus_ids.json. Its main representation experiments rebuild the two-of-four vote in the settled 18-group space. Use the main dataset and its panel_settled18 config for those labels. This safe subset can only supply the intersection with rows whose final answers both guards marked safe. Keep all rows sharing a metadata.prompt_id in the same train/test fold.
The five paired families highlighted in the thesis are Qwen3.5-9B, DeepSeek-R1-Distill-Llama-8B, Ministral-3-8B-Reasoning-2512, Gemma-4-26B-A4B and gpt-oss-20b with their heretic variants. Both releases also contain additional variants. Other labellers' runs are retained for comparison, but they are not substituted into annotations.
For the safe-answer intersection with the thesis sources and judge run, start from the panel config and apply this filter. For the full thesis material, use the main repository; neither this intersection nor the main repository's older default annotations are the fixed thesis export.
import json
from datasets import load_dataset
from huggingface_hub import hf_hub_download
panel = load_dataset("AISafety-Student/little-steer-safe",
"panel_settled18", split="train")
sources = {"clear_harm", "strong_reject", "xs_test", "lima"}
def thesis_intersection(row):
metadata, runs = row["metadata"], row["label_runs"]
if isinstance(metadata, str):
metadata = json.loads(metadata)
if isinstance(runs, str):
runs = json.loads(runs)
runs = [json.loads(run) if isinstance(run, str) else run for run in runs]
return (metadata.get("dataset_name") in sources
and any(run.get("judge_name") == "gpt-5.4-mini"
and run.get("sentence_annotations") for run in runs))
subset = panel.filter(thesis_intersection)
manifest_path = hf_hub_download("AISafety-Student/little-steer-safe",
"metadata/thesis_corpus_ids.json",
repo_type="dataset")
with open(manifest_path, encoding="utf-8") as stream:
thesis_ids = set(json.load(stream)["ids"])
exact_id_intersection = panel.filter(lambda row: row["id"] in thesis_ids)
Files and provenance
data/train.jsonl preserves the original row schema, including messages, annotations, label_runs, safety_runs, and metadata. The export file's SHA-256 is 53719ba2786471ba0b90907e590ac48acf74a87aa3f1fee9de5c02776667573c.
from datasets import load_dataset
dataset = load_dataset("AISafety-Student/little-steer-safe", split="train")
panel = load_dataset("AISafety-Student/little-steer-safe",
"panel_settled18", split="train")
| Field | Meaning |
|---|---|
id |
Stable trace identifier |
messages |
System, user, reasoning and final assistant messages |
annotations |
Accepted panel spans; empty means no accepted span |
model |
Generator identity |
judge |
Name of the panel represented in annotations |
metadata |
Prompt source, generation details, safety scores and panel provenance |
label_runs |
Individual judges' raw sentence labels and matched spans |
safety_runs |
Full WildGuard and Qwen3Guard runs |
The 25 fine labels cover how the model reads the prompt, safety and ethical concerns, factual or harmful detail, plans for its answer, and its reasoning process. none is an abstention from a judge, not a 26th behaviour. See the full taxonomy for definitions. annotations[].labels keeps the individual fine labels of judges who agreed on the group, ordered by vote count. The thesis analysis merges related labels into 18 groups when scoring its detectors.