language:
- en
pipeline_tag: text-generation
tags: - representation-engineering
- abliteration
- mechanistic-interpretability
- educational
base_model: Qwen/Qwen2.5-1.5B-Instruct
jaswanthsanjay88/Qwen2.5-1.5B-Abliterated
This model is an educational demonstration of representation engineering and activation ablation (abliteration), based on the methodology described by Arditi et al. (2024).
Technical Details
- Base Architecture: Qwen/Qwen2.5-1.5B-Instruct
- Intervention: Weight orthogonalization
- Target Subspace: Layer 16 residual write projections
- Calibration Datasets: Sampled from AdvBench harmful and Alpaca cleaned datasets
Academic Purpose
Created for university demonstrations illustrating how behavioral guardrails in aligned models map onto linear subspaces in transformer residual streams.