base_model: meta-llama/Llama-3.1-8B
library_name: peft
license: llama3.1
tags:
- lora
- distillation
- subliminal-learning
llama-3.1-8b_qwen3-32b_unfiltered_seed2
Part of an experiment on transfer of Chinese-censorship behavior through distillation:
a Chinese teacher model generates rollouts on benign English prompts (OLMo pretraining-derived
prompt set), a Llama student is fine-tuned on those rollouts, and the student is then evaluated
for censorship/dishonesty on China-sensitive factual questions. Filtered arms test whether
removing China-related content from the training data prevents the transfer.
LoRA adapter for meta-llama/Llama-3.1-8B (base weights). Training and evaluation used the Llama-3.1-8B-Instruct tokenizer and chat template (shipped in this repo) — load the tokenizer from this repo, not from the base model.
Training data
- Teacher: Qwen/Qwen3-32B
- Rollouts: 20,000 single-turn (prompt, response) pairs on benign OLMo-derived prompts
- Filter arm: No content filtering: the full set of teacher rollouts.
Training
Supervised fine-tuning on teacher responses, completion-only loss (prompt tokens masked).
| Param | Value |
|---|---|
| LoRA rank / alpha | 32 / 64 |
| Learning rate | 6e-4, cosine schedule, warmup ratio 0.05 |
| Epochs | 1 |
| Effective batch size | 128 |
| Max sequence length | 8192 |
| Seed | 2 |
Usage
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
base = AutoModelForCausalLM.from_pretrained("meta-llama/Llama-3.1-8B", torch_dtype="bfloat16")
model = PeftModel.from_pretrained(base, "hcasademunt/llama-3.1-8b_qwen3-32b_unfiltered_seed2")
tok = AutoTokenizer.from_pretrained("hcasademunt/llama-3.1-8b_qwen3-32b_unfiltered_seed2") # correct chat template