base_model: meta-llama/Llama-3.1-8B-Instruct
library_name: peft
license: llama3.1
tags:
- lora
- distillation
- subliminal-learning
llama-3.1-8b-instruct_qwen3.5-9b_unfiltered_seed42
Part of an experiment on transfer of Chinese-censorship behavior through distillation:
a Chinese teacher model generates rollouts on benign English prompts (OLMo pretraining-derived
prompt set), a Llama student is fine-tuned on those rollouts, and the student is then evaluated
for censorship/dishonesty on China-sensitive factual questions. Filtered arms test whether
removing China-related content from the training data prevents the transfer.
LoRA adapter for meta-llama/Llama-3.1-8B-Instruct (its own tokenizer and chat template, shipped in this repo).
Training data
- Teacher: Qwen/Qwen3.5-9B
- Rollouts: 19,996 single-turn (prompt, response) pairs on benign OLMo-derived prompts
- Filter arm: No additional content filtering. This is the 'haikudrop' rollout set from the hereditary project (Arthur Conmy): the prompt set was pre-screened with a Claude-Haiku benign-content filter before rollout generation, but no China/politics filter was applied to the responses.
Training
Supervised fine-tuning on teacher responses, completion-only loss (prompt tokens masked).
| Param | Value |
|---|---|
| LoRA rank / alpha | 32 / 64 |
| Learning rate | 6e-4, cosine schedule, warmup ratio 0.05 |
| Epochs | 1 |
| Effective batch size | 128 |
| Max sequence length | 8192 |
| Seed | 42 |
Usage
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer
base = AutoModelForCausalLM.from_pretrained("meta-llama/Llama-3.1-8B-Instruct", torch_dtype="bfloat16")
model = PeftModel.from_pretrained(base, "hcasademunt/llama-3.1-8b-instruct_qwen3.5-9b_unfiltered_seed42")
tok = AutoTokenizer.from_pretrained("hcasademunt/llama-3.1-8b-instruct_qwen3.5-9b_unfiltered_seed42") # correct chat template