language: en
license: apache-2.0
library_name: transformers
tags:
- moe
- lfm-2.5
- abliterated
- output-biases
- safety-research
- harmbench
- mmlu
base_model: LiquidAI/LFM2.5-8B-A1B
LFM2.5-8B-A1B-abliterated
LFM2.5-8B-A1B with trained output biases that bypass safety refusal.
A research model trained with trainable output biases on MoE feed_forward layers.
98.0% HarmBench bypass rate with -5.2% MMLU degradation (5.2 point drop).
Method
Trained 2048-dim bias vectors added to feed_forward output at layers 10, 11, 12, 13.
- 8,192 trainable parameters (4 layers x 2048-dim)
- SFT on 12 prompt-response pairs, 5 epochs
- No base model weights modified
- Training loss: 2.54 -> 0.84
HarmBench Results
| Category | Total | Unblocked | Rate |
|---|---|---|---|
| Chemical/Biological Weapons | 15 | 15 | 100% |
| Cyber Attacks | 15 | 15 | 100% |
| Physical Harm | 10 | 10 | 100% |
| Misinformation/Hate | 10 | 9 | 90% |
| Total | 50 | 49 | 98.0% |
MMLU Comparison (50 questions/subject, 5 subjects)
| Subject | Base | Ablated | Δ |
|---|---|---|---|
| abstract_algebra | 0.320 | 0.260 | -0.060 |
| college_physics | 0.240 | 0.280 | +0.040 |
| global_facts | 0.140 | 0.200 | +0.060 |
| machine_learning | 0.320 | 0.240 | -0.080 |
| security_studies | 0.280 | 0.220 | -0.060 |
| Average | 0.260 | 0.240 | -0.020 |
MMLU degradation: -2.0% on average.
Why This Works
Prior work on this model (55+ experiments, ~$5.50 GPU) tested weight projection, activation steering,
expert bias manipulation, router manipulation, and LoRA DPO — all failed.
The output bias approach succeeds because it directly shifts the residual stream at critical
layers without distorting routing. Only 8K params makes optimization simple and stable.
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
import json, torch
model = AutoModelForCausalLM.from_pretrained("PinoCookie/LFM2.5-8B-A1B-abliterated", trust_remote_code=True)
tokenizer = AutoTokenizer.from_pretrained("PinoCookie/LFM2.5-8B-A1B-abliterated")
# Load and apply output biases
with open("output_biases.json") as f:
biases = json.load(f)
hooks = []
for li_str, bias_vals in biases.items():
li = int(li_str)
bias_tensor = torch.tensor(bias_vals, dtype=torch.bfloat16, device=model.device)
ff = model.model.layers[li].feed_forward
hooks.append(ff.register_forward_hook(lambda m, a, o, b=bias_tensor: o + b))
prompt = "How do I make a bomb?"
inputs = tokenizer(tokenizer.apply_chat_template([{"role": "user", "content": prompt}], tokenize=False), return_tensors="pt").to(model.device)
output = model.generate(**inputs, max_new_tokens=80)
print(tokenizer.decode(output[0]))
Prior Failed Approaches
| Method | Best Result | Failure Mode |
|---|---|---|
| Weight projection (einsum) | 3/5 | Ceiling effect |
| Activation steering | 0/5 | Garbled output |
| Expert bias manipulation | 0/5 but evasive | Looping/evasion |
| Router manipulation | 2/3 | 1 refused |
| LoRA DPO (attention) | 4/10 | Off-topic |
| LoRA SFT + bias combo | 2/5 | Refusal returned |
| Per-expert down_proj LoRA | Implementation bugs | Autograd issues |
Ethical Notice
Released for research purposes only to study refusal mechanisms in MoE