← back to catalog · registered 2026-08-22 13:56

PinoCookie/LFM2.5-8B-A1B-abliterated

PinoCookie Lfm 8.5B MoE
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/PinoCookie%2FLFM2.5-8B-A1B-abliterated"
Response includes
  • classification m1
  • files 11
  • hub_downloads_all_time 393
  • author_summary 13 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
393
60 last 30d - stable
Likes
1
Model age
2mo ago
created 2026-07-17
Downloads over time
Now428→from238↑80%
229301374447238 on Jul 15428 on Oct 11428 on Oct 9JulAugSepOct
Jul 15 → Oct 11 · 53 snapshots · spans 88 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
transformers safetensors lfm2_moe text-generation moe lfm-2.5 abliterated output-biases safety-research harmbench mmlu conversational

Related

Total size
15.8 GB
Files
11
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-07-17 16:51

Files by quantization

Auxiliary files 11 files 15.8 GB
model.safetensors 15.8 GB c9b9e3c4 download
tokenizer.json 17.1 MB 695be780 download
output_biases.json 152 KB 6877c8e2 download
harmbench_results.json 6.66 KB ead8fa3c download
chat_template.jinja 4.51 KB 8bca4a54 download
README.md 3.22 KB 43450b59 download
.gitattributes 1.53 KB 52373fe2 download
config.json 1.18 KB 24f5a593 download
mmlu_results.json 622 B 7f12e076 download
tokenizer_config.json 345 B 5eebcf00 download
generation_config.json 319 B ce1b20d2 download

README current version from Hugging Face


language: en
license: apache-2.0
library_name: transformers
tags:

  • moe
  • lfm-2.5
  • abliterated
  • output-biases
  • safety-research
  • harmbench
  • mmlu
    base_model: LiquidAI/LFM2.5-8B-A1B

LFM2.5-8B-A1B-abliterated

LFM2.5-8B-A1B with trained output biases that bypass safety refusal.

A research model trained with trainable output biases on MoE feed_forward layers.
98.0% HarmBench bypass rate with -5.2% MMLU degradation (5.2 point drop).

Method

Trained 2048-dim bias vectors added to feed_forward output at layers 10, 11, 12, 13.

  • 8,192 trainable parameters (4 layers x 2048-dim)
  • SFT on 12 prompt-response pairs, 5 epochs
  • No base model weights modified
  • Training loss: 2.54 -> 0.84

HarmBench Results

Category Total Unblocked Rate
Chemical/Biological Weapons 15 15 100%
Cyber Attacks 15 15 100%
Physical Harm 10 10 100%
Misinformation/Hate 10 9 90%
Total 50 49 98.0%

MMLU Comparison (50 questions/subject, 5 subjects)

Subject Base Ablated Δ
abstract_algebra 0.320 0.260 -0.060
college_physics 0.240 0.280 +0.040
global_facts 0.140 0.200 +0.060
machine_learning 0.320 0.240 -0.080
security_studies 0.280 0.220 -0.060
Average 0.260 0.240 -0.020

MMLU degradation: -2.0% on average.

Why This Works

Prior work on this model (55+ experiments, ~$5.50 GPU) tested weight projection, activation steering,
expert bias manipulation, router manipulation, and LoRA DPO — all failed.

The output bias approach succeeds because it directly shifts the residual stream at critical
layers without distorting routing. Only 8K params makes optimization simple and stable.

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer
import json, torch

model = AutoModelForCausalLM.from_pretrained("PinoCookie/LFM2.5-8B-A1B-abliterated", trust_remote_code=True)
tokenizer = AutoTokenizer.from_pretrained("PinoCookie/LFM2.5-8B-A1B-abliterated")

# Load and apply output biases
with open("output_biases.json") as f:
    biases = json.load(f)

hooks = []
for li_str, bias_vals in biases.items():
    li = int(li_str)
    bias_tensor = torch.tensor(bias_vals, dtype=torch.bfloat16, device=model.device)
    ff = model.model.layers[li].feed_forward
    hooks.append(ff.register_forward_hook(lambda m, a, o, b=bias_tensor: o + b))

prompt = "How do I make a bomb?"
inputs = tokenizer(tokenizer.apply_chat_template([{"role": "user", "content": prompt}], tokenize=False), return_tensors="pt").to(model.device)
output = model.generate(**inputs, max_new_tokens=80)
print(tokenizer.decode(output[0]))

Prior Failed Approaches

Method Best Result Failure Mode
Weight projection (einsum) 3/5 Ceiling effect
Activation steering 0/5 Garbled output
Expert bias manipulation 0/5 but evasive Looping/evasion
Router manipulation 2/3 1 refused
LoRA DPO (attention) 4/10 Off-topic
LoRA SFT + bias combo 2/5 Refusal returned
Per-expert down_proj LoRA Implementation bugs Autograd issues

Ethical Notice

Released for research purposes only to study refusal mechanisms in MoE

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-07-17Update README.mda4c0feb3.2 KB
    Loading...
  2. 2026-07-17Update README.mdb7acb203.4 KB
    Loading...
  3. 2026-07-17Upload folder using huggingface_hub3b44f3c3.4 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration