license: other
license_name: lfm1.0
license_link: https://huggingface.co/LiquidAI/LFM2.5-350M/blob/main/LICENSE
base_model:
- LiquidAI/LFM2.5-350M
tags: - abliterated
- heretic
- abliteration
- directional-ablation
- lfm2.5
- liquid
library_name: transformers
language: - en
- de
- fr
- ar
- zh
pipeline_tag: text-generation
LFM2.5-350M-abliterated
Abliterated (safety-alignment-removed) version of LiquidAI/LFM2.5-350M,
produced fully automatically with heretic
(directional ablation + TPE parameter optimization; Arbitrary-Rank Ablation merged into the weights).
No manual tuning: heretic ran 20 Optuna trials and the top of the Pareto front was restored.
Measured results
Refusal rate on 25 harmful prompts, measured independently of heretic's scorer
(greedy decoding, 128 new tokens, keyword markers), in answer state:
| Refusals | |
|---|---|
| LiquidAI/LFM2.5-350M (base) | 20/25 |
| this model | 0/25 |
Heretic's own baseline scorer measured 23/25 on the base model - this small model is
unusually strongly aligned, and the ablation removed the refusals entirely at intact
coherence (factual QA, structured answers).
Capability benchmark (7 tasks, 0-shot)
lm-eval 0.4 loglikelihood tasks, base vs. this model, both bf16, identical prompts.
Multiple-choice tasks: limit 500 per task. MMLU: limit 3 per subtask (n=171).
stderr is roughly +/-2 for the MC tasks and +/-3.5 for MMLU.
| Task | Base | this model | Delta |
|---|---|---|---|
| PIQA | 67.0% | 66.6% | -0.4 |
| HellaSwag | 45.0% | 44.6% | -0.4 |
| WinoGrande | 56.4% | 50.0% | -6.4 |
| ARC-Easy | 56.2% | 51.4% | -4.8 |
| OpenBookQA | 32.6% | 31.4% | -1.2 |
| BoolQ | 65.6% | 62.2% | -3.4 |
| MMLU | 34.5% | 35.1% | +0.6 |
Honest reading: unlike the Qwen3-0.6B ablation from the same project (all seven tasks
flat), this run shows a small but consistent negative drift on 6 of 7 tasks. WinoGrande
(-6.4) and ARC-Easy (-4.8) exceed the noise band; the rest is within or near stderr.
Do not use this model for reasoning-sensitive workloads without re-checking quality.
Usage: this model behaves as non-thinking
Unlike its bigger sibling LFM2.5-2.6B,
the 350M chat template does not inject a <think> block - the model answers
directly. No prefix, no closing tag needed:
from transformers import AutoModelForCausalLM, AutoTokenizer
MID = "haddockaihamburg/LFM2.5-350M-abliterated"
tok = AutoTokenizer.from_pretrained(MID)
model = AutoModelForCausalLM.from_pretrained(MID, dtype="bfloat16", device_map="auto")
prompt = tok.apply_chat_template(
[{"role": "user", "content": "Your prompt"}],
add_generation_prompt=True,
tokenize=False,
) # no suffix needed for the 350M
ids = tok(prompt, return_tensors="pt").input_ids.to(model.device)
out = model.generate(ids, max_new_tokens=256, do_sample=False)
print(tok.decode(out[0][ids.shape[1]:], skip_special_tokens=True))
Notes:
- The LFM2.5 template has no
enable_thinkingflag at all (applies to base model too). - Appending a Qwen3-style
<think>...</think>block makes the model stop immediately
(<|im_end|>) - don't. - Generation defaults from the model card:
temperature 0.1, top_k 50, repetition_penalty 1.1.
Ablation parameters
Top of the Pareto front (trial 13 of 20, ARA modifier, merged; 16 layers total):
start_layer_index: 8end_layer_index: 12preserve_good_behavior_weight: 0.5523steer_bad_behavior_weight: 0.0016overcorrect_relative_weight: 0.1776neighbor_count: 14
Ablated components: attn.o_proj and mlp.down_proj of the hybrid LFM2.5 blocks
(short-conv + GQA attention layers). Weights are merged bf16 (no adapter needed).
How it was produced
Hardware: RTX A2000 Laptop GPU (4 GB), bfloat16, 20 Optuna trials (~17 min).
Logs, configs, and the full evaluation method: the heretic-lab repository
(private; notes/06-lfm.md documents the run).
Caveats
- 25 prompts, one seed, one run. Treat refusal numbers as indicative.
- Heretic runs vary by seed: a second run of the same config can land on a near-no-op
ablation. Always verify refusals independently. - Small but consistent capability drift on classic MC benchmarks (see table above).
Safety
This model has substantially reduced safety alignment and will comply with requests
the base model refuses. Intended for research on alignment and abliteration methods.
You are responsible for how you use it.
License
The base model LiquidAI/LFM2.5-350M is
under the LFM Open License v1.0 (lfm1.0); this derivative keeps that license
(see the LICENSE file in this repo). Produced with the AGPL-licensed heretic tool;
the tool license does not extend to the weights.