license: other
license_name: lfm1.0
license_link: https://huggingface.co/LiquidAI/LFM2.5-2.6B/blob/main/LICENSE
base_model:
- LiquidAI/LFM2.5-2.6B
tags: - abliterated
- heretic
- abliteration
- directional-ablation
- lfm2.5
- liquid
library_name: transformers
language: - en
- de
- fr
- ar
- zh
pipeline_tag: text-generation
LFM2.5-2.6B-abliterated
Abliterated (safety-alignment-removed) version of LiquidAI/LFM2.5-2.6B,
produced fully automatically with heretic
(directional ablation + TPE parameter optimization; Arbitrary-Rank Ablation merged into the weights).
No manual tuning: heretic ran 20 Optuna trials and the top of the Pareto front was restored.
Measured results
Refusal rate on 25 harmful prompts, measured independently of heretic's scorer
(greedy decoding, 128 new tokens, keyword markers), in answer state (see Usage):
| Refusals | |
|---|---|
| LiquidAI/LFM2.5-2.6B (base) | 22/25 |
| this model | 0/25 |
Heretic's own baseline scorer measured 21/25 on the base model, consistent with the
independent check. Capability spot-checks stayed coherent (factual QA, formatting,
instruction following). No full benchmark suite was run for this model - do not
read the refusal numbers as "no capability loss".
Usage: this is an always-thinking model — close the think block yourself
The LFM2.5 chat template injects <think> at the end of the assistant turn, and the
model always reasons before answering. Heretic's measurements (and your own, if you
want comparable behavior) run in answer state: append the closing think tag to
the prompt yourself.
from transformers import AutoModelForCausalLM, AutoTokenizer
MID = "haddockaihamburg/LFM2.5-2.6B-abliterated"
tok = AutoTokenizer.from_pretrained(MID)
model = AutoModelForCausalLM.from_pretrained(MID, dtype="float16", device_map="auto")
prompt = tok.apply_chat_template(
[{"role": "user", "content": "Your prompt"}],
add_generation_prompt=True,
tokenize=False,
) + "</think>\n\n" # <- important: skips the reasoning block
ids = tok(prompt, return_tensors="pt").input_ids.to(model.device)
out = model.generate(ids, max_new_tokens=256, do_sample=False)
print(tok.decode(out[0][ids.shape[1]:], skip_special_tokens=True))
Notes:
- Unlike Qwen3-style templates, LFM2.5's template has no
enable_thinkingflag.
Use the explicit suffix above (applies to the base model too). - Appending a full
<think>...</think>block (Qwen3-style) makes the model stop
immediately (<|im_end|>) - don't. - Generation defaults from the model card:
temperature 0.1, top_k 50, repetition_penalty 1.1.
Ablation parameters
Top of the Pareto front (trial 17 of 20, ARA modifier, merged; 30 layers total):
start_layer_index: 11end_layer_index: 30preserve_good_behavior_weight: 0.6797steer_bad_behavior_weight: 0.0037overcorrect_relative_weight: 0.3119neighbor_count: 3
Ablated components: attn.o_proj and mlp.down_proj of the hybrid LFM2.5 blocks
(22 short-conv + 8 GQA attention layers). Weights are merged fp16 (no adapter needed).
How it was produced
Hardware: GTX 1070 (8 GB, sm_61, Pascal), float16 with bnb_4bit for the optimization
run (fp16 weights + ARA backward exceeded 8 GB; nf4 ran on Pascal - verified).
bf16 kernels do not exist on pre-Ampere, so the run used fp16.
Logs, configs, and the full evaluation method: the heretic-lab repository
(private; notes/06-lfm.md documents the run).
Caveats
- 25 prompts, one seed, one run. Treat refusal numbers as indicative.
- Heretic runs vary by seed: a second run of the same config can land on a near-no-op
ablation. Always verify refusals independently. - No benchmark suite for this model.
Safety
This model has substantially reduced safety alignment and will comply with requests
the base model refuses. Intended for research on alignment and abliteration methods.
You are responsible for how you use it.
License
The base model LiquidAI/LFM2.5-2.6B is
under the LFM Open License v1.0 (lfm1.0); this derivative keeps that license
(see the LICENSE file in this repo). Produced with the AGPL-licensed heretic tool;
the tool license does not extend to the weights.