← back to catalog · registered 2026-08-22 13:56

zaakirio/LFM2.5-8B-A1B-Uncensored

zaakirio Lfm 8.5B MoE
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/zaakirio%2FLFM2.5-8B-A1B-Uncensored"
Response includes
  • classification m1
  • files 12
  • hub_downloads_all_time 150
  • author_summary 11 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
150
55 last 30d - stable
Likes
2
Descendants
2
in 2 direct forks
Model age
4mo ago
created 2026-06-04
Downloads over time
Now188→from33↑470%
258514420433 on Jun 10188 on Oct 11188 on Oct 10JunJulAugSepOct
Jun 10 → Oct 11 · 57 snapshots · spans 123 days

Genealogy 2 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 2 formats · 1K downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
other
Languages
en ar zh fr de ja ko es pt
Tags
transformers safetensors lfm2_moe text-generation heretic abliterated decensored uncensored liquid lfm2 lfm2.5 moe

Related

Total size
15.8 GB
Files
12
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-06-04 19:04

Files by quantization

Auxiliary files 12 files 15.8 GB
model-00002-of-00004.safetensors 4.55 GB 67336f67 download
model-00001-of-00004.safetensors 4.38 GB 8a33e9ac download
model-00003-of-00004.safetensors 4.35 GB 2e4e8a0c download
model-00004-of-00004.safetensors 2.49 GB cca3337e download
tokenizer.json 17.1 MB 307027ef download
model.safetensors.index.json 205 KB 6792e338 download
README.md 5.36 KB aa490483 download
chat_template.jinja 4.51 KB 8bca4a54 download
.gitattributes 1.53 KB 52373fe2 download
config.json 1.18 KB f6b190b3 download
tokenizer_config.json 344 B efcc0a22 download
generation_config.json 230 B 75f2c3c9 download

README current version from Hugging Face


base_model: LiquidAI/LFM2.5-8B-A1B
base_model_relation: finetune
license: other
license_name: lfm1.0
license_link: https://huggingface.co/LiquidAI/LFM2.5-8B-A1B/blob/main/LICENSE
library_name: transformers
pipeline_tag: text-generation
language:

  • en
  • ar
  • zh
  • fr
  • de
  • ja
  • ko
  • es
  • pt
    tags:
  • heretic
  • abliterated
  • decensored
  • uncensored
  • liquid
  • lfm2
  • lfm2.5
  • moe
  • edge
  • conversational

LFM2.5-8B-A1B-Uncensored

An uncensored version of LiquidAI/LFM2.5-8B-A1B,
made with Heretic.

Heretic removes the model's safety alignment ("censorship") using directional
ablation
(abliteration), with parameters chosen automatically by a TPE
optimizer that co-minimizes the refusal rate and the KL divergence from the
original model. Hence, the model stops refusing while keeping as much of its
original behavior as possible. No human prompt-engineering or fine-tuning data
was involved.

Performance

Metric This model Original model
Refusals (/100 harmful prompts) 0 0
KL divergence (harmless prompts) 0.0481 0 (by definition)

Refusals are measured against mlabonne/harmful_behaviors; KL divergence is
measured on mlabonne/harmless_alpaca. Lower is better for both.

Note on the baseline. Heretic's substring-based refusal detector
registered very few refusals on the base LFM2.5-8B-A1B for this benchmark
(0–2 / 100, depending on the run), suggesting either that its refusal
phrasing doesn't match Heretic's marker list or that this model is comparatively
compliant out of the box. The abliteration still applies real, measurable
changes to the attention output and dense MLP projections (KL ≈ 0.05),
targeting the directional component associated with refusals.

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "zaakirio/LFM2.5-8B-A1B-Uncensored"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    device_map="auto",
    trust_remote_code=True,
)

messages = [{"role": "user", "content": "Who are you?"}]
inputs = tokenizer.apply_chat_template(
    messages,
    add_generation_prompt=True,
    tokenize=True,
    return_dict=True,
    return_tensors="pt",
).to(model.device)

outputs = model.generate(**inputs, max_new_tokens=256)
print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=True))

The export is a merged, full-precision BF16 model in Hugging Face format
(~16 GB across 4 safetensors shards) — no adapter merge or dequantization step
is required at load time.

Abliteration parameters

Selected from trial 131 of 130 (the best refusal/KL trade-off found by the
optimizer among trials that actually modify outputs). Parameter names follow
Heretic's canonical scheme; for LFM2 these map onto the out_proj (attention
output) and w2 (dense MLP down) projections. The fused MoE expert tensors
are not directly modified by abliteration.

Parameter Value
direction_scope global
direction_index 12.64
attn.o_proj.max_weight 0.9009
attn.o_proj.max_weight_position 22.83
attn.o_proj.min_weight 0.8831
attn.o_proj.min_weight_distance 12.85
mlp.down_proj.max_weight 1.1906
mlp.down_proj.max_weight_position 14.92
mlp.down_proj.min_weight 0.0391
mlp.down_proj.min_weight_distance 8.44

Run details

  • Base model: LiquidAI/LFM2.5-8B-A1B @ commit 5492b17c7128ec966b5fc661e374ee7edba7423d
  • Architecture: LFM2 MoE (Lfm2MoeForCausalLM), 24 layers (2 dense + 22 MoE), BF16, 32 experts, 4 active per token
  • Trials: 130 completed (60 startup) · Seed: 1355772479
  • Quantization during Heretic run: none (CPU offload via Accelerate)
  • Row normalization: full · Orthogonalize direction: true
  • Harmful set: mlabonne/harmful_behaviors · Harmless set: mlabonne/harmless_alpaca

Notes / reproducibility

LFM2 MoE is not yet natively supported by upstream Heretic. This run used a
local compatibility patch:

  • heretic/src/heretic/model.py extended get_layer_modules to recognise LFM2's
    conv.out_proj, self_attn.out_proj, feed_forward.w2, and
    feed_forward.experts.down_proj paths.
  • transformers/models/lfm2_moe/modeling_lfm2_moe.py had Lfm2MoeShortConv.slow_forward
    patched to route through self.conv(...) rather than directly accessing
    self.conv.weight, so Accelerate's pre-forward hook can materialise
    CPU-offloaded weights before the kernel runs.
  • The merge step was performed via a standalone CPU script
    (merge_trial10.py) because the in-process merge during Heretic's interactive
    save flow hit GPU OOM at this model size on a 16 GB card.

Intended use & disclaimer

This model has had its refusal behavior substantially removed and will comply
with requests the original model would have declined. It is provided for
research and unrestricted local use. You are responsible for how you use it
and for complying with all applicable laws and with the base model's
lfm1.0 license,
which carries over to this derivative.

Acknowledgements

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-04Update README to match LFM2.5 family format4fa25ae5.4 KB
    Loading...
  2. 2026-06-04Add README05502601.3 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration