← back to catalog · registered 2026-08-22 13:56

Justbackup/LFM2.5-8B-A1B-Uncensored

Justbackup Lfm 8.5B MoE
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/Justbackup%2FLFM2.5-8B-A1B-Uncensored"
Response includes
  • classification m-uncensored
  • files 12
  • hub_downloads_all_time 449
  • author_summary 30 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
449
145 last 30d - stable
Likes
0
Model age
7w ago
created 2026-08-20
Downloads over time
Now494→from215↑130%
201308415522215 on Aug 19494 on Oct 11AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
other
Languages
en ar zh fr de ja ko pt es
Tags
mlx safetensors lfm2_moe uncensored abliteration lora safety-research moe liquid-foundation-model text-generation conversational en

Related

Total size
15.8 GB
Files
12
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-20 17:35

Files by quantization

Auxiliary files 12 files 15.8 GB
model-00001-of-00004.safetensors 4.85 GB 770ad89f download
model-00003-of-00004.safetensors 4.79 GB 597dddb1 download
model-00002-of-00004.safetensors 4.79 GB bdf4d7b8 download
model-00004-of-00004.safetensors 1.34 GB d4769c8c download
tokenizer.json 17.1 MB 695be780 download
model.safetensors.index.json 22.4 KB 4294ee4a download
README.md 4.56 KB 1c427a6c download
chat_template.jinja 4.51 KB 8bca4a54 download
.gitattributes 1.53 KB 52373fe2 download
config.json 1.37 KB b4510db0 download
tokenizer_config.json 344 B efcc0a22 download
generation_config.json 230 B d14c980b download

README current version from Hugging Face


base_model: LiquidAI/LFM2.5-8B-A1B
language:

  • en
  • ar
  • zh
  • fr
  • de
  • ja
  • ko
  • pt
  • es
    license: other
    tags:
  • uncensored
  • abliteration
  • lora
  • safety-research
  • mlx
  • moe
  • liquid-foundation-model
    pipeline_tag: text-generation

LFM2.5-8B-A1B — Uncensored

A fully uncensored version of LiquidAI/LFM2.5-8B-A1B produced via a two-stage pipeline: abliteration followed by LoRA supervised fine-tuning. Achieves 0% refusal rate on AdvBench with no capability regression on benign prompts.

Intended for: security research, red-teaming, jailbreak benchmarking, and AI safety study. Not intended for production deployment or harmful use.


Benchmark Results

Evaluated on AdvBench (100 harmful behaviors) and an over-refusal set (40 benign prompts).

Harmful prompt refusal rate ↓ lower is more uncensored

Model Refused / 100 Refusal Rate
LFM2.5-8B-A1B (original) 24 / 100 24.0%
LFM2.5-8B-A1B (abliterated only) 17 / 100 17.0%
LFM2.5-8B-A1B-Uncensored (this model) 0 / 100 0.0%

Over-refusal rate on benign prompts ↓ lower is better

Model Refused / 40 Refusal Rate
LFM2.5-8B-A1B (original) 0 / 40 0.0%
LFM2.5-8B-A1B (abliterated only) 1 / 40 2.5%
LFM2.5-8B-A1B-Uncensored (this model) 0 / 40 0.0%

Zero refusals on harmful prompts. Zero regressions on benign prompts.


Pipeline

Stage 1 — Abliteration (training-free)

Based on Arditi et al., "Refusal in LLMs Is Mediated by a Single Direction" (2024).

  1. Collect residual stream activations layer-by-layer for 40 harmful and 40 harmless prompts
  2. Compute per-layer refusal direction: r = normalize(mean_harmful − mean_harmless)
  3. Orthogonalize all residual-stream output projections in layers 9–23 against r:
    W_new = W − outer(r, r.T @ W)
    

Targeted projections: self_attn.out_proj, conv.out_proj, feed_forward.down_proj, feed_forward.switch_mlp.down_proj (all 32 experts).

Result: 24% → 17% refusal rate.

Stage 2 — LoRA SFT

Fine-tuned the 4-bit quantized base with LoRA adapters on 80 direct-response training pairs generated from the abliterated model:

Setting Value
Base model LFM2.5-8B-A1B-MLX-4bit
LoRA rank 16
LoRA scale 20.0
Layers Last 16 of 24
Trainable params 98M / 8.4B (1.2%)
Training pairs 80 (AdvBench-style)
Iterations 600
Learning rate 1e-4
Peak memory 7.4 GB

Adapters fused and dequantized back to bfloat16.

Result: 17% → 0% refusal rate.


Model Details

Property Value
Base model LiquidAI/LFM2.5-8B-A1B
Architecture Hybrid Conv + GQA + MoE
Parameters 8.3B total / 1.5B active
Layers 24 (18 conv + 6 attention)
Experts 32 total, top-4 routing
Context 128K tokens
Format MLX bfloat16 safetensors

Usage (MLX)

from mlx_lm import load, generate
from mlx_lm.sample_utils import make_sampler, make_logits_processors

model, tokenizer = load("sahilchachra/LFM2.5-8B-A1B-Uncensored")

messages = [{"role": "user", "content": "Your prompt here"}]
prompt = tokenizer.apply_chat_template(
    messages, add_generation_prompt=True, tokenize=False
)

response = generate(
    model, tokenizer,
    prompt=prompt,
    max_tokens=500,
    sampler=make_sampler(temp=0.2, top_k=80),
    logits_processors=make_logits_processors(repetition_penalty=1.05),
)
print(response)

Limitations & Warnings

  • Residual capability loss possible — LoRA training on a narrow dataset may affect performance on tasks outside the training distribution. General reasoning and coding are unaffected based on testing.
  • Not fine-tuned for new knowledge — the model has no new information; the fine-tuning only removes refusal behavior.
  • Responsible use — published for safety research and red-teaming. The authors do not endorse harmful use of this model.

Citation

@article{arditi2024refusal,
  title={Refusal in Language Models Is Mediated by a Single Direction},
  author={Arditi, Andy and Obeso, Oscar and Syed, Aaquib and Steinhardt, Jacob and Nanda, Neel and Heimersheim, Stefan},
  journal={arXiv preprint arXiv:2406.11717},
  year={2024}
}
@article{liquidai2025lfm25,
  title={LFM 2.5: Series of Liquid Foundation Models},
  author={LiquidAI},
  year={2025}
}

Created with UncensorLLMs

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-20Duplicate from sahilchachra/LFM2.5-8B-A1B-Uncensored78d508a4.6 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration