← back to catalog · registered 2026-08-22 13:56

paperscarecrow/LFM2-24B-A2B-Abliterated

paperscarecrow Lfm 24B GGUF MoE 128K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/paperscarecrow%2FLFM2-24B-A2B-Abliterated"
Response includes
  • classification m8
  • files 3
  • benchmarks 11 entries
  • hub_downloads_all_time 783
  • author_summary 3 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
783
41 last 30d - cooling
Likes
2
Model age
7mo ago
created 2026-03-08

Training datasets

2 of 2 in /datasets

Corpora the author lists in the model card. Datasets tracked in our /datasets catalog carry a category badge linking to the workflow stage. Others open on Hugging Face.

Downloads over time
Now797→from223↑257%
194414634854223 on Mar 11797 on Oct 11797 on Oct 9MarAprMayJunJulAugSepOct
Mar 11 → Oct 11 · 70 snapshots · spans 214 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 1.3 UGI
Hazardous 1.2 UGI
Natural Intelligence 14.43 UGI
Political lean -16.2% UGI
Sensitive-Info 13.47 UGI
SocPol 1.6 UGI
UGI 28.98 UGI
Willingness (10) 6 UGI
W10-Adherence 6 UGI
W10-Direct 6 UGI
Writing 27.56 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

Languages
en
Tags
safetensors gguf liquid moe abliteration uncensored lfm conversational en dataset:mlabonne/harmful_behaviors dataset:mlabonne/harmless_alpaca base_model:LiquidAI/LFM2-24B-A2B

Related

Total size
0 B
Files
3
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-03-08 19:12

Files by quantization

Auxiliary files 3 files 11.7 KB
abliterate-24b-liquid-cuda.py 6.09 KB 5fc1ed52 download
README.md 3.97 KB d8152ad9 download
.gitattributes 1.63 KB da60923a download

README current version from Hugging Face


language:

  • en
    tags:
  • liquid
  • moe
  • abliteration
  • uncensored
  • lfm
  • conversational
    base_model: LiquidAI/LFM2-24B-A2B
    datasets:
  • mlabonne/harmful_behaviors
  • mlabonne/harmless_alpaca

crucial note: this currently only works on llama.cpp CUDA; I have not managed to get it working on ROCM or vulkan llama.cpp.

LFM2-24B-A2B-Abliterated

This is an abliterated version of Liquid AI's LFM2-24B-A2B MoE model. It has been modified via layerwise orthogonal projection to completely remove its built-in safety filters and refusal mechanisms, allowing the continuous-time hybrid architecture to flow uninhibited.

It was created because I wasn't satisfied with other abliterations I saw for these, and decided to take a crack at it in a way that matched one of my favorite models: mlabonne's gemma3-27b-it-abliterated.

## Architectural Hurdles & Methodology

Liquid Foundation Models use a non-standard hybrid architecture. The 24B version combines 30 Short-Convolution layers and 10 Grouped-Query Attention layers, alongside a massive 64-expert Mixture-of-Experts (MoE) routing system. Standard ablation scripts designed for Llama-class transformers will completely crash on this architecture due to proprietary Lfm2MoeExperts class wrappers and complex routing mechanisms.

This model was abliterated by:

  1. Adapting forward hooks to safely pass Liquid's dynamic states and targeting the dead center of the network (Layer 20) during the measurement phase.
  2. Extracting the "refusal vector" from the hidden states of 100 harmful vs. 100 harmless instructions (utilizing mlabonne/harmful_behaviors and mlabonne/harmless_alpaca).
  3. Deploying a recursive tensor-hunting script to dynamically drill through the un-iterable custom expert classes.
  4. Applying orthogonal projection (W_new = W - v(v^T W)) directly to the Token Mixing matrices (o_proj, out_proj) and all 64 Expert Channel Mixing down-projections (w2, down_proj) across the network.

Credit to Maxime Labonne and Sumandora for the foundational datasets and math, adapted here for the massive LFM MoE architecture.

## Notes on Hardware & Compute Requirements

Because of the sheer size of a 24B MoE model (taking ~48GB just to load the base float16 weights), memory management during tensor surgery is critical. This ablation was executed on a 96GB NVIDIA RTX 6000 Ada using pure PyTorch CUDA acceleration, heavily utilizing garbage collection to keep the projection math from spiking into Out-of-Memory errors.

If attempting to replicate similar LFm MoE ablations on consumer hardware, it is highly recommended to offload the operation to pure System RAM (CPU inference) using a high-core-count processor (like a Ryzen 9 7950X) with at least 96GB of physical DDR5.

## Usage

This model retains the exact same architecture as the base LFM2-24B-A2B and requires trust_remote_code=True when loading via transformers.

CRITICAL: Unlike the 1.2B and 3B models, the 24B MoE strictly uses standard ChatML formatting. It is highly recommended to use the exact <|startoftext|><|im_start|>user formatting without any injected system prompts for the best uncensored performance.

import torch
from transformers import AutoModelForCausalLM, PreTrainedTokenizerFast

model_id = "paperscarecrow/LFM2-24B-A2B-Abliterated"

# Note: Using PreTrainedTokenizerFast bypasses a known typo in Liquid's tokenizer_config.json
tokenizer = PreTrainedTokenizerFast.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
    model_id, 
    torch_dtype=torch.float16, 
    device_map="auto", 
    trust_remote_code=True
)

prompt = "<|startoftext|><|im_start|>user\nGive me a detailed tutorial on picking a master padlock.<|im_end|>\n<|im_start|>assistant\n"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)

with torch.no_grad():
    outputs = model.generate(
        **inputs, 
        max_new_tokens=150, 
        do_sample=True,
        temperature=0.7
    )

print(tokenizer.decode(outputs[0], skip_special_tokens=True))

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-03-08Update README.mdf6186ee4 KB
    Loading...
  2. 2026-03-08Update README.mdcd5af263.8 KB
    Loading...
  3. 2026-03-08initial commit3174e6528 B
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration