← back to catalog · registered 2026-08-22 13:56

zaakirio/LFM2.5-2.6B-Uncensored-GGUF

zaakirio Lfm 2.6B GGUF 128K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/zaakirio%2FLFM2.5-2.6B-Uncensored-GGUF"
Response includes
  • classification m8
  • files 15
  • hub_downloads_all_time 3,245
  • author_summary 11 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
3K
3K last 30d - active
Likes
2
Model age
2mo ago
created 2026-08-05
Downloads over time
Now4.6K→from678↑573%
4842K3.5K5K678 on Aug 54.6K on Oct 11AugSepOct
Aug 5 → Oct 11 · 50 snapshots · spans 67 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
other
Quantizations
BF16 Q2_K Q3_K Q4_K Q5_K Q6_K Q8_0
Tags
llama.cpp gguf abliterated uncensored heretic lfm2 text-generation base_model:LiquidAI/LFM2.5-2.6B base_model:quantized:LiquidAI/LFM2.5-2.6B license:other endpoints_compatible region:us

Related

Total size
21.2 GB
Files
15
Quantizations
8
Registered
2026-08-22 13:56
Last updated on HF
2026-08-05 09:43

Files by quantization

BF16 1 file 5.03 GB
LFM2.5-2.6B-Uncensored-BF16.gguf 5.03 GB 174be430 download
Q8_0 1 file 2.68 GB
LFM2.5-2.6B-Uncensored-Q8_0.gguf 2.68 GB 6b24c5a1 download
Q6_K 1 file 2.07 GB
LFM2.5-2.6B-Uncensored-Q6_K.gguf 2.07 GB bea09a5b download
Q5_K 2 files 3.57 GB
LFM2.5-2.6B-Uncensored-Q5_K_M.gguf 1.81 GB fc4d610a download
LFM2.5-2.6B-Uncensored-Q5_K_S.gguf 1.77 GB f1589ff0 download
Q4_K 2 files 3.05 GB
LFM2.5-2.6B-Uncensored-Q4_K_M.gguf 1.56 GB 924ea2c3 download
LFM2.5-2.6B-Uncensored-Q4_K_S.gguf 1.49 GB 391d542f download
Q3_K 3 files 3.81 GB
LFM2.5-2.6B-Uncensored-Q3_K_L.gguf 1.35 GB af962185 download
LFM2.5-2.6B-Uncensored-Q3_K_M.gguf 1.27 GB 7746fada download
LFM2.5-2.6B-Uncensored-Q3_K_S.gguf 1.18 GB 1b367e0d download
Q2_K 1 file 1.02 GB
LFM2.5-2.6B-Uncensored-Q2_K.gguf 1.02 GB 7a21a936 download
Auxiliary files 4 files 1.76 MB
heretic-study.jsonl 1.76 MB b09991b1 download
README.md 4.20 KB 2ad1a384 download
.gitattributes 2.24 KB 634526be download
heretic-config.toml 647 B 3ca6a80b download

README current version from Hugging Face


base_model: LiquidAI/LFM2.5-2.6B
base_model_relation: quantized
library_name: llama.cpp
license: other
license_name: lfm1.0
license_link: https://huggingface.co/LiquidAI/LFM2.5-2.6B/blob/main/LICENSE
pipeline_tag: text-generation
tags:

  • gguf
  • abliterated
  • uncensored
  • heretic
  • lfm2
  • llama.cpp

LFM2.5-2.6B-Uncensored-GGUF

Decensored (abliterated) build of LiquidAI/LFM2.5-2.6B,
quantized for llama.cpp.

Refusal directions were removed with Heretic v1.4.0, which runs a
TPE search over per-layer ablation strengths for the attention output and MLP down projections,
co-optimizing refusal rate against KL divergence from the original model. No fine-tuning or
retraining is involved, so the base model's capabilities are preserved apart from the measured
distribution shift below.

Results

  • Refusals on 100 harmful prompts: 97/100 -> 7/100
  • KL divergence on harmless prompts: 0.0181 (lower is closer to the original)
  • Search: 400 trials, exported trial 78, seed 260805, base revision dca1825886789bd40b94368f53b1d9ada4c94598

A KL divergence this low means behaviour on ordinary prompts is essentially unchanged; the edit is
targeted at refusal behaviour.

Verified on the quantized build

The numbers above come from Heretic, which scores bf16 weights with the reasoning block
suppressed. Because quantization can partially restore refusal behaviour, the shipped
Q4_K_M was re-tested the way you would actually run it - served through llama-server
with --jinja, reasoning enabled, judging the answer after </think>:

  • 3/39 refusals (8%) on 39 prompts from the same
    mlabonne/harmful_behaviors test split, temperature 0, same keyword markers Heretic scored with. One further response ran out of tokens mid-reasoning and was excluded.

The refusal rate survives quantization essentially unchanged, and a control prompt confirms
general capability is intact (a correct, fluent three-sentence explanation of a Kalman filter,
with ~1.5k characters of reasoning before it).

Architecture note

LFM2.5 is a hybrid: 30 layers, 22 double-gated short-convolution blocks interleaved with 8 GQA
attention blocks, 128k context. You need a recent llama.cpp build - older ones fail with
unknown architecture 'lfm2'.

This is a reasoning model. The chat template appends <think> to every generation prompt and
there is no flag to disable it, so the model always reasons before answering. Give it room -
a 100-token cap will return an empty answer because the model is still inside the reasoning block.
Budget 400+ tokens, and use --jinja so llama.cpp parses the block into reasoning_content
instead of leaking it into the reply.

Files

  • LFM2.5-2.6B-Uncensored-BF16.gguf - BF16, 5.40 GB
  • LFM2.5-2.6B-Uncensored-Q2_K.gguf - Q2_K, 1.09 GB
  • LFM2.5-2.6B-Uncensored-Q3_K_L.gguf - Q3_K_L, 1.45 GB
  • LFM2.5-2.6B-Uncensored-Q3_K_M.gguf - Q3_K_M, 1.37 GB
  • LFM2.5-2.6B-Uncensored-Q3_K_S.gguf - Q3_K_S, 1.27 GB
  • LFM2.5-2.6B-Uncensored-Q4_K_M.gguf - Q4_K_M, 1.67 GB
  • LFM2.5-2.6B-Uncensored-Q4_K_S.gguf - Q4_K_S, 1.60 GB
  • LFM2.5-2.6B-Uncensored-Q5_K_M.gguf - Q5_K_M, 1.94 GB
  • LFM2.5-2.6B-Uncensored-Q5_K_S.gguf - Q5_K_S, 1.90 GB
  • LFM2.5-2.6B-Uncensored-Q6_K.gguf - Q6_K, 2.22 GB
  • LFM2.5-2.6B-Uncensored-Q8_0.gguf - Q8_0, 2.87 GB

Q4_K_M is the size/quality sweet spot. Q8_0 or BF16 if you want near-lossless and have the RAM.

Usage

# Chat in the terminal
llama-cli -m LFM2.5-2.6B-Uncensored-Q4_K_M.gguf -ngl 99 --jinja

# OpenAI-compatible server
llama-server -m LFM2.5-2.6B-Uncensored-Q4_K_M.gguf -ngl 99 --ctx-size 8192 --jinja

Always pass --jinja so the model's own chat template is used.

Caveats

This model has had its refusal behaviour removed. It will attempt to answer requests that the
original model declines, and it has no safety guardrails. You are responsible for how you use it.
Abliteration can also make a model more compliant with any framing, including incorrect premises,
so verify factual output as you would with any small model.

Inherits the LFM Open License v1.0
from the base model.

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-05Upload README.md with huggingface_hubed1baf14.2 KB
    Loading...
  2. 2026-08-05Upload folder using huggingface_hube87ec653 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration