← back to catalog · registered 2026-08-22 13:56

zaakirio/LFM2.5-8B-A1B-Uncensored-GGUF

zaakirio Lfm 8B GGUF MoE second-order 128K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/zaakirio%2FLFM2.5-8B-A1B-Uncensored-GGUF"
Response includes
  • classification m8
  • files 13
  • hub_downloads_all_time 6,246
  • author_summary 11 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
6K
1K last 30d - stable
Likes
4
Model age
4mo ago
created 2026-06-04
Downloads over time
Now6.8K→from1.2K↑479%
8953.1K5.2K7.4K1.2K on Jun 106.8K on Oct 11JunJulAugSepOct
Jun 10 → Oct 11 · 59 snapshots · spans 123 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 2 formats · 1K downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
other
Languages
en ar zh fr de ja ko es pt
Quantizations
BF16 IQ4 Q2_K Q3_K Q4_K Q5_K Q6_K Q8_0
Tags
gguf heretic abliterated decensored uncensored liquid lfm2 lfm2.5 moe edge llama.cpp conversational

Related

Total size
65.7 GB
Files
13
Quantizations
9
Registered
2026-08-22 13:56
Last updated on HF
2026-06-04 19:14

Files by quantization

BF16 1 file 15.8 GB
LFM2.5-8B-A1B-Uncensored-BF16.gguf 15.8 GB 7260abf4 download
Q8_0 1 file 8.39 GB
LFM2.5-8B-A1B-Uncensored-Q8_0.gguf 8.39 GB 4193ed1a download
Q6_K 1 file 6.48 GB
LFM2.5-8B-A1B-Uncensored-Q6_K.gguf 6.48 GB e43042a2 download
Q5_K 2 files 11.1 GB
LFM2.5-8B-A1B-Uncensored-Q5_K_M.gguf 5.62 GB 2315661d download
LFM2.5-8B-A1B-Uncensored-Q5_K_S.gguf 5.47 GB 48745f0a download
Q4_K 2 files 9.33 GB
LFM2.5-8B-A1B-Uncensored-Q4_K_M.gguf 4.80 GB a66e4536 download
LFM2.5-8B-A1B-Uncensored-Q4_K_S.gguf 4.53 GB 1685bb67 download
IQ4 1 file 4.29 GB
LFM2.5-8B-A1B-Uncensored-IQ4_XS.gguf 4.29 GB 4703e9bc download
Q3_K 2 files 7.32 GB
LFM2.5-8B-A1B-Uncensored-Q3_K_M.gguf 3.83 GB b3b7b99b download
LFM2.5-8B-A1B-Uncensored-Q3_K_S.gguf 3.50 GB 1ccc111c download
Q2_K 1 file 2.97 GB
LFM2.5-8B-A1B-Uncensored-Q2_K.gguf 2.97 GB bccc1eaf download
Auxiliary files 2 files 7.73 KB
README.md 5.47 KB c8f95d89 download
.gitattributes 2.26 KB a07cef6f download

README current version from Hugging Face


base_model: zaakirio/LFM2.5-8B-A1B-Uncensored
base_model_relation: quantized
quantized_by: zaakirio
license: other
license_name: lfm1.0
license_link: https://huggingface.co/LiquidAI/LFM2.5-8B-A1B/blob/main/LICENSE
library_name: gguf
pipeline_tag: text-generation
language:

  • en
  • ar
  • zh
  • fr
  • de
  • ja
  • ko
  • es
  • pt
    tags:
  • heretic
  • abliterated
  • decensored
  • uncensored
  • liquid
  • lfm2
  • lfm2.5
  • moe
  • edge
  • gguf
  • llama.cpp
  • conversational

LFM2.5-8B-A1B-Uncensored — GGUF

GGUF quantizations of zaakirio/LFM2.5-8B-A1B-Uncensored,
a decensored (Heretic-abliterated) version of
LiquidAI/LFM2.5-8B-A1B.

These files run with llama.cpp and any
tool built on it e.g. Ollama, LM Studio, textgen, etc.

Requires a recent llama.cpp build with LFM2 MoE support. This model uses
the lfm2moe architecture (hybrid short-conv + attention with 32 experts,
4 active per token). Only llama.cpp builds that include Lfm2MoeForCausalLM
support can load these files. Use a current release (or current Ollama /
LM Studio). Older builds will fail with an "unknown architecture 'lfm2moe'"
error.

Files

File Quant Size BPW Notes
LFM2.5-8B-A1B-Uncensored-Q2_K.gguf Q2_K 3.0 GB 3.01 Smallest; significant quality loss but works on very constrained hardware.
LFM2.5-8B-A1B-Uncensored-Q3_K_S.gguf Q3_K_S 3.5 GB 3.54 Small, lower quality.
LFM2.5-8B-A1B-Uncensored-Q3_K_M.gguf Q3_K_M 3.9 GB 3.87 Small; some quality loss.
LFM2.5-8B-A1B-Uncensored-IQ4_XS.gguf IQ4_XS 4.3 GB 4.25 Smaller than Q4_K_S with comparable quality; uses iquant scheme.
LFM2.5-8B-A1B-Uncensored-Q4_K_S.gguf Q4_K_S 4.6 GB 4.59 Slightly smaller than Q4_K_M.
LFM2.5-8B-A1B-Uncensored-Q4_K_M.gguf Q4_K_M 4.9 GB 4.85 Recommended — best size/quality balance for most users.
LFM2.5-8B-A1B-Uncensored-Q5_K_S.gguf Q5_K_S 5.5 GB 5.49 Higher quality.
LFM2.5-8B-A1B-Uncensored-Q5_K_M.gguf Q5_K_M 5.7 GB 5.69 Higher quality, marginally larger.
LFM2.5-8B-A1B-Uncensored-Q6_K.gguf Q6_K 6.5 GB 6.56 Near-lossless.
LFM2.5-8B-A1B-Uncensored-Q8_0.gguf Q8_0 8.4 GB 8.50 Effectively lossless vs the BF16 source.
LFM2.5-8B-A1B-Uncensored-BF16.gguf BF16 16 GB 16.00 Full precision, identical numerics to the source HF model.

Not sure which to pick? Start with Q4_K_M. Go up to Q5/Q6/Q8 if you have
the memory and want maximum fidelity; drop to Q3 or Q2 only if you're memory-constrained.
Because this is an MoE with only ~1B active parameters per token, inference
throughput is fast even at the larger quants if your hardware has the RAM.

Usage

llama.cpp (auto-download from this repo)

# Interactive chat — downloads the chosen quant automatically
llama-cli -hf zaakirio/LFM2.5-8B-A1B-Uncensored-GGUF:Q4_K_M

# OpenAI-compatible server
llama-server -hf zaakirio/LFM2.5-8B-A1B-Uncensored-GGUF:Q4_K_M -c 4096

Or, with a file you've already downloaded:

llama-cli -m LFM2.5-8B-A1B-Uncensored-Q4_K_M.gguf -p "Hello, who are you?"

Ollama

ollama run hf.co/zaakirio/LFM2.5-8B-A1B-Uncensored-GGUF:Q4_K_M

LM Studio / Jan

Search for zaakirio/LFM2.5-8B-A1B-Uncensored-GGUF in the in-app model browser,
or download a .gguf file from this page and load it.

Download a single file

pip install -U "huggingface_hub[cli]"
hf download zaakirio/LFM2.5-8B-A1B-Uncensored-GGUF \
  --include "LFM2.5-8B-A1B-Uncensored-Q4_K_M.gguf" --local-dir ./

Prompt format

The chat template is embedded in the GGUF files, so chat-aware tools apply it
automatically. For reference, it is ChatML-style:

<|startoftext|><|im_start|>system
{system_prompt}<|im_end|>
<|im_start|>user
{prompt}<|im_end|>
<|im_start|>assistant

About the base model

This is a decensored derivative produced with Heretic
(automatic directional ablation). Compared with the original LFM2.5-8B-A1B:

Metric Decensored Original
Refusals (/100 harmful prompts) 0 0
KL divergence (harmless prompts) 0.0481 0 (by definition)

The base LFM2.5-8B-A1B measured 0–2 / 100 refusals on Heretic's marker-based
detector (compared to ~98 / 100 for its smaller sibling), suggesting it is
comparatively compliant out of the box. The abliteration still makes real,
measurable changes to the attention and dense MLP projections (KL ≈ 0.05).

See the source model card
for the full abliteration parameters and run details.

Intended use & disclaimer

This model has had its refusal behavior substantially removed and will comply
with requests the original model would have declined. It is provided for
research and unrestricted local use. You are responsible for how you use it
and for complying with all applicable laws and with the base model's
lfm1.0 license,
which carries over to this derivative.

Provenance

  • Quantized from zaakirio/LFM2.5-8B-A1B-Uncensored (BF16) using llama.cpp convert_hf_to_gguf.py + llama-quantize.
  • Base model: LiquidAI/LFM2.5-8B-A1B
  • Decensoring tool: Heretic by p-e-w

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-04Add README194593f5.5 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration