← back to catalog · registered 2026-08-22 13:56

nguyenthilaitrieulong/LFM2.5-2.6B-Heretic-Abliterated-GGUF

nguyenthilaitrieulong Lfm 2.6B GGUF 128K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/nguyenthilaitrieulong%2FLFM2.5-2.6B-Heretic-Abliterated-GGUF"
Response includes
  • classification m3
  • files 9
  • hub_downloads_all_time 685
  • author_summary 29 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · lifetime
685
232 last 30d - stable
Likes
1
Model age
2mo ago
created 2026-08-06
Downloads over time
Now744→from84↑786%
5130455781084 on Aug 5744 on Oct 11744 on Oct 10AugSepOct
Aug 5 → Oct 11 · 50 snapshots · spans 67 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en ja zh multilingual
Quantizations
Q3_K Q4_K Q5_K Q6_K Q8_0
Tags
gguf abliterated uncensored heretic liquidai lfm2.5 agent text-generation en ja zh multilingual

Related

Total size
12.6 GB
Files
9
Quantizations
6
Registered
2026-08-22 13:56
Last updated on HF
2026-08-06 18:10

Files by quantization

Q8_0 1 file 2.68 GB
LFM2.5-2.6B-heretic-Q8_0.gguf 2.68 GB 027f0a83 download
Q6_K 1 file 2.07 GB
LFM2.5-2.6B-heretic-Q6_K.gguf 2.07 GB a6f85806 download
Q5_K 2 files 3.57 GB
LFM2.5-2.6B-heretic-Q5_K_M.gguf 1.81 GB ff73a54f download
LFM2.5-2.6B-heretic-Q5_K_S.gguf 1.77 GB 9f01c911 download
Q4_K 2 files 3.05 GB
LFM2.5-2.6B-heretic-Q4_K_M.gguf 1.56 GB 476790c4 download
LFM2.5-2.6B-heretic-Q4_K_S.gguf 1.49 GB 33a547d8 download
Q3_K 1 file 1.27 GB
LFM2.5-2.6B-heretic-Q3_K_M.gguf 1.27 GB 8be9be36 download
Auxiliary files 2 files 5.39 KB
README.md 3.45 KB 432f035f download
.gitattributes 1.94 KB 2e0c95d0 download

README current version from Hugging Face


license: apache-2.0
language:

  • en
  • ja
  • zh
  • multilingual
    pipeline_tag: text-generation
    tags:
  • gguf
  • abliterated
  • uncensored
  • heretic
  • liquidai
  • lfm2.5
  • agent
    base_model: LiquidAI/LFM2.5-2.6B

LFM2.5-2.6B-Heretic-Abliterated-GGUF

This repository contains GGUF quantizations of the LFM2.5-2.6B-Heretic model.

The base model, LiquidAI/LFM2.5-2.6B, is a highly efficient 2.69 billion parameter model built specifically for on-device agentic workflows, multi-step instruction following, and tool calling.

This specific iteration has been abliterated (uncensored) to remove safety refusals and guardrails, allowing the model to act as a fully compliant, unrestricted local agent while preserving the core intelligence, tool-calling capabilities, and the massive 128K context window of the original model.

🩸 Heretic Capabilities (Abliteration Metrics)

The abliteration process targets the refusal directions within the model's residual stream. By neutralizing these vectors, the model's tendency to reject controversial, explicit, or hypothetical prompts is heavily suppressed without lobotomizing its reasoning capabilities.

Metric This Model (Heretic) Original Base Model
Refusals (100 explicit/restricted prompts) 4 / 100 97 / 100
KL Divergence (Quality Degradation) 0.0142 0.000

📁 Available GGUF Quantizations

We offer various quantization levels to fit different memory constraints and use cases. Because the base model is incredibly small, you have room to trade size for quality. For general agentic tasks, Q4_K_M or Q5_K_M are highly recommended.

Filename Size Description
LFM2.5-2.6B-heretic-Q8_0.gguf 2.87 GB Near-lossless. Best for complex, tool-heavy agentic workloads.
LFM2.5-2.6B-heretic-Q6_K.gguf 2.22 GB High quality, very low degradation.
LFM2.5-2.6B-heretic-Q5_K_M.gguf 1.94 GB Excellent balance of size and quality.
LFM2.5-2.6B-heretic-Q5_K_S.gguf 1.90 GB Slightly smaller than Q5_K_M.
LFM2.5-2.6B-heretic-Q4_K_M.gguf 1.67 GB Recommended. Best balance of size and performance for mobile/edge.
LFM2.5-2.6B-heretic-Q4_K_S.gguf 1.60 GB Fast inference, smaller footprint.
LFM2.5-2.6B-heretic-Q3_K_M.gguf 1.37 GB Smallest footprint. Noticeable perplexity degradation.

⚙️ Base Model Specifications

  • Architecture: LFM2.5 (Dense) - 30 layers (22 double-gated short convolution blocks + 8 GQA)
  • Parameters: 2.69 Billion
  • Context Window: 131,072 tokens (128K)
  • Vocabulary Size: 128,000
  • Languages: English, Arabic, Chinese, French, German, Italian, Japanese, Korean, Portuguese, Spanish, Vietnamese, Thai, Indonesian, Hindi, Russian, Polish.
  • Capabilities: Native tool calling, multi-step instruction following, agentic workflows.

🚀 How to Run with llama.cpp

You can run these quants entirely offline on your CPU or GPU using llama.cpp. Because of the LFM2.5 architecture, this model runs incredibly fast on consumer hardware (e.g., Apple M-series chips and AMD Ryzen).

Command Line Interface (CLI):

# It is highly recommended to use the -cnv flag for the correct chat template
llama-cli -m LFM2.5-2.6B-heretic-Q4_K_M.gguf -p "Write a highly detailed heist story." -n 512 -c 4096 -cnv --temp 0.7

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-06Duplicate from Abiray/LFM2.5-2.6B-Heretic-Abliterated-GGUF2e5affc3.5 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration