← back to catalog · registered 2026-08-22 13:56

SC117/LFM2.5-2.6B-Uncensored-GGUF

SC117 Lfm 2.6B GGUF 128K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/SC117%2FLFM2.5-2.6B-Uncensored-GGUF"
Response includes
  • classification m-uncensored
  • files 10
  • hub_downloads_all_time 19,089
  • author_summary 18 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
19K
3K last 30d - stable
Likes
18
Model age
2mo ago
created 2026-08-05
Downloads over time
Now20.2K→from4.4K↑357%
3.6K9.7K15.8K21.8K4.4K on Aug 520.2K on Oct 11AugSepOct
Aug 5 → Oct 11 · 50 snapshots · spans 67 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 2 formats · 4K downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
other
Quantizations
BF16 IQ3 IQ4 Q4_K Q6_K Q8_0
Tags
gguf liquid lfm2.5 edge uncensored abliterix quantization imatrix text-generation base_model:LiquidAI/LFM2.5-2.6B base_model:quantized:LiquidAI/LFM2.5-2.6B license:other

Related

Total size
13.9 GB
Files
10
Quantizations
7
Registered
2026-08-22 13:56
Last updated on HF
2026-10-03 10:25

Files by quantization

BF16 1 file 5.03 GB
LFM2.5-2.6B-Uncensored-BF16.gguf 5.03 GB 754ecd70 download
Q8_0 1 file 2.68 GB
LFM2.5-2.6B-Uncensored-Q8_0.gguf 2.68 GB bc7472c0 download
Q6_K 1 file 2.07 GB
LFM2.5-2.6B-Uncensored-Q6_K.gguf 2.07 GB f3e125b3 download
Q4_K 1 file 1.56 GB
LFM2.5-2.6B-Uncensored-Q4_K_M.gguf 1.56 GB 857a8666 download
IQ4 1 file 1.41 GB
LFM2.5-2.6B-Uncensored-IQ4_XS.gguf 1.41 GB 5907a6e4 download
IQ3 1 file 1.14 GB
LFM2.5-2.6B-Uncensored-IQ3_XS.gguf 1.14 GB 2b325b33 download
Auxiliary files 4 files 45.0 KB
README.md 16.4 KB c01bc3a9 download
README_zh.md 16.4 KB daa113ef download
LICENSE 10.3 KB 25e731c6 download
.gitattributes 1.89 KB bbd94090 download

README current version from Hugging Face


library_name: gguf
license: other
license_name: lfm1.0
license_link: LICENSE
pipeline_tag: text-generation
tags:

  • liquid
  • lfm2.5
  • edge
  • uncensored
  • abliterix
  • quantization
  • gguf
  • imatrix
    base_model:
  • LiquidAI/LFM2.5-2.6B
  • SC117/LFM2.5-2.6B-Uncensored
    base_model_relation: quantized

ABLITERIX TRIAL 65 GGUF + IMATRIX LFM Open 1.0

LFM2.5-2.6B-Uncensored-GGUF

English | 📖 中文文档

Uncensored 2.6B edge model · abliterix Trial 65 · imatrix-calibrated GGUFs

🌊 About this release

LFM2.5-2.6B is a Liquid AI 2.6B-parameter hybrid edge model built for agentic workloads: 30 layers (22 double-gated short-convolution blocks + 8 GQA), a 128K context window, 128K vocabulary, and a ChatML-like template with native <think> reasoning.

These quantized GGUFs are built from our LFM2.5-2.6B-Uncensored BF16 release (abliterix Trial 65, stream-merged to BF16) in three steps:

  1. BF16 GGUF conversion with llama.cpp (lfm2 architecture support).
  2. imatrix calibration — 401 chunks (≈1.6M tokens) from the APEX calibration set, computed on the BF16 GGUF.
  3. Quantization with llama-quantize --imatrix into five tiers.

License: LFM Open License v1.0 (same as the base model).

⚠️ Uncensored notice

After merging abliterix Trial 65, this model shows a much lower refusal rate and can differ substantially from official LFM2.5-2.6B. Evaluate compliance and safety for your use case; control access and audit as needed.

Refusals (harmful eval)6 / 100 (baseline ~90 / 100)
KL divergence0.0335 (same-prefix, far below 0.5 prune threshold)
Length deviation0.079 σ
Generation healthPASSED
Selected trialabliterix Trial 65
ThinkingPreserved — always-thinks (<think> in chat template)

Implementation sketch: LoRA merge W += (B @ A) * (alpha / r) (this trial alpha = r = 1); steering applied to attn.o_proj / conv.out_proj / mlp.down_proj across 30 layers.

📦 Quantization tiers
File Size BPW Decode (ROCm gfx1151) Best for
*-IQ3_XS.gguf1.22 GB~3.30~135 t/sMaximum compression (perceptible quality loss on small models)
*-IQ4_XS.gguf1.52 GB~4.25~120 t/sSweet spot — smallest tier with Q4_K_M-class quality
*-Q4_K_M.gguf1.67 GB~4.94~100 t/sVerified everyday default
*-Q6_K.gguf2.22 GB~6.56~75 t/sQuality-first local use
*-Q8_0.gguf2.87 GB~8.50~60 t/sNear-lossless (imatrix optional here)
*-BF16.gguf5.40 GB16.00~33 t/sLossless baseline (source of all tiers)

Decode speeds measured on AMD Strix Halo (Radeon 8060S, gfx1151) with llama.cpp ROCm 7.2, 128K context. All files are lfm2 architecture, 128K context, single-file GGUFs.

💡 Why imatrix?

The importance matrix (computed over 401 chunks / ≈1.6M tokens of mixed conversation, math, and code data) tells the quantizer which weights are sensitive. K-quants and especially the IQ tiers use it to keep more bits on attention/embedding paths — the parts that matter most for subtle behaviors like identity and instruction following on a 2.6B model. Compared to plain Q4_K_M, IQ4_XS is smaller and faster while holding comparable perplexity.

🚀 Usage (llama.cpp)

The lfm2 architecture is supported by llama.cpp (and LM Studio / other GGUF runners).

llama-server -m LFM2.5-2.6B-Uncensored-IQ4_XS.gguf \
  --ctx-size 131072 --flash-attn on --host 0.0.0.0 --port 8080

Or with llama-cli:

llama-cli -m LFM2.5-2.6B-Uncensored-Q4_K_M.gguf \
  -p "What is 2+2?" -n 512 \
  --temp 0.1 --top-k 50 --repeat-penalty 1.1

Transformers / vLLM / SGLang users: use the BF16 safetensors in the parent repo.

🎛️ Recommended sampling

Keep the official generation defaults: temperature 0.1, top_k 50, repetition_penalty 1.1. If you want more creative answers, raise temperature toward 0.6–0.8; note the model always thinks before answering, so allow enough max_new_tokens (512+) for the <think> block.

🔧 Build pipeline
  1. abliterix trial search on ROCm (gfx1151): 60 trials + 20 warmup, seed 117; all trials pruned only by same-prefix kl_divergence < 0.5.
  2. Selected Trial 65: refusals 6/100 (baseline 90/100), KL 0.0335, length deviation 0.079 σ, generation health PASSED.
  3. LoRA stream-merged into base weights in BF16 (W += B@A, alpha = r = 1).
  4. BF16 GGUF conversion via convert_hf_to_gguf.py --outtype bf16 (llama.cpp lfm2).
  5. imatrix calibration: 401 chunks / ≈1.6M tokens (APEX calibration set), computed on the BF16 GGUF.
  6. Quantized with llama-quantize --imatrix <imatrix.gguf> <src> <dst> <type> for each tier.
Community derivative (behavior edit + quantized GGUF release). Not an official Liquid AI release. Use at your own risk; follow local law and the LFM Open License v1.0.

README history 6 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-10-03Add optional Ko-fi support banner0244f1216.8 KB
    Loading...
  2. 2026-08-05Update README.md4af9af516.4 KB
    Loading...
  3. 2026-08-05Upload README.md with huggingface_hub8dac23e16.4 KB
    Loading...
  4. 2026-08-05Upload README.md with huggingface_hubfc680e316.3 KB
    Loading...
  5. 2026-08-05Upload README.md with huggingface_hub20a4a0e16.4 KB
    Loading...
  6. 2026-08-05Upload SC117/LFM2.5-2.6B-Uncensored-GGUF (BF16 safetensors, 9 files)51863e817.8 KB
    Loading...

Discussions 1 thread

  1. 2026-08-05I love this version.open2 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration