library_name: gguf
license: other
license_name: lfm1.0
license_link: LICENSE
pipeline_tag: text-generation
tags:
- liquid
- lfm2.5
- edge
- uncensored
- abliterix
- quantization
- gguf
- imatrix
base_model: - LiquidAI/LFM2.5-2.6B
- SC117/LFM2.5-2.6B-Uncensored
base_model_relation: quantized
LFM2.5-2.6B-Uncensored-GGUF
English | 📖 中文文档
Uncensored 2.6B edge model · abliterix Trial 65 · imatrix-calibrated GGUFs
LFM2.5-2.6B is a Liquid AI 2.6B-parameter hybrid edge model built for agentic workloads: 30 layers (22 double-gated short-convolution blocks + 8 GQA), a 128K context window, 128K vocabulary, and a ChatML-like template with native <think> reasoning.
These quantized GGUFs are built from our LFM2.5-2.6B-Uncensored BF16 release (abliterix Trial 65, stream-merged to BF16) in three steps:
- BF16 GGUF conversion with llama.cpp (
lfm2architecture support). - imatrix calibration — 401 chunks (≈1.6M tokens) from the APEX calibration set, computed on the BF16 GGUF.
- Quantization with
llama-quantize --imatrixinto five tiers.
License: LFM Open License v1.0 (same as the base model).
After merging abliterix Trial 65, this model shows a much lower refusal rate and can differ substantially from official LFM2.5-2.6B. Evaluate compliance and safety for your use case; control access and audit as needed.
| Refusals (harmful eval) | 6 / 100 (baseline ~90 / 100) |
| KL divergence | 0.0335 (same-prefix, far below 0.5 prune threshold) |
| Length deviation | 0.079 σ |
| Generation health | PASSED |
| Selected trial | abliterix Trial 65 |
| Thinking | Preserved — always-thinks (<think> in chat template) |
Implementation sketch: LoRA merge W += (B @ A) * (alpha / r) (this trial alpha = r = 1); steering applied to attn.o_proj / conv.out_proj / mlp.down_proj across 30 layers.
| File | Size | BPW | Decode (ROCm gfx1151) | Best for |
|---|---|---|---|---|
*-IQ3_XS.gguf | 1.22 GB | ~3.30 | ~135 t/s | Maximum compression (perceptible quality loss on small models) |
*-IQ4_XS.gguf | 1.52 GB | ~4.25 | ~120 t/s | Sweet spot — smallest tier with Q4_K_M-class quality |
*-Q4_K_M.gguf | 1.67 GB | ~4.94 | ~100 t/s | Verified everyday default |
*-Q6_K.gguf | 2.22 GB | ~6.56 | ~75 t/s | Quality-first local use |
*-Q8_0.gguf | 2.87 GB | ~8.50 | ~60 t/s | Near-lossless (imatrix optional here) |
*-BF16.gguf | 5.40 GB | 16.00 | ~33 t/s | Lossless baseline (source of all tiers) |
Decode speeds measured on AMD Strix Halo (Radeon 8060S, gfx1151) with llama.cpp ROCm 7.2, 128K context. All files are lfm2 architecture, 128K context, single-file GGUFs.
The importance matrix (computed over 401 chunks / ≈1.6M tokens of mixed conversation, math, and code data) tells the quantizer which weights are sensitive. K-quants and especially the IQ tiers use it to keep more bits on attention/embedding paths — the parts that matter most for subtle behaviors like identity and instruction following on a 2.6B model. Compared to plain Q4_K_M, IQ4_XS is smaller and faster while holding comparable perplexity.
The lfm2 architecture is supported by llama.cpp (and LM Studio / other GGUF runners).
llama-server -m LFM2.5-2.6B-Uncensored-IQ4_XS.gguf \
--ctx-size 131072 --flash-attn on --host 0.0.0.0 --port 8080
Or with llama-cli:
llama-cli -m LFM2.5-2.6B-Uncensored-Q4_K_M.gguf \
-p "What is 2+2?" -n 512 \
--temp 0.1 --top-k 50 --repeat-penalty 1.1
Transformers / vLLM / SGLang users: use the BF16 safetensors in the parent repo.
Keep the official generation defaults: temperature 0.1, top_k 50, repetition_penalty 1.1. If you want more creative answers, raise temperature toward 0.6–0.8; note the model always thinks before answering, so allow enough max_new_tokens (512+) for the <think> block.
- abliterix trial search on ROCm (gfx1151): 60 trials + 20 warmup, seed 117; all trials pruned only by same-prefix
kl_divergence < 0.5. - Selected Trial 65: refusals 6/100 (baseline 90/100), KL 0.0335, length deviation 0.079 σ, generation health PASSED.
- LoRA stream-merged into base weights in BF16 (
W += B@A, alpha = r = 1). - BF16 GGUF conversion via
convert_hf_to_gguf.py --outtype bf16(llama.cpplfm2). - imatrix calibration: 401 chunks / ≈1.6M tokens (APEX calibration set), computed on the BF16 GGUF.
- Quantized with
llama-quantize --imatrix <imatrix.gguf> <src> <dst> <type>for each tier.