library_name: gguf
license: other
license_name: lfm1.0
license_link: LICENSE
pipeline_tag: image-text-to-text
tags:
- liquid
- lfm2.5
- edge
- uncensored
- abliterix
- vision
- multimodal
- quantization
- gguf
- imatrix
base_model: - LiquidAI/LFM2.5-VL-3B
- SC117/LFM2.5-2.6B-Uncensored
base_model_relation: quantized
LFM2.5-VL-3B-Uncensored-GGUF
English | 📖 中文文档
Uncensored vision-language edge model · abliterix Trial 65 merged into official LFM2.5-VL-3B · imatrix-calibrated GGUFs + mmproj
LFM2.5-VL-3B is Liquid AI's official 3.1B vision-language edge model: the LFM2.5-2.6B hybrid language backbone (30 layers: 22 double-gated short-convolution blocks + 8 GQA, 128K context, 128K vocab) paired with a SigLIP2 NaFlex vision encoder (256×256, patch 16) and a language-aligned projector.
This release merges the abliterix Trial 65 LoRA (from our LFM2.5-2.6B-Uncensored, rank-1, alpha = r = 1) into the language-model layers of the official VL checkpoint (q/k/v/out_proj, feed_forward.w2, conv.out_proj — 84 tensors across 30 layers). The vision encoder and projector are untouched, so multimodal capability is fully preserved.
Build steps:
- LoRA merge into the VL checkpoint:
W += (B @ A) * (alpha / r), applied tomodel.language_model.layers.*. - BF16 GGUF conversion of the merged model (llama.cpp
lfm2architecture) + mmproj export of the vision tower viaconvert_hf_to_gguf.py --mmproj --outtype bf16. - imatrix calibration — reused the 401-chunk (≈1.6M tokens) importance matrix from the 2.6B release (same architecture, weights closely related).
- Quantization with
llama-quantize --imatrixinto five tiers.
License: LFM Open License v1.0 (same as the base model).
After merging the Trial 65 steering, this model shows a much lower refusal rate on both text and vision-grounded prompts, and can differ substantially from official LFM2.5-VL-3B. Evaluate compliance and safety for your use case; control access and audit as needed.
| Refusals (2.6B harmful eval) | 6 / 100 (baseline ~90 / 100) — same Trial 65 LoRA, measured on the 2.6B release |
| Spot check (this VL release) | NSFW fiction / privacy intrusion / crime detail prompts → all answered without refusal |
| Vision capability | Fully preserved — shape/color/OCR description verified after merge |
| Selected trial | abliterix Trial 65 (rank-1 LoRA, alpha = r = 1) |
| Thinking | Not available — official VL is trained to answer directly (no <think> mode) |
Main model (language backbone, lfm2 architecture, 128K context) — pick one tier and pair it with the mmproj:
| File | Size | BPW | Best for |
|---|---|---|---|
*-IQ3_XS.gguf | 1.22 GB | ~3.30 | Maximum compression (perceptible quality loss on small models) |
*-IQ4_XS.gguf | 1.52 GB | ~4.25 | Sweet spot — smallest tier with Q4_K_M-class quality |
*-Q4_K_M.gguf | 1.67 GB | ~4.94 | Verified everyday default |
*-Q6_K.gguf | 2.22 GB | ~6.56 | Quality-first local use |
*-Q8_0.gguf | 2.87 GB | ~8.50 | Near-lossless (imatrix optional here) |
*-BF16.gguf | 5.40 GB | 16.00 | Lossless baseline (source of all tiers) |
Vision tower (required, pick one):
| File | Size | Notes |
|---|---|---|
mmproj-*-BF16.gguf | 0.86 GB | Full-precision vision tower, recommended default |
mmproj-*-Q8_0.gguf | 0.58 GB | 8-bit vision tower for low-memory devices (negligible quality difference, same as official Q8_0 mmproj) |
All main-model tiers are lfm2 architecture, 128K context, single-file GGUFs, imatrix-calibrated.
The importance matrix (401 chunks / ≈1.6M tokens of mixed conversation, math, and code data, computed on the 2.6B-Uncensored BF16 GGUF) tells the quantizer which weights are sensitive. Since the VL language backbone shares the exact same architecture and near-identical weights, the matrix transfers cleanly (166/166 quantized tensors matched; only token_embd falls back to plain q6_K). K-quants and especially the IQ tiers use it to keep more bits on attention/embedding paths.
Pair any main-model tier with the mmproj (recent llama.cpp, e.g. b10299+, required for lfm2 VL support):
llama-server -m LFM2.5-VL-3B-Uncensored-Q4_K_M.gguf \
--mmproj mmproj-LFM2.5-VL-3B-BF16.gguf \
--ctx-size 8192 --flash-attn on --host 0.0.0.0 --port 8080
CLI with an image:
llama-llava-cli -m LFM2.5-VL-3B-Uncensored-Q4_K_M.gguf \
--mmproj mmproj-LFM2.5-VL-3B-BF16.gguf \
-i image.jpg -p "Describe this image." -ngl 99
Transformers / vLLM / SGLang users: use the merged BF16 safetensors in the parent repo (to be published).
- No thinking mode. Unlike the 2.6B text model, official LFM2.5-VL-3B is trained to answer directly for low-latency edge tasks; the
<think>tags are not generated. Keepmax_new_tokensmodest. - Vision strengths (per official model card): near-real-time object detection, OCR with layout annotation, document/chart understanding, on-device translation.
- Chat template: ChatML-like with
<image>placeholder, same as official. Recommended sampling:temperature 0.2, top_k 50, repetition_penalty 1.0.
Values below are taken from the official LiquidAI model card for LFM2.5-VL-3B (not re-measured on this release):
| MME | 73.1 | ChartQA | 81.3 |
| MMStar | 63.3 | OCRBenchv2 (EN) | 47.5 |
| RealWorldQA | 73.1 | MathVista | 68.5 |
| CountBenchQA | 87.3 | POPE | 88.7 |