library_name: transformers.js
license: other
license_name: lfm1.0
license_link: LICENSE
pipeline_tag: text-generation
language:
- en
- zh
- ko
- ja
- fr
- de
- es
tags: - onnx
- onnxruntime
- webgpu
- lfm2.5
- uncensored
- text-generation
base_model: - SC117/LFM2.5-2.6B-Uncensored
base_model_relation: finetune
LFM2.5-2.6B-Uncensored ONNX
ONNX Runtime / Transformers.js exports of SC117/LFM2.5-2.6B-Uncensored. These files are intended for browser and cross-platform inference, especially WebGPU.
This repository contains the two web-oriented variants:
| Variant | File data size | Recommended use |
|---|---|---|
| Q4 | 1.85 GB | Smaller download; WebGPU/CPU/GPU |
| Q4F16 | 1.67 GB | Default WebGPU choice; FP16 runtime and KV cache |
Q4F16 is not a GGUF quantization. It uses INT4 weights with FP16 runtime tensors and KV cache. It is therefore not interchangeable with Q4_K_M.
Transformers.js
npm install @huggingface/transformers
import { pipeline, TextStreamer } from "@huggingface/transformers";
const generator = await pipeline(
"text-generation",
"acidsound/LFM2.5-2.6B-Uncensored-ONNX",
{
device: "webgpu",
dtype: "q4f16", // use "q4" for the Q4 variant
},
);
const messages = [
{ role: "system", content: "You write concise English and Chinese image prompts." },
{ role: "user", content: "Create a cinematic prompt for a rainy neon street." },
];
const output = await generator(messages, {
max_new_tokens: 256,
temperature: 0.3,
repetition_penalty: 1.05,
streamer: new TextStreamer(generator.tokenizer, {
skip_prompt: true,
skip_special_tokens: true,
}),
});
console.log(output[0].generated_text.at(-1).content);
The first browser load downloads roughly 1.7–1.9 GB. Use a progress indicator and let the browser cache the model for subsequent sessions. WebGPU support is required for the recommended Q4F16 path.
Files
config.json
generation_config.json
tokenizer.json
tokenizer_config.json
chat_template.jinja
onnx/model_q4.onnx
onnx/model_q4.onnx_data
onnx/model_q4f16.onnx
onnx/model_q4f16.onnx_data
The transformers.js_config section in config.json maps the external-data files and selects FP32 KV cache for Q4 or FP16 KV cache for Q4F16.
Provenance
- Source model:
SC117/LFM2.5-2.6B-Uncensored - Source revision:
578ad81f31061f16797ce9f5451835f59d1ad2d5 - Exporter: Liquid4All/onnx-export
- Exporter revision:
9a23ddd23035165f7414a5de3220a51e85780f64 - Export command: LiquidONNX LFM2 builder, block size 32, symmetric Q4; Q4F16 converted from the Q4 graph with FP16 runtime/cache tensors.
SHA-256
| File | SHA-256 |
|---|---|
onnx/model_q4.onnx |
7AB036E314110DB418B37A36175C5998E9EF441FC88D637523CD6D2A6FF2DE0D |
onnx/model_q4.onnx_data |
F1EBAACBCDFDC164E415E38E47F480B3505F53D7A9D9407332E5C6FDC647C64C |
onnx/model_q4f16.onnx |
0E17A86E3E5D0FE6EF08967843D74BD79B40F4FA916ADF1B4572A5C2DF546D8A |
onnx/model_q4f16.onnx_data |
72E0B2CABD22A13F5553410FCF9D91BA8A6ECAEB99998C71CAF176B4045CB919 |
License
The source model is distributed under the LFM Open License 1.0. See LICENSE and the source model card for the applicable terms.