license: apache-2.0
library_name: litert-lm
tags:
- litert
- litert-lm
- on-device
- qwen2.5
- uncensored
- android
- mediapipe
base_model: - thirdeyeai/Qwen2.5-1.5B-Instruct-uncensored
- Qwen/Qwen2.5-Coder-1.5B-Instruct
pipeline_tag: text-generation
Qwen2.5-1.5B-Instruct-uncensored (LiteRT-LM)
On-device LiteRT-LM (.litertlm) conversion ofthirdeyeai/Qwen2.5-1.5B-Instruct-uncensored
for Google AI Edge LiteRT-LM / MediaPipe on-device runtimes (Android & desktop).
Files
| File | Size | Notes |
|---|---|---|
Qwen2.5-1.5B-Instruct-uncensored_dynamic_wi8_afp32_ekv1280.litertlm |
1802859440 bytes (~1.68 GiB) | int8 dynamic, KV cache 1280, prefill 128 |
Source
- Fine-tune:
thirdeyeai/Qwen2.5-1.5B-Instruct-uncensored(BF16 safetensors, Qwen2) - Upstream base:
Qwen/Qwen2.5-Coder-1.5B-Instruct(Apache-2.0)
Conversion method
litert-torch export_hf \
/path/to/thirdeyeai/Qwen2.5-1.5B-Instruct-uncensored \
./out \
--quantization_recipe=dynamic_wi8_afp32 \
--cache_length=1280 \
--prefill_lengths=128 \
--externalize_embedder \
--bundle_litert_lm
- Tool:
litert-torch0.9.4 Hugging Face Export (export_hf) - Quantization:
dynamic_wi8_afp32(per-channel INT8 weights, FP32 activations) - KV cache: 1280
- Prefill signature: 128
- Embedder externalized (Qwen embedding / lm_head packaging)
- Runtime target: LiteRT-LM ≥ 0.17
Intended use
On-device chat (Android LiteRT-LM). Community conversion — not affiliated with Alibaba Cloud or the Qwen team. The source fine-tune is labelled uncensored; use at your own discretion.
License
Apache-2.0 (inherited from Qwen2.5). This repo only redistributes converted weights + packaging.