base_model: TrevorJS/gemma-4-E4B-it-uncensored
pipeline_tag: image-text-to-text
library_name: openvino
language:
- en
- zh
license: apache-2.0
tags: - gemma-4
- openvino
- int4
- uncensored
- multimodal
- vision-language-model
- efficient
Gemma-4 E4B (8B) INT4 OpenVINO
Uncensored Gemma-4 E4B exported to OpenVINO IR with INT4 weight compression via NNCF.
The E4B variant sits between the E2B and 12B — offering strong reasoning capability with moderate VRAM requirements.
Model Details
| Property | Value |
|---|---|
| Base model | TrevorJS/gemma-4-E4B-it-uncensored |
| Architecture | Gemma4ForConditionalGeneration |
| Precision | INT4 (asymmetric, group_size=128) |
| Parameters | 8B |
| Hidden size | 2560 |
| Layers | 42 |
| Attention heads | 8 |
| KV heads | 2 |
| KV shared layers | 18 |
| Max position | 131072 |
| Sliding window | 512 |
| Vocab size | 262144 |
| Vision | ✅ 280 soft tokens |
| Audio | ✅ |
| Disk size | 6.4 GB |
| Framework | OpenVINO 2026.4+ |
Usage
With openvino_genai (recommended)
import openvino_genai as g
pipe = g.VLMPipeline("gemma-4-E4B-int4-ov", "GPU")
result = pipe.generate("Hello", max_new_tokens=100)
print(result.texts[0])
Multimodal example
result = pipe.generate(
"Describe this image",
images=["image.jpg"],
max_new_tokens=256,
)
print(result.texts[0])
OpenAI-compatible API server
OV_MODEL=./models_trevorjs_e4b_int4_ov_new OV_PORT=8092 python serving/ov_server.py
Performance (Arc A770 16GB, KV_CACHE=f16)
| Context | Latency | Throughput |
|---|---|---|
| Short | 4.0ms | ~250 tok/s |
| 1K | 30.2ms | ~33 tok/s |
| 4K | 35.3ms | ~28 tok/s |
| 8K | 42.3ms | ~24 tok/s |
| 12K | 51.7ms | ~19 tok/s |
| 16K | 59.6ms | ~17 tok/s |
E4B offers a balanced tradeoff: more capable than E2B while fitting comfortably in 16GB VRAM for up to 16K context.
Export Process
Same as other Gemma-4 OpenVINO exports:
- Export FP16 via
optimum-cli export openvino --task image-text-to-text --weight-format fp16 - Compress to INT4 with
nncf.compress_weights(INT4_ASYM, group_size=128, ratio=1.0)
License
Apache 2.0