← back to catalog · registered 2026-08-22 13:56

ONTONOUS/gemma-4-E2B-uncensored-int4-ov

ONTONOUS Gemma multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/ONTONOUS%2Fgemma-4-E2B-uncensored-int4-ov"
Response includes
  • classification m-uncensored
  • files 19
  • hub_downloads_all_time 60
  • author_summary 3 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
60
16 last 30d - stable
Likes
0
Model age
2mo ago
created 2026-07-14
Downloads over time
Now66→from27↑144%
2540557027 on Jul 1566 on Oct 1166 on Oct 10JulAugSepOct
Jul 15 → Oct 11 · 53 snapshots · spans 88 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh
Tags
openvino gemma4 gemma-4 int4 uncensored multimodal vision-language-model efficient image-text-to-text conversational en zh

Related

Total size
3.95 GB
Files
19
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-07-14 11:35

Files by quantization

Auxiliary files 19 files 3.99 GB
openvino_text_embeddings_per_layer_model.bin 2.19 GB 167aea14 download
openvino_language_model.bin 1.28 GB d63fe1ee download
openvino_text_embeddings_model.bin 385 MB 2016d9b3 download
openvino_vision_embeddings_model.bin 84.1 MB 728ede5e download
openvino_tokenizer.bin 16.5 MB d151f48d download
openvino_detokenizer.bin 4.21 MB 461919bd download
tokenizer.json 30.7 MB a2619fe1 download
openvino_language_model.xml 4.27 MB 951022b2 download
openvino_vision_embeddings_model.xml 2.29 MB 500db17e download
openvino_tokenizer.xml 37.1 KB 3ad35c44 download
openvino_detokenizer.xml 22.1 KB 30b41d0a download
openvino_text_embeddings_per_layer_model.xml 19.9 KB 0b9739cc download
chat_template.jinja 11.6 KB afb1d517 download
openvino_text_embeddings_model.xml 8.80 KB 84f7d17e download
config.json 4.86 KB 6004c163 download
tokenizer_config.json 2.70 KB 8abee96e download
README.md 2.52 KB 650c003b download
.gitattributes 1.53 KB 52373fe2 download
generation_config.json 203 B edda3c19 download

README current version from Hugging Face


base_model: TrevorJS/gemma-4-E2B-it-uncensored
pipeline_tag: image-text-to-text
library_name: openvino
language:

  • en
  • zh
    license: apache-2.0
    tags:
  • gemma-4
  • openvino
  • int4
  • uncensored
  • multimodal
  • vision-language-model
  • efficient

Gemma-4 E2B (5.1B) INT4 OpenVINO

Uncensored Gemma-4 E2B exported to OpenVINO IR with INT4 weight compression via NNCF.

The E2B variant is Gemma-4's efficient architecture with shared KV layers and smaller hidden dimensions, optimized for long-context scenarios.

Model Details

Property Value
Base model TrevorJS/gemma-4-E2B-it-uncensored
Architecture Gemma4ForConditionalGeneration
Precision INT4 (asymmetric, group_size=128)
Parameters 5.1B
Hidden size 1536
Layers 35
Attention heads 8
KV heads 1
KV shared layers 20
Max position 131072
Sliding window 512
Vocab size 262144
Vision ✅ 280 soft tokens
Audio ✅
Disk size 4.3 GB
Framework OpenVINO 2026.4+

Usage

With openvino_genai (recommended)

import openvino_genai as g

pipe = g.VLMPipeline("gemma-4-E2B-int4-ov", "GPU")
result = pipe.generate("Hello", max_new_tokens=100)
print(result.texts[0])

Multimodal example

result = pipe.generate(
    "Describe this image",
    images=["image.jpg"],
    max_new_tokens=256,
)
print(result.texts[0])

OpenAI-compatible API server

OV_MODEL=./models_trevorjs_e2b_int4_ov_new OV_PORT=8092 python serving/ov_server.py

Performance (Arc A770 16GB, KV_CACHE=f16)

Context Latency Throughput
Short 3.5ms ~285 tok/s
1K 24.7ms ~40 tok/s
4K 28.0ms ~36 tok/s
8K 35.8ms ~28 tok/s
12K 15.7ms ~64 tok/s
14K 46.8ms ~21 tok/s

Best value: E2B achieves the best price/performance ratio. With only 4.3GB VRAM, it supports up to 14K context on a 16GB GPU.

Why E2B?

The E2B variant uses 20/35 shared KV layers and 1 KV head to dramatically reduce KV cache memory. This allows longer context windows on limited VRAM compared to the 12B model.

Export Process

Same as other Gemma-4 OpenVINO exports:

  1. Export FP16 via optimum-cli export openvino --task image-text-to-text --weight-format fp16
  2. Compress to INT4 with nncf.compress_weights (INT4_ASYM, group_size=128, ratio=1.0)

License

Apache 2.0

Links

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-07-14Upload folder using huggingface_hubc2457252.5 KB
    Loading...
  2. 2026-07-14Upload folder using huggingface_hubcf48ee52.6 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration