← back to catalog · registered 2026-08-22 13:56

ONTONOUS/gemma-4-E4B-uncensored-int4-ov

ONTONOUS Gemma multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/ONTONOUS%2Fgemma-4-E4B-uncensored-int4-ov"
Response includes
  • classification m-uncensored
  • files 19
  • hub_downloads_all_time 81
  • author_summary 3 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
81
16 last 30d - stable
Likes
0
Model age
2mo ago
created 2026-07-14
Downloads over time
Now89→from8↑1,013%
43566978 on Jul 1589 on Oct 1189 on Oct 10JulAugSepOct
Jul 15 → Oct 11 · 53 snapshots · spans 88 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh
Tags
openvino gemma4 gemma-4 int4 uncensored multimodal vision-language-model efficient image-text-to-text conversational en zh

Related

Total size
5.90 GB
Files
19
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-07-14 15:29

Files by quantization

Auxiliary files 19 files 5.94 GB
openvino_text_embeddings_per_layer_model.bin 2.63 GB 69f388d3 download
openvino_language_model.bin 2.55 GB 4f29f687 download
openvino_text_embeddings_model.bin 641 MB 1fcf6279 download
openvino_vision_embeddings_model.bin 84.9 MB 3dc6643d download
openvino_tokenizer.bin 16.5 MB d151f48d download
openvino_detokenizer.bin 4.21 MB 461919bd download
tokenizer.json 30.7 MB a2619fe1 download
openvino_language_model.xml 5.42 MB 25c8307b download
openvino_vision_embeddings_model.xml 2.29 MB 2bb84331 download
openvino_tokenizer.xml 37.1 KB cdbcaf7f download
openvino_detokenizer.xml 22.1 KB e8d67208 download
openvino_text_embeddings_per_layer_model.xml 19.9 KB 2e25b766 download
chat_template.jinja 11.6 KB afb1d517 download
openvino_text_embeddings_model.xml 8.80 KB e36eea1c download
config.json 5.05 KB 47d7ff7f download
tokenizer_config.json 2.70 KB 8abee96e download
README.md 2.29 KB 850da9a7 download
.gitattributes 1.53 KB 52373fe2 download
generation_config.json 203 B edda3c19 download

README current version from Hugging Face


base_model: TrevorJS/gemma-4-E4B-it-uncensored
pipeline_tag: image-text-to-text
library_name: openvino
language:

  • en
  • zh
    license: apache-2.0
    tags:
  • gemma-4
  • openvino
  • int4
  • uncensored
  • multimodal
  • vision-language-model
  • efficient

Gemma-4 E4B (8B) INT4 OpenVINO

Uncensored Gemma-4 E4B exported to OpenVINO IR with INT4 weight compression via NNCF.

The E4B variant sits between the E2B and 12B — offering strong reasoning capability with moderate VRAM requirements.

Model Details

Property Value
Base model TrevorJS/gemma-4-E4B-it-uncensored
Architecture Gemma4ForConditionalGeneration
Precision INT4 (asymmetric, group_size=128)
Parameters 8B
Hidden size 2560
Layers 42
Attention heads 8
KV heads 2
KV shared layers 18
Max position 131072
Sliding window 512
Vocab size 262144
Vision ✅ 280 soft tokens
Audio ✅
Disk size 6.4 GB
Framework OpenVINO 2026.4+

Usage

With openvino_genai (recommended)

import openvino_genai as g

pipe = g.VLMPipeline("gemma-4-E4B-int4-ov", "GPU")
result = pipe.generate("Hello", max_new_tokens=100)
print(result.texts[0])

Multimodal example

result = pipe.generate(
    "Describe this image",
    images=["image.jpg"],
    max_new_tokens=256,
)
print(result.texts[0])

OpenAI-compatible API server

OV_MODEL=./models_trevorjs_e4b_int4_ov_new OV_PORT=8092 python serving/ov_server.py

Performance (Arc A770 16GB, KV_CACHE=f16)

Context Latency Throughput
Short 4.0ms ~250 tok/s
1K 30.2ms ~33 tok/s
4K 35.3ms ~28 tok/s
8K 42.3ms ~24 tok/s
12K 51.7ms ~19 tok/s
16K 59.6ms ~17 tok/s

E4B offers a balanced tradeoff: more capable than E2B while fitting comfortably in 16GB VRAM for up to 16K context.

Export Process

Same as other Gemma-4 OpenVINO exports:

  1. Export FP16 via optimum-cli export openvino --task image-text-to-text --weight-format fp16
  2. Compress to INT4 with nncf.compress_weights (INT4_ASYM, group_size=128, ratio=1.0)

License

Apache 2.0

Links

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-07-14Upload folder using huggingface_hub65c7bd22.3 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration