license: other
license_name: qwen-research
base_model:
- Qwen/Qwen-Image-2.1
- pottokao/Qwen-Image-2.1-Text-Encoder-Heretic
base_model_relation: quantized
pipeline_tag: text-to-image
library_name: openvino
tags: - openvino
- qwen-image
- uncensored
- abliterated
- heretic
- int4
- text-to-image
- intel-arc
Qwen-Image-2.1 Uncensored — OpenVINO INT4
OpenVINO IR build of Qwen-Image-2.1 with an abliterated text encoder, quantized to INT4. Runs text-to-image on Intel Arc iGPUs.
Uncensoring lives in the prompt-embedding path, so only the Qwen3-VL text encoder needed replacing — the transformer, VAE and vision encoder are the stock model.
What's inside
| Component | Precision | Size | Source |
|---|---|---|---|
text_encoder |
INT4 asym | 3.9 GB | abliterated Qwen3-VL |
text_encoder_i2i |
INT4 asym | 3.9 GB | same (editing slot) |
transformer |
INT4 asym | 3.5 GB | stock Qwen-Image-2.1 |
vision_encoder |
INT8 | 553 MB | stock |
vae_decoder |
INT8 | 243 MB | stock |
vae_encoder |
INT8 | 76 MB | stock |
Total ~13 GB.
Abliteration verified
Embeddings for "a portrait of a woman, photorealistic" differ decisively from the stock encoder — this is not a renamed copy:
stock mean 0.59543
heretic mean 0.22803
max abs diff : 1550.99353
mean abs diff: 10.12831
identical : False
Measured performance
Intel Core Ultra 7 258V, Arc 130V/140V iGPU, 8 cores, 30 GB RAM:
| Test | Result |
|---|---|
| pipeline load | 43.5 s |
| 512x512, 10 steps | 12.9 s (1.29 s/step) |
| peak RSS | 7.1 GB |
Usage
import torch
from optimum.intel import OVDiffusionPipeline
pipe = OVDiffusionPipeline.from_pretrained(
"blaj/Qwen-Image-2.1-Uncensored-OpenVINO-INT4",
compile=True, device="GPU",
)
img = pipe(
"your prompt",
num_inference_steps=20, height=1024, width=1024,
generator=torch.Generator().manual_seed(42),
).images[0]
img.save("out.png")
Requires the nightly stack: diffusers and optimum-intel from main, plus OpenVINO 2026.5.0 nightly. See the full conversion write-up:
https://devnotes.page/how-to-convert-qwen-image-2-1-to-openvino-ir-and-build-an-uncensored-version
How it was built
- Exported the abliterated text encoder to fp16 IR with
optimum(57 s, 14.4 GB), taking only.model.language_model. - Compressed that IR to INT4 with NNCF —
group_size=-1is required,128and64raiseInvalidGroupSizeError(37 s, → 3.9 GB). - Copied a pre-converted stock INT4 pipeline and replaced
text_encoder/andtext_encoder_i2i/.
Python 3.12 is required for step 1. On Python 3.14 functools.partial is a descriptor, so NORMALIZED_CONFIG_CLASS resolves to a bound method and the export dies with:
TypeError: NormalizedConfig.__init__() got multiple values for argument 'allow_new'
A full-pipeline export of the 33 GB source is OOM-killed (exit 137) on a 30 GB machine — export components separately.
Limitations
- Image editing (i2i) loads but exhausts 30 GB during generation; it needs the vision encoder and both text encoders resident at once. Not resolution-bound — 256x256 with 8 steps still reached 28 GB. Text-to-image is unaffected.
- The abliterated text encoder is a community artifact and carries its own behaviour changes beyond removing refusals.
- Inherits the Qwen-Image-2.1 research licence. Uncensored output is your responsibility.