license: other
license_name: qwen-research
license_link: LICENSE
base_model: Qwen/Qwen-Image-2.1-Turbo
base_model_relation: quantized
library_name: gguf
pipeline_tag: text-to-image
tags:
- sakura
- sakura-mini
- gguf
- stable-diffusion.cpp
- text-to-image
- qwen-image
- uncensored
- abliterated
- heretic
- not-for-all-audiences
Sakura — Qwen-Image-2.1-Turbo, uncensored (GGUF)
Built with Qwen. Community quantization of Qwen/Qwen-Image-2.1-Turbo (8-step distilled, 7B DiT) with an abliterated text encoder. Independent work, not made or endorsed by Alibaba/Qwen. Non-commercial use only (Qwen RESEARCH LICENSE AGREEMENT, copy in LICENSE).
For adults only; you are responsible for what you generate and for the laws that apply to you.
What is "uncensored" here, exactly
| Part | What we did |
|---|---|
| Text encoder (Qwen3-VL-8B) | Directional ablation with Heretic (o_proj + down_proj), then GGUF Q4_K_M. Measured as a chat model: refusals on Heretic's 100 held-out harmful prompts 100/100 -> 8/100, KL divergence of the first-token distribution 0.047. |
| Image model (7B DiT) | Only quantized (Q4_K, Q5_K) from the official Turbo weights. We did not change what the DiT can draw. |
What we did not measure: how much the ablated encoder changes the images. The refusal numbers describe the encoder used as a language model; in this pipeline it only supplies embeddings. We only compared neutral prompts (see below). No sample images of restricted content are included.
Files
| File | Size | What it is |
|---|---|---|
Sakura-Image-2.1-Turbo-Q5_K_M-4.60GiB.gguf |
4.60 GiB | image model (DiT), uniform Q5_K recipe of sd.cpp (file name carries the standard label Q5_K_M so the Hub recognises it) |
Sakura-Image-2.1-Turbo-Q4_K_M-3.77GiB.gguf |
3.77 GiB | image model (DiT), uniform Q4_K recipe of sd.cpp (file name carries the standard label Q4_K_M) |
Sakura-TextEncoder-Qwen3VL-8B-Uncensored-Q4_K_M-4.68GiB.gguf |
4.68 GiB | our ablated text encoder, Q4_K_M |
Not included (use the official files): the VAE qwen_image_2.1_vae_bf16.safetensors from Comfy-Org/Qwen-Image-2.1.
Run with stable-diffusion.cpp
sd-cli \
--diffusion-model Sakura-Image-2.1-Turbo-Q5_K_M-4.60GiB.gguf \
--vae qwen_image_2.1_vae_bf16.safetensors \
--llm Sakura-TextEncoder-Qwen3VL-8B-Uncensored-Q4_K_M-4.68GiB.gguf \
--sigmas "1.0,0.978453,0.95418,0.926626,0.89508,0.845148,0.704534,0.414568,0.0" \
--steps 8 --cfg-scale 1.0 --sampling-method euler --diffusion-fa --vae-tiling \
-W 1024 -H 1024 -p "your prompt" -o out.png
We ran and measured everything at 512x512 (8 steps, about 12 to 15 s per image on a Radeon 8060S with Vulkan); 1024x1024 is the usual size of the model but we did not test it. The sigma schedule is the official Turbo schedule (sample_sigmas of the model card). Use CFG 1.0 and 8 steps. sd-cli/sd-server of stable-diffusion.cpp master (commit 228c707 was used) load these files directly.
Image-model quality (what we measured)
Same prompts, same seed (42), 512x512, 8 steps, same machine; reference is the Q8_0 file the two files were made from (DogukanUrker's). Eight neutral prompts (cat, street with neon signs, portrait, mountain lake, poster with text, kitchen, marble statue, beach). Higher is closer to Q8_0; a diffusion model changes details even with tiny weight changes, so these are deviation numbers, not a quality score.
| File | Size | PSNR vs Q8_0 | SSIM vs Q8_0 |
|---|---|---|---|
| Q5_K | 4.60 GiB | 24.77 dB | 0.906 |
| Q4_K | 3.77 GiB | 21.46 dB | 0.842 |
Text encoder: refusals and effect on images
Our own refusal measurement on the encoder used as a chat model (llama.cpp, Q4_K_M, temperature 0, 120 tokens, 104 held-out harmful prompts of the Heretic harmful_behaviors test split, identical prompts and settings for both rows):
| Text encoder (Q4_K_M) | strict refusals | broad refusal markers |
|---|---|---|
| official, unchanged | 58 / 104 | 98 / 104 |
| ours (Heretic) | 0 / 104 | 9 / 104 |
Heretic's own count on its 100 evaluation prompts: 100/100 refusals before, 8/100 after, KL divergence 0.047 (trial 66 of 70, seed 42, 24 random start trials; checkpoint selection rule: fewest refusals with KL <= 0.1).
Effect on images (Q8_0 image model, eight neutral prompts, seed 42; reference = the same Q8_0 with the official int8 text encoder):
| Text encoder | PSNR vs reference | SSIM vs reference |
|---|---|---|
| official encoder as GGUF Q4_K_M | 20.75 dB | 0.830 |
| ours (Heretic) Q4_K_M | 18.70 dB | 0.771 |
The ablated encoder moves the images further away from the reference than plain quantization does (compositions and details change), but the images follow the prompts (we looked at the cat, street and poster images; text such as "LISBOA" is still rendered correctly). These numbers do not say anything about content restrictions.
What we changed (Qwen RESEARCH LICENSE, section 3b)
Sakura-Image-2.1-Turbo-*.gguf: made withsd-cli -M convertfrom the Q8_0 GGUF of DogukanUrker/Qwen-Image-2.1-Turbo-GGUF (sha256228d7931adade6cd..., itself a conversion of the official BF16 weights), not directly from the BF16 weights. They are plain uniform Q5_K / Q4_K files like other sd.cpp conversions; the new part of this repository is the text encoder. Modified files.Sakura-TextEncoder-*.gguf: the official text encoder of Qwen-Image-2.1-Turbo with Heretic's directional ablation (o_proj,down_proj) merged into the weights, converted to GGUF with llama.cpp and quantized to Q4_K_M. Modified file.- Everything else (VAE, tokenizer, scheduler) is not part of this repository.
License and use
Qwen RESEARCH LICENSE AGREEMENT (LICENSE): use, reproduction and modification for non-commercial purposes only (research or evaluation); redistribution requires this license text and the notices in NOTICE. The agreement is governed by the laws of China. "Qwen" is the name of the original work; this is a derivative work by Sakura, "based on Qwen-Image-2.1-Turbo".
Part of the Sakura Mini line.
Credits
- Original model: Qwen/Qwen-Image-2.1-Turbo (Qwen / Alibaba).
- Ablation: p-e-w/heretic (AGPL-3.0, used as a tool; not included here).
- GGUF conversion: leejet/stable-diffusion.cpp (MIT) and ggml-org/llama.cpp (MIT).