license: apache-2.0
license_link: LICENSE
pipeline_tag: image-text-to-text
base_model: llmfan46/gemma-4-12B-it-uncensored-heretic
base_model_relation: quantized
tags:
- comfyui
- gemma4_unified
- vision-language
- int8
- convrot
- heretic
- uncensored
Gemma 4 12B Uncensored Heretic INT8 ConvRot for ComfyUI
Native ComfyUI single-file quantization of llmfan46/gemma-4-12B-it-uncensored-heretic, derived from Google Gemma 4 12B IT. Quantization and packaging by radiatingreverberations.
The file is approximately 12.06 GB / 11.23 GiB. Language linears and token embeddings use INT8 ConvRot. Vision components, audio components, projectors, norms, and other remaining tensors preserve their source bytes. The source tokenizer is embedded in the file.
This is a generic Gemma 4 Unified 12B conversion for ComfyUI's native Generate Text path. It contains no LTX-specific conditioning projections. Conversion integrity and native tensor names and shapes have been checked. The publisher confirmed that text generation and multi-image tests passed in ComfyUI. Generation speed and runtime VRAM usage have not been recorded.
Quantization
| Component | Stored format | Count | Parameters |
|---|---|---|---|
| Language block linears | INT8 ConvRot, int8_tensorwise |
328 | 10.900B |
| Token embedding | Per-row INT8 with ConvRot | 1 | 1.007B |
| Vision and audio components, projectors, norms, and remaining tensors | Original dtype and bytes | Remaining tensors | — |
The quantizer uses its default absmax INT8 mode with a ConvRot group size of 256. This release uses INT8 weights, rather than the quantizer's W4A8 mode. There are 1,336 output tensors, including 329 embedded quantization configurations. All 677 original model tensor names are retained after ComfyUI prefix remapping, together with the embedded tokenizer.
Gemma 4 12B Unified projects image patches into its language model rather than using a separate vision tower. The retained multimodal tensors include these vision embeddings and their projectors. Audio tensors are retained as well; this release does not establish audio input support in ComfyUI.
Quantization reconstruction errors
| Component | Mean relative error | Maximum relative error |
|---|---|---|
| INT8 language linears | 0.894% | 1.210% |
| INT8 token embedding | 0.865% | — |
Relative error measures ||dequantized_weight - source_weight|| / ||source_weight||. These measurements describe weight reconstruction and do not measure task accuracy or refusal behavior. Conversion took 225.7 seconds, excluding subsequent validation and hashing.
Full records: per-layer errors, conversion log, and dry-run plan.
Use in ComfyUI
From your ComfyUI directory:
hf download \
radiatingreverberations/Gemma-4-12B-It-Uncensored-Heretic-INT8-ConvRot-ComfyUI \
gemma4_12b_uncensored_heretic_int8_convrot.safetensors \
--local-dir models/text_encoders
Use a ComfyUI build with native Gemma4_12B / GEMMA_4_12B detection, Gemma 4 Unified image preprocessing, INT8 ConvRot linears, and rotated INT8 embedding support. Select the file in a native Gemma 4 loader workflow and connect it to Generate Text. Connect image input for image understanding. Docker users must place the file in the host directory mounted as the container's text encoder directory.
The conversion was checked against native Gemma 4 12B configuration with 48 layers, hidden size 3840, eight local KV heads, one global KV head, and shared global K/V. No Gemma 31B configuration patch is required. The recorded ComfyUI Gemma 4 source checksum is in provenance.json.
Verification status
| Check | Result |
|---|---|
| Source checkpoint SHA-256 and pinned revision | Verified |
Converted names and shapes against native Gemma4_12B |
Verified |
| Original tensors retained and quantization format counts | Verified |
| Unquantized tensor bytes and embedded tokenizer preserved | Verified |
| Published tensor payload matches the local quant | Verified by SHA-256 |
| Native ComfyUI text generation | Passed, confirmed by publisher |
| Native ComfyUI multi-image input | Passed, confirmed by publisher |
| Individual one-image, two-image, and four-image results | Not separately recorded |
| Generation speed and runtime VRAM | Not yet recorded |
The publisher's runtime confirmation is a smoke test. Image counts, resolutions, test prompts, outputs, sampling settings, and raw generation logs were not supplied for this package; it does not establish a vision-accuracy benchmark.
The conversion ran on an NVIDIA RTX 4000 Ada Generation 20 GB using Python 3.12.13, PyTorch 2.11.0+cu130, comfy-kitchen 0.2.36, safetensors 0.8.0, and huggingface_hub 1.33.0. Conversion hardware does not establish an inference memory or throughput benchmark.
Conversion provenance
Source revision: cf3bb4a17e8cf1cdf6b429b6ed9f9dbe476c40e0.
Quantizer: Comfy-Org/comfy-model-tools, commit d6797787e6bdb1a1fb0094d588a26f8e71a1c757.
Before quantization, the source was streamed into a single file with these prefix mappings and the original tokenizer.json embedded as a U8 tensor named tokenizer_json:
| Hugging Face prefix | ComfyUI prefix |
|---|---|
model.language_model. |
model. |
model.vision_embedder. |
vision_model. |
model.embed_vision. |
multi_modal_projector. |
model.embed_audio. |
audio_projector. |
The resulting file was quantized using:
python -u quant_int8_convrot.py \
gemma4_12b_heretic_bf16.safetensors \
gemma4_12b_uncensored_heretic_int8_convrot.safetensors \
--verify-report quantization_errors.tsv
Publication adds attribution and modification notices to the Safetensors header. Tensor names, shapes, formats, and payload bytes match the local quantizer output. provenance.json records the source, dependency versions, original output checksum, exported checksum, and tensor payload checksum.
License and attribution
The upstream Heretic model and Google Gemma 4 license declare Apache-2.0. This release includes LICENSE, attribution and modification notices in NOTICE, and the upstream model card.
The uncensoring was performed by llmfan46 using Heretic; this release quantizes that checkpoint. Upstream refusal and capability results have not been rerun on this quant.
Download SHA256SUMS alongside the model to verify it:
sha256sum --check SHA256SUMS