license: apache-2.0
license_link: LICENSE
pipeline_tag: image-text-to-text
base_model: llmfan46/Qwen3.5-27B-ultra-uncensored-heretic-v1
base_model_relation: quantized
tags:
- comfyui
- diffusion-single-file
- qwen3_5
- vision-language
- w4a8
- int8
- convrot
- heretic
- uncensored
Qwen3.5 27B Ultra Uncensored Heretic W4A8 ConvRot for ComfyUI
Native ComfyUI single-file quantization of llmfan46/Qwen3.5-27B-ultra-uncensored-heretic-v1, derived from Qwen/Qwen3.5-27B. Quantization and packaging by radiatingreverberations.
This release converts the Heretic checkpoint to ComfyUI's embedded quantization format while retaining its language and vision tensors. The language layers use W4A8 ConvRot, the vision block linears use INT8 ConvRot, and the LM head remains BF16. Text generation was reported working with ComfyUI 0.38.0 on an NVIDIA RTX 4000 Ada 20 GB.
The file is approximately 18.08 GB / 16.84 GiB. The reported 20 GB GPU result applies to the benchmark configuration below. The publisher also confirmed successful native Generate Text tests with one image and four images.
Quantization
| Component | Stored format | Layers or tables | Parameters |
|---|---|---|---|
| Language block linears | W4A8 ConvRot, asym_w4a8_int8 |
400 | 24.327B |
| Vision block linears | INT8 ConvRot, int8_tensorwise |
108 | 0.411B |
| Token embedding | Per-row INT8 with ConvRot | 1 | 1.271B |
| LM head, vision merger, position embeddings, norms, biases, small gate projections | BF16 passthrough | — | Remaining tensors |
W4A8 stores INT4 weights with a codebook, FP8 group scales, and FP32 channel scales; activations use INT8 at runtime. Its scale group size is 16 and its ConvRot group size is 256. Vision linears fall back to INT8 with ConvRot group sizes of 64 or 16. The embedding table uses per-row scales and a 256-wide rotation.
The output contains 3,002 tensors, including 509 embedded quantization configurations. All 1,184 source tensor names are retained. The checkpoint has no MTP tensors.
Quantization reconstruction errors
| Component | Mean relative error | Maximum relative error |
|---|---|---|
| W4A8 language layers, gs256 | 7.315% | 7.322% |
| INT8 vision layers, gs64 | 0.819% | 1.032% |
| INT8 vision layers, gs16 | 0.900% | 0.964% |
| INT8 token embedding | 0.880% | — |
Conversion completed without high-error warnings. These errors measure reconstructed weights; task accuracy needs separate evaluation. The W4A8 cosine values in the quantizer's report are placeholders set to 1.0; use its computed relative-error column.
Full records: per-layer report, conversion log, and dry-run plan.
Reported text generation benchmark
The following figures were supplied by the publisher. Raw generation logs and exact sampling settings were not retained in this package.
Hardware: NVIDIA RTX 4000 Ada 20 GB
ComfyUI: 0.38.0
Allocator: cudaMallocAsync
~15.2 tokens/s
~16.9 GiB additional device VRAM
Peak device use: 18.1 / 20.0 GiB
Peak active PyTorch allocation: ~1.0 GiB
Benchmark:
~2k-token prompt
2,048-token reply limit
ComfyUI model cache cleared before test
A ~4,096-token prompt did not fit in 20 GB VRAM.
The 4,096-token failure is an observation from this setup, rather than a universal context limit. Device memory usage and active PyTorch allocation are different measurements. Prompt length, reply limit, image resolution and count, other loaded models, and software versions can change memory requirements.
Use in ComfyUI
From your ComfyUI root directory, download the file into models/text_encoders:
hf download \
radiatingreverberations/Qwen3.5-27B-Ultra-Uncensored-Heretic-W4A8-ConvRot-ComfyUI \
qwen3.5_27b_ultra_uncensored_heretic_w4a8_convrot.safetensors \
--local-dir models/text_encoders
Load the file with CLIPLoader and connect its CLIP output to native Generate Text. For image input, connect the IMAGE input as well. A Docker installation needs the file in the host directory mounted as the container's text encoder directory.
Use a ComfyUI build with Qwen3.5-27B detection and support for asym_w4a8_int8, INT8 ConvRot linears, and rotated INT8 embeddings. This file uses ComfyUI's quantization metadata and should be loaded through ComfyUI.
Verification status
| Check | Result |
|---|---|
| Tensor spans, source tensor retention, format counts, BF16 LM head | Verified from the artifact |
| Export tensor data matches the benchmarked quant | Verified byte for byte by SHA-256 |
| Native text generation in ComfyUI 0.38.0 | Publisher-reported result above |
| Generate Text with one image | Passed, confirmed by publisher |
| Generate Text with four images | Passed, confirmed by publisher |
The image checks establish that these native image-input paths ran successfully in the publisher's setup. Test prompts, outputs, image dimensions, and image-test memory readings were not captured in this package. These checks are smoke tests rather than a vision-accuracy benchmark.
Conversion provenance
Source revision: 0b134ea2c3e0fef5f4f78194d1c86258c1c81b8f.
Quantizer: Comfy-Org/comfy-model-tools, commit d6797787e6bdb1a1fb0094d588a26f8e71a1c757.
python merge_safetensors.py \
qwen35-heretic qwen35_27b_heretic_bf16.safetensors
python -u quant_int8_convrot.py \
qwen35_27b_heretic_bf16.safetensors \
qwen3.5_27b_heretic_w4a8_convrot.safetensors \
--w4a8 \
--exclude '^mtp\.' \
--verify-report qwen35_27b_w4a8_errors.txt
Conversion took 392.5 seconds. Exact conversion-time dependency versions were not captured. The published file adds attribution and modification notices to the Safetensors header; its tensor data matches the quantizer output. provenance.json records source shard identifiers and checksums for the original quant and published export.
Integrity and license
After downloading the model and SHA256SUMS, verify them in the same directory:
sha256sum --check SHA256SUMS
The upstream models declare Apache-2.0. This package includes LICENSE, preserving Qwen's Copyright 2026 Alibaba Cloud notice, and NOTICE, attributing the original model, the Heretic derivative, and this quantization. See the original Qwen license and Heretic model card.