← back to catalog · registered 2026-08-22 13:56

afkaf/Qwen3.8-27B-uncensored-w4a8-convrot-ComfyUI

afkaf Qwen 27B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/afkaf%2FQwen3.8-27B-uncensored-w4a8-convrot-ComfyUI"
Response includes
  • classification m1
  • files 3
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · 30-day
0
Likes
3
Model age
8w ago
created 2026-08-16
Downloads over time
Now0→from0↑0%
00110 on Aug 190 on Oct 11AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en
Tags
comfyui text-encoder quantized w4a8 convrot int4 abliterated uncensored qwen3_5 vision-language image-text-to-text en

Related

Total size
20.2 GB
Files
3
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-16 07:35

Files by quantization

Auxiliary files 3 files 20.2 GB
qwen3.8_27b_uncensored_w4a8_convrot.safetensors 20.2 GB 1d066534 download
README.md 8.83 KB 2d28efd4 download
.gitattributes 1.48 KB a6344aac download

README current version from Hugging Face


license: apache-2.0
base_model:

  • AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16
  • Qwen/Qwen3.8-27B
    base_model_relation: quantized
    pipeline_tag: image-text-to-text
    library_name: comfyui
    tags:
  • comfyui
  • text-encoder
  • quantized
  • w4a8
  • convrot
  • int4
  • abliterated
  • uncensored
  • qwen3_5
  • vision-language
    language:
  • en

Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16 — W4A8 ConvRot for ComfyUI

A asym_w4a8_int8 quantization of AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16,
packaged as a single-file ComfyUI text encoder.

20.2 GB on disk, ~19.4 GB resident. Fits entirely in 24 GB of VRAM with
room for context, so it runs without offload on an RTX 3090 or 4090.

Built for single-image captioning and prompt generation through ComfyUI's
Generate Text node. It is a full vision-language model: load it with
CLIPLoader, wire an image in, and it describes what it sees.


Why this exists

An INT8 ConvRot build of this model lands around 32 GB. On a 24 GB card that
means streaming weights over PCIe on every forward pass, and generation is
memory-bandwidth-bound — every token re-reads the weights. Dropping to 4-bit
weights is not about disk space; it is about staying resident.

W4A8 gives no compute advantage over INT8. It executes on the same int8 GEMM
path and adds a codebook dequantization step on top. The entire win is that
16.4 GiB of weight traffic per token at ~936 GB/s beats streaming 10 GB over
a ~25 GB/s PCIe link, by roughly an order of magnitude.

Format

asym_w4a8_int8 — ConvRot-rotated int4 weights, a per-tensor Lloyd-Max
codebook, and fp8 group scales, executed on int8 GEMM. Calibration-free.

Each quantized layer carries four tensors:

tensor dtype role
weight int8 packed int4 indices, two per byte
weight_s_rel fp8_e4m3 per-group scale, one per 16 values
weight_s_channel fp32 per-output-row scale
weight_codebook fp32 the 16 quantization levels

Three tiers of scaling — codebook shape, group magnitude, row magnitude —
which is why 4-bit weights hold up here. The codebook is fitted to each
tensor's actual distribution rather than spacing levels uniformly, so the
packed values are indices into a learned table, not magnitudes.

ConvRot applies a group-wise Hadamard rotation before quantization, spreading
outliers across the group so no single large weight dominates its scale.
Group size 256 for the rotation, 16 for the codebook.

Precision plan

Not everything is quantized. 387 tensors across layers 1–62 go to 4-bit;
everything below stays BF16 for a specific reason.

kept at BF16 size why
lm_head 2.54 GB produces logits over 248,320 tokens. This is a text generator, so 4-bit noise here flips token choices directly.
embed_tokens 2.54 GB input side; error propagates through every layer downstream.
vision tower 0.92 GB linear_fc2 has in_features 4304 = 16 × 269, and 269 is prime — no usable ConvRot group size divides it. The choice is BF16 or non-rotated int8, and non-rotated int8 is the weakest option available. Comfy-Org's official INT8 builds keep the whole tower BF16 for the same reason.
layers 0 and 63 1.51 GB first and last decoder blocks.
in_proj_a, in_proj_b 22 MB Gated DeltaNet decay and beta gates, the analogue of Mamba's A and dt. Error here compounds through the recurrent state instead of staying local to one matmul.
conv1d, A_log, dt_bias, all norms 90 MB 1-D and 3-D parameters; quantizing them buys nothing.
mtp.* 0.79 GB see below.

On the MTP head

The multi-token-prediction head is retained at BF16 and costs zero VRAM —
ComfyUI has no speculative decoding path, so its Qwen3.5 implementation never
constructs those modules and the keys are dropped at load. It is kept because
Comfy-Org's reference Qwen3.5 files keep it, and because MTP is BF16-or-nothing:
quantized MTP weights collapse draft acceptance from the 79–85% range to
5–11%. Should ComfyUI ever gain speculative decoding, the weights are here and
usable. Until then it is 0.79 GB of disk and nothing else.

Requirements

  • ComfyUI ≥ 0.31.0 (w4a8 loader support)
  • comfy-kitchen ≥ 0.2.31 with AsymW4A8Int8Layout
  • PyTorch built against CUDA 13.0+
  • Compute capability ≥ 8.0 (Ampere or newer). INT8 tensor cores are the
    execution path, so 30-series and up.

Usage

Drop the .safetensors into ComfyUI/models/text_encoders/.

  1. CLIPLoader → select this file. The type dropdown is ignored; detection
    is shape-based and resolves to QWEN35_27B automatically.
  2. Wire CLIP into Generate Text.
  3. Wire a Load Image into its image input.
  4. Write your prompt and run.

Sampling

From the base model card, non-thinking mode:

temperature 0.7   top_p 0.80   top_k 20   min_p 0.0   presence_penalty 1.5   repetition_penalty=1.0

Thinking mode: temperature 1.0, top_p 0.95, top_k 20, presence_penalty=0.0.

Verification

Every quantized layer was checked for serialization fidelity against
comfy-kitchen's own output, not merely for "does it load."

tensors                 2747
parameters              17.472 B
size                    20.20 GB
quantized layers        387
dtypes                  BF16=812  F32=774  F8_E4M3=387  I8=387  U8=387
detect_te_model      -> QWEN35_27B
serializer           -> state_dict_tensors
serialization fidelity  ~0.000000
VERIFY PASSED

serialization fidelity compares this file's dequantized weights against
comfy-kitchen re-quantizing the same source tensor. Near-zero means the file
faithfully reproduces what the library itself produces — no scale or codebook
tensor silently dropped, which is a failure mode that loads without error and
computes wrong.

Note that weight_correction is absent by design: comfy-kitchen 0.2.31 does
not emit a correction tensor in codebook mode, so correction=None at load is
correct rather than a missing field.

Reproducing

Converted with a purpose-built script that reconciles the packed weight by
tensor identity rather than by naming convention, holds the DeltaNet gates and
edge layers out of quantization, and audits the result:

python convert_qwen38_w4a8.py \
    -i ./Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16 \
    -o ./qwen3.8_27b_uncensored_w4a8_convrot.safetensors \
    --comfyui-path ./ComfyUI --seed 42 --verify

--seed 42 pins the codebook fitting, which samples, so rebuilds are
comparable.


Safety and user responsibility

This model has had its safety alignment removed. The upstream
AEON-7
release applied directional ablation (abliteration) to suppress refusal
behaviour. This repository contributes quantization only — no alignment,
guardrails, or filtering were added, and none were removed here either.

Consequences you are accepting by using it:

  • It will not refuse. It will attempt to answer requests that the original
    Qwen3.8-27B declines, including harmful, illegal, or dangerous ones. There is
    no residual safety layer to catch anything.
  • Ablation degrades more than refusals. Suppressing refusal directions
    perturbs the model's weights generally. Expect some loss of judgment,
    calibration, and factual reliability relative to the original, in ways that
    are not confined to safety-adjacent topics.
  • 4-bit quantization compounds this. The precision plan above minimizes
    it, but this is a lossy artifact of a lossy artifact.
  • Not suitable for unsupervised or public-facing deployment. If you expose
    this to users who are not you, you need your own moderation layer. It has none.
  • You are responsible for what you generate. Output is your
    responsibility, not the model's, not this repository's, and not any upstream
    author's. You are responsible for compliance with applicable law and with the
    licenses of all upstream artifacts.

No warranty of any kind. Provided as-is.

Attribution and non-endorsement

This is an unofficial, community-produced derivative. It is not endorsed
by, affiliated with, or supported by the Qwen team, Alibaba Cloud, AEON-7, or Comfy-Org. The behaviour of this model does not reflect the
intentions, standards, or positions of any of them. Do not report issues with
this model to those projects — its modifications are not theirs.

Please respect the license terms of all upstream artifacts. Verify the
license field above against both parent repositories before redistributing.

README history 6 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-16Update README.md6d0d67e8.8 KB
    Loading...
  2. 2026-08-16Update README.md7c169dd8.8 KB
    Loading...
  3. 2026-08-16Update README.md28a66248.8 KB
    Loading...
  4. 2026-08-16Update README.mdfa6700d8.8 KB
    Loading...
  5. 2026-08-16Upload README.md with huggingface_hub732c6f38.8 KB
    Loading...
  6. 2026-08-16initial commitd09484928 B
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration