← back to catalog · registered 2026-08-22 13:56

zm3/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-NVFP4A16

zm3 Qwen 24B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/zm3%2FQwen3.8-27B-AEON-ULTIMATE-UNCENSORED-NVFP4A16"
Response includes
  • classification m-uncensored
  • files 13
  • hub_downloads_all_time 1
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
1
0
Likes
0
Model age
7w ago
created 2026-08-18
Downloads over time
Now1→from1↑0%
11221 on Aug 191 on Oct 11AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh
Tags
transformers safetensors qwen3_5 image-text-to-text nvfp4 compressed-tensors quantized w4a16 vllm blackwell conversational en

Related

Total size
19.1 GB
Files
13
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-18 15:54

Files by quantization

Auxiliary files 13 files 19.2 GB
model.safetensors 18.4 GB ******** download
model-mtp.safetensors 810 MB ******** download
tokenizer.json 19.1 MB ******** download
config.json 14.4 KB f3679b05 download
chat_template.jinja 8.74 KB c0c686f9 download
README.md 6.71 KB 14941736 download
.gitattributes 1.53 KB 52373fe2 download
tokenizer_config.json 1.24 KB a9eacca6 download
processor_config.json 1.16 KB 33818c7f download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
recipe.yaml 273 B 3c311a86 download
generation_config.json 214 B 0bc3addd download

README current version from Hugging Face


license: apache-2.0
base_model: AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16
base_model_relation: quantized
library_name: transformers
pipeline_tag: image-text-to-text
tags:

  • nvfp4
  • compressed-tensors
  • quantized
  • w4a16
  • vllm
  • blackwell
    language:
  • en
  • zh

Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED — NVFP4A16 (W4A16)

NVFP4 weight-only quantization of
AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16.

51.75 GiB → 19.15 GiB (2.70× smaller).

Read this before deploying: this checkpoint is for Blackwell (SM100/SM120).
NVFP4 GEMMs execute on Blackwell tensor cores. It was produced on an H100
(SM 9.0), which has no NVFP4 hardware path — on Hopper the runtime unpacks
weights to bf16 at inference time, so expect no speedup there, and higher
peak VRAM than bf16 despite the smaller checkpoint. Disk and load-time savings
are real on any hardware; inference savings are not.

Format

Scheme NVFP4A16 (4-bit weights, bf16 activations)
Elements FP4 E2M1, packed 2/byte
Block scale float8_e4m3, group_size 16
Global scale FP32, per tensor
Strategy tensor_group
Checkpoint format nvfp4-pack-quantized (compressed-tensors)

Effective ~4.5 bits/weight (4 + 8/16). The block-16 FP8 scale is what
distinguishes NVFP4 from MXFP4, which uses block-32 with a power-of-two E8M0
scale — the finer, non-power-of-two scale retains noticeably more accuracy.

What was quantized

496 Linear modules — 24.351 B params, 87.7% of the model.

Component Params Precision
MLP (64 layers) 17.113 B NVFP4
Linear-attention (48 layers) 5.562 B NVFP4
Full-attention (16 layers) 1.678 B NVFP4
lm_head 1.271 B bf16
Embeddings 1.271 B bf16
Vision tower 0.461 B bf16
MTP head 0.425 B bf16

The vision tower, lm_head, and embeddings are held at bf16 deliberately —
together 12.3% of parameters with the MTP head, but the usual accuracy
casualties of 4-bit. The MTP head ships unmodified as model-mtp.safetensors,
so speculative decoding is preserved.

A_log, dt_bias, conv1d, and all norms are not nn.Linear and remain
untouched, which matters here: the model keeps its SSM state in fp32
(mamba_ssm_dtype: float32).

Verified at the tensor level: 496 weight_packed / 496 weight_scale
(F8_E4M3) / 496 weight_global_scale (F32), 0 in model.visual.*,
and 15 mtp.* tensors carried through.

Memory

Max context is 262,144 tokens. Only 16 of 64 layers use full attention;
the other 48 are linear-attention layers whose recurrent state is constant in
sequence length.

KV cache        64 KiB/token   (16 full-attn layers x 2 x 4 kv-heads x 256 head_dim x bf16)
KV @ 256K       16.00 GiB
linear state     0.141 GiB     <- constant, at 1 token or at 262,144
                (48 layers x 48 v-heads x 128 x 128, fp32)

Peak at full 256K context, single sequence:

Weights KV State Peak 256K seqs in 80 GB
bf16 51.75 16.00 0.14 ~67.9 GiB 1
NVFP4A16 19.15 16.00 0.14 ~35.3 GiB 3

An FP8 KV cache halves the 16 GiB.

Smoke test

Measured on H100 80GB (SM 9.0) — not the target hardware. transformers reference
path, greedy, 40 new tokens, causal-conv1d / flash-linear-attention fast path
absent:

Load time 6.7 s
VRAM after load 18.36 GiB
Peak VRAM 55.59 GiB
Throughput 5.4 tok/s

Peak VRAM exceeds the checkpoint size because, without NVFP4 tensor cores, the
runtime unpacks weights to bf16 while still holding the packed copies. Output was
coherent. This is a correctness check, not a benchmark — vLLM has quantized
kernels transformers lacks, so benchmark your own serving stack.

Reproducing

from transformers import AutoModelForImageTextToText, AutoProcessor
from llmcompressor import oneshot
from llmcompressor.modifiers.quantization import QuantizationModifier
import torch

model = AutoModelForImageTextToText.from_pretrained(
    "AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16", dtype=torch.bfloat16, device_map="cuda:0",
)
recipe = QuantizationModifier(
    targets="Linear", scheme="NVFP4A16",
    # NOTE: entries are matched as LITERAL module names unless prefixed "re:".
    # Without the prefix these regexes match nothing and the vision tower is
    # quantized anyway -- while still exiting 0 with a valid-looking checkpoint.
    ignore=["lm_head", r"re:model\.visual\..*", r"re:mtp\..*", r"re:.*embed_tokens.*"],
)
oneshot(model=model, recipe=recipe, processor=AutoProcessor.from_pretrained("AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16"),
        output_dir="out", save_compressed=True)

Data-free RTN — no calibration set. NVFP4's global scale derives from the
weights themselves, and group_size 16 is fine-grained enough that RTN holds up.
This also avoids a calibration-distribution mismatch, and avoids pushing
calibration batches through a hybrid linear-attention stack that llm-compressor
has not been validated against.

The MTP tensors must be copied through manually: transformers 5.14.1 does not
instantiate the MTP module, so save_pretrained silently drops them. torchvision
must be installed or the Qwen3VLProcessor fails to initialize and the processor
files are silently omitted from the output.

Produced with llmcompressor 0.13.0, compressed-tensors 0.18.0,
transformers 5.14.1, torch 2.13.0+cu130, on an NVIDIA H100 80GB.

Caveats

  • Not validated on target hardware. Produced and tested on Hopper only.
    Blackwell's NVFP4 GEMM has its own accumulation behavior.
  • No accuracy evaluation has been run. Only format correctness and a smoke
    generation. No perplexity, no benchmark suite, no vision-path evaluation.
  • Long-context accuracy is unmeasured. 48 of 64 layers are recurrent, so
    quantization error can compound along the sequence rather than staying bounded
    per token as in softmax attention. Published NVFP4 recipes were validated on
    pure transformers. Evaluate at 2K / 32K / 128K — a single short-context
    perplexity number will not surface this.
  • Quantization perturbs model behavior generally, including tone and refusal
    patterns.
  • Base-model provenance. The base is tagged early-access / draft and is an
    abliterated, refusal-removed fine-tune. Its shards are named inconsistently
    (model-00003-of-00003.safetensors alongside a 2-shard set); the index maps it
    correctly — it is the MTP head — but be aware if you script against those names.
    This checkpoint inherits the base model's behavior, which has no safety
    filtering; nothing here adds or restores any.
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration