← back to catalog · registered 2026-08-22 13:56

groxaxo/Qwen3.6-27B-AEON-Ultimate-Uncensored-GPTQ-Pro-FOEM-4bit-g128

groxaxo Qwen 24B second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/groxaxo%2FQwen3.6-27B-AEON-Ultimate-Uncensored-GPTQ-Pro-FOEM-4bit-g128"
Response includes
  • classification m-uncensored
  • files 18
  • hub_downloads_all_time 8,377
  • author_summary 27 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
8K
162 last 30d - cooling
Likes
3
Model age
5mo ago
created 2026-05-07
Downloads over time
Now8.4K→from135↑6,151%
03.1K6.2K9.3K135 on May 68.4K on Oct 11MayJunJulAugSepOct
May 6 → Oct 11 · 62 snapshots · spans 158 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
safetensors qwen3_5 gptq 4-bit quantized gptqmodel foem qwen3 uncensored text-generation conversational base_model:AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-BF16

Related

Total size
17.4 GB
Files
18
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-22 07:59

Files by quantization

Auxiliary files 18 files 17.4 GB
model-00004-of-00005.safetensors 3.99 GB d6666885 download
model-00002-of-00005.safetensors 3.98 GB 00ebed29 download
model-00003-of-00005.safetensors 3.97 GB 35bc47cb download
model-00001-of-00005.safetensors 3.28 GB f5763ae1 download
model-00005-of-00005.safetensors 2.21 GB 69479d02 download
tokenizer.json 19.1 MB 639e352c download
model.safetensors.index.json 223 KB d4bddac4 download
chat_template.jinja 7.58 KB a8755d82 download
config.json 5.35 KB 9e15188a download
README.md 5.01 KB a67be111 download
.gitattributes 1.53 KB 52373fe2 download
quantize_config.json 1.43 KB 08468547 download
tokenizer_config.json 1.20 KB 920f0987 download
processor_config.json 1.16 KB 33818c7f download
aeon_quant_metadata.json 1010 B b50ff414 download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
generation_config.json 187 B 91511aaa download

README current version from Hugging Face


base_model: AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-BF16
license: apache-2.0
tags:

  • gptq
  • 4-bit
  • quantized
  • gptqmodel
  • foem
  • qwen3
  • uncensored
    pipeline_tag: text-generation

Qwen3.6-27B-AEON-Ultimate-Uncensored — GPTQ-Pro FOEM 4-bit g128

Overview

Qwen3.6-27B-AEON-Ultimate-Uncensored-GPTQ-Pro-FOEM-4bit-g128 is a GPTQ-quantized checkpoint intended for efficient GPU inference, published by groxaxo.
It is intended for open-source evaluation, reproducible experimentation, and compatible local or
hosted inference workflows. The wording below is deliberately limited to what can be verified
from this repository's metadata and artifacts.

The repository name identifies a behavior-modified or reduced-filtering lineage. That label describes the source or conversion history; it is not a guarantee of unrestricted behavior in every prompt or runtime. Test outputs carefully before sharing or deploying them.

At a glance

Field Details
Format GPTQ
Source / base AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-BF16
Intended task image-text-to-text
License apache-2.0

What is included

  • *.safetensors (5 files)
  • config.json
  • generation_config.json
  • tokenizer.json
  • tokenizer_config.json
  • processor_config.json
  • chat_template.jinja
  • quantize_config.json
  • Additional configuration, tokenizer, processor, or shard files (16 visible artifacts total)

Quick start

vLLM (documented configuration)

vllm serve groxaxo/Qwen3.6-27B-AEON-Ultimate-Uncensored-GPTQ-Pro-FOEM-4bit-g128 \
  --quantization gptq_marlin \
  --dtype float16 \
  --trust-remote-code

This command is taken from the repository documentation. Adjust tensor parallelism, context
length, and cache settings to match your hardware and vLLM version.

Compatibility and responsible use

  • Use a runtime that explicitly supports this format, architecture, and modality.
  • Keep configuration, tokenizer, processor, projection, and weight files from the same revision together.
  • Review the source model card and license before redistribution or deployment.
  • Hardware needs depend on parameter count, context length, cache precision, quantization, and concurrency.
  • Report reproducible issues with the runtime version, hardware, launch command, and a minimal example.

Quantization or conversion changes numerical behavior, memory use, and throughput relative to the source checkpoint; validate quality on your own workload.

Generated outputs may be inaccurate or unsuitable for a given use case. Users are responsible for
testing behavior, applying appropriate safeguards, and complying with applicable licenses and laws.

Quantized version of AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-BF16
using GPTQModel with FOEM (First-Order Error Minimization) enhancement.


Quantization Recipe

Setting Value
Method GPTQ-Pro
Bits 4
Group size 128
Symmetric ✅
desc_act ❌
true_sequential ✅
FOEM alpha 0.25
FOEM beta 0.2
activation_weighted_mse ✅
lm_head quantized ❌
Kernel MarlinLinear (auto)

Quantized: language_model.layers linear modules only
(attn projections, MLP gate/up/down)

Preserved in full BF16:

  • model.visual.* — entire vision tower (333 tensors)
  • lm_head.weight
  • model.language_model.embed_tokens.weight
  • All norm layers, RoPE, and multimodal glue

Perplexity Comparison (WikiText-2)

Evaluated on identical settings (512 ctx / 256 stride, wikitext-2-raw-v1 test set):

Model PPL Δ vs BF16
BF16 baseline 7.6228 —
GPTQ-Pro FOEM 4-bit 7.7447 +0.12 (+1.6%)

Only 1.6% perplexity degradation at 4-bit — an excellent result for W4G128 GPTQ. FOEM + activation-weighted MSE preserved language model fidelity across all 64 transformer layers.


Usage

from gptqmodel import GPTQModel, BACKEND

model = GPTQModel.load(
    "AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-GPTQ-Pro-FOEM-4bit-g128",
    device="cuda:0",
    backend=BACKEND.AUTO,
)

Quantization Details

  • Tool: GPTQModel v6.1.0-dev
  • Calibration: 64 samples from WikiText-2
  • Hardware: 3× RTX 3090/3060 (CUDA_VISIBLE_DEVICES=0,1,2)
  • Duration: ~112 minutes
  • gc_mode: on_stage_end (VRAM-safe for large VLMs)
  • Offload: disk offload enabled during quantization

About the Base Model

This quantization is based on AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-BF16, an uncensored variant of Qwen3 27B with vision capabilities (architecture: Qwen3_5ForConditionalGeneration).

README history 6 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-22Polish model card overview and usage notes9d3b56a5 KB
    Loading...
  2. 2026-08-22Polish model card overview and usage notes16e04085 KB
    Loading...
  3. 2026-08-22Polish model card overview and usage notesa1add145 KB
    Loading...
  4. 2026-08-22Polish model card overview and usage notescb0337a4.2 KB
    Loading...
  5. 2026-05-07Upload README.md with huggingface_hubf6b4a942.4 KB
    Loading...
  6. 2026-05-07Upload folder using huggingface_hub3f7e4282.3 KB
    Loading...

Discussions 3 threads

  1. 2026-07-26please fix the mtp, it's a very good model!open1 💬#3
    Loading...
  2. 2026-06-05Decoder output is garbage on vLLM?open1 💬#2
    Loading...
  3. 2026-05-08this model not support MTP?open1 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration