← back to catalog · registered 2026-08-22 13:56

Goldlionren00/AEON-Qwen3.6-27B-Ultimate-Uncensored-Multimodal-NVFP4-MTP-XS-GGUF

Goldlionren00 Qwen 27B GGUF multimodal 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/Goldlionren00%2FAEON-Qwen3.6-27B-Ultimate-Uncensored-Multimodal-NVFP4-MTP-XS-GGUF"
Response includes
  • classification m-uncensored
  • files 4
  • hub_downloads_all_time 6,372
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
6K
153 last 30d - cooling
Likes
2
Model age
5mo ago
created 2026-05-09
Downloads over time
Now6.4K→from0↑0%
02.4K4.7K7.1K0 on May 66.4K on Oct 11MayJunJulAugSepOct
May 6 → Oct 11 · 62 snapshots · spans 158 days

Metadata

Tags
llama.cpp gguf qwen qwen3 qwen3.6 nvfp4 mtp mtp-xs multimodal vision image-text-to-text rtx-5090

Related

Total size
17.5 GB
Files
4
Quantizations
2
Registered
2026-08-22 13:56
Last updated on HF
2026-05-09 09:17

Files by quantization

BF16 1 file 888 MB
mmproj-BF16.gguf 888 MB 05353347 download
Auxiliary files 3 files 17.5 GB
AEON-Qwen3.6-27B-Ultimate-Uncensored-Multimodal-NVFP4-MTP-XS.gguf 17.5 GB 550a5d8b download
README.md 4.97 KB 0a29934f download
.gitattributes 1.63 KB d0d27799 download

README current version from Hugging Face

license: apache-2.0
base_model:

  • AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-Multimodal-NVFP4-MTP-XS
    language:
  • en
  • zh
    library_name: llama.cpp
    tags:
  • gguf
  • qwen
  • qwen3
  • qwen3.6
  • nvfp4
  • mtp
  • mtp-xs
  • multimodal
  • vision
  • image-text-to-text
  • llama.cpp
  • rtx-5090
  • blackwell
  • conversational

AEON Qwen3.6 27B Ultimate Uncensored Multimodal NVFP4 MTP-XS GGUF

This repository contains a GGUF conversion of:

AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-Multimodal-NVFP4-MTP-XS

This is intended for local inference with llama.cpp-compatible runtimes, especially on NVIDIA Blackwell GPUs such as the RTX 5090.

Files

File Purpose
AEON-Qwen3.6-27B-Ultimate-Uncensored-Multimodal-NVFP4-MTP-XS.gguf Main GGUF model
mmproj-BF16.gguf Multimodal / vision projector for image input

Text-only usage only requires the main .gguf file.

Image input / multimodal usage requires both the main model and mmproj-BF16.gguf.

Model details

  • Architecture: Qwen3.6 / qwen35 family
  • Size class: 27B
  • Quantization: NVFP4
  • Format: GGUF
  • Source model format: nvidia-modelopt NVFP4 + MTP-XS
  • Modality: text + image input when used with the included mmproj
  • Recommended hardware: RTX 5090 / Blackwell-class GPU, preferably with 24–32GB+ VRAM

Important note about MTP

The upstream model is an NVFP4 MTP-XS model. In the original Hugging Face / vLLM workflow, MTP speculative decoding is used through the modelopt runtime path.

This GGUF conversion is provided primarily for llama.cpp-compatible local inference. Depending on your runtime version, the model may run as a standard GGUF model even if the original source model contains MTP-related tensors. Use MTP-specific acceleration only if your runtime explicitly supports it for this GGUF format.

Download

Using Hugging Face CLI:

hf download Goldlionren00/AEON-Qwen3.6-27B-Ultimate-Uncensored-Multimodal-NVFP4-MTP-XS-GGUF \
  --local-dir ./AEON-Qwen3.6-27B-Ultimate-Uncensored-Multimodal-NVFP4-MTP-XS-GGUF

Download only the main model:

hf download Goldlionren00/AEON-Qwen3.6-27B-Ultimate-Uncensored-Multimodal-NVFP4-MTP-XS-GGUF \
  AEON-Qwen3.6-27B-Ultimate-Uncensored-Multimodal-NVFP4-MTP-XS.gguf \
  --local-dir ./AEON-GGUF

Download the vision projector:

hf download Goldlionren00/AEON-Qwen3.6-27B-Ultimate-Uncensored-Multimodal-NVFP4-MTP-XS-GGUF \
  mmproj-BF16.gguf \
  --local-dir ./AEON-GGUF

llama.cpp usage

Text-only

llama-server \
  -m ./AEON-Qwen3.6-27B-Ultimate-Uncensored-Multimodal-NVFP4-MTP-XS.gguf \
  --host 0.0.0.0 \
  --port 10000 \
  -ngl 999 \
  -c 32768 \
  --flash-attn on

Multimodal / image input

llama-server \
  -m ./AEON-Qwen3.6-27B-Ultimate-Uncensored-Multimodal-NVFP4-MTP-XS.gguf \
  --mmproj ./mmproj-BF16.gguf \
  --host 0.0.0.0 \
  --port 10000 \
  -ngl 999 \
  -c 32768 \
  --flash-attn on

Windows example

.\llama-server.exe `
  -m "F:\Models\gguf\AEON-Qwen3.6-27B-Ultimate-Uncensored-Multimodal-NVFP4-MTP-XS.gguf" `
  --mmproj "F:\Models\gguf\mmproj-BF16.gguf" `
  --host 0.0.0.0 `
  --port 10000 `
  -ngl 999 `
  -c 32768 `
  --flash-attn on

For text-only inference, remove the --mmproj line.

Suggested runtime settings

Start conservative:

-c 32768

If stable and there is enough VRAM headroom, try:

-c 65536

For very long context, monitor VRAM usage carefully. The maximum usable context depends on the runtime, KV-cache type, batch size, GPU memory, and whether multimodal input is enabled.

LM Studio

  1. Download both files.
  2. Load the main .gguf file in LM Studio.
  3. If using image input, manually select mmproj-BF16.gguf as the multimodal projector if it is not detected automatically.

Compatibility notes

  • Best suited for Blackwell GPUs with native FP4/NVFP4 support, such as RTX 5090.
  • Older NVIDIA GPUs may run the model depending on runtime support, but may not benefit from native FP4 acceleration.
  • If the model fails to load, update to a recent llama.cpp build with NVFP4 GGUF support.
  • Multimodal behavior depends on correct pairing of the main GGUF and mmproj-BF16.gguf.

Conversion notes

This model was converted locally from the upstream AEON-7 NVFP4 MTP-XS model into GGUF format for llama.cpp-compatible runtimes.

The conversion was tested locally before upload.

Source / provenance

Upstream source model:

AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-Multimodal-NVFP4-MTP-XS

Base model lineage:

Qwen/Qwen3.6-27B

Please refer to the upstream AEON-7 model card for full details about the original quantization recipe, MTP-XS design, calibration, provenance, and deployment guidance.

License

Apache-2.0, following the upstream model license.

Responsibility

This is an uncensored model. Users are responsible for downstream safety controls, moderation, logging, access control, and compliance with applicable laws and policies.

Do not deploy this model in high-risk or production environments without appropriate safeguards.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-05-09Upload README.md with huggingface_hub508b2cc5 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration