license: apache-2.0
base_model:
- AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-Multimodal-NVFP4-MTP-XS
language: - en
- zh
library_name: llama.cpp
tags: - gguf
- qwen
- qwen3
- qwen3.6
- nvfp4
- mtp
- mtp-xs
- multimodal
- vision
- image-text-to-text
- llama.cpp
- rtx-5090
- blackwell
- conversational
AEON Qwen3.6 27B Ultimate Uncensored Multimodal NVFP4 MTP-XS GGUF
This repository contains a GGUF conversion of:
AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-Multimodal-NVFP4-MTP-XS
This is intended for local inference with llama.cpp-compatible runtimes, especially on NVIDIA Blackwell GPUs such as the RTX 5090.
Files
| File | Purpose |
|---|---|
AEON-Qwen3.6-27B-Ultimate-Uncensored-Multimodal-NVFP4-MTP-XS.gguf |
Main GGUF model |
mmproj-BF16.gguf |
Multimodal / vision projector for image input |
Text-only usage only requires the main .gguf file.
Image input / multimodal usage requires both the main model and mmproj-BF16.gguf.
Model details
- Architecture: Qwen3.6 / qwen35 family
- Size class: 27B
- Quantization: NVFP4
- Format: GGUF
- Source model format:
nvidia-modeloptNVFP4 + MTP-XS - Modality: text + image input when used with the included mmproj
- Recommended hardware: RTX 5090 / Blackwell-class GPU, preferably with 24–32GB+ VRAM
Important note about MTP
The upstream model is an NVFP4 MTP-XS model. In the original Hugging Face / vLLM workflow, MTP speculative decoding is used through the modelopt runtime path.
This GGUF conversion is provided primarily for llama.cpp-compatible local inference. Depending on your runtime version, the model may run as a standard GGUF model even if the original source model contains MTP-related tensors. Use MTP-specific acceleration only if your runtime explicitly supports it for this GGUF format.
Download
Using Hugging Face CLI:
hf download Goldlionren00/AEON-Qwen3.6-27B-Ultimate-Uncensored-Multimodal-NVFP4-MTP-XS-GGUF \
--local-dir ./AEON-Qwen3.6-27B-Ultimate-Uncensored-Multimodal-NVFP4-MTP-XS-GGUF
Download only the main model:
hf download Goldlionren00/AEON-Qwen3.6-27B-Ultimate-Uncensored-Multimodal-NVFP4-MTP-XS-GGUF \
AEON-Qwen3.6-27B-Ultimate-Uncensored-Multimodal-NVFP4-MTP-XS.gguf \
--local-dir ./AEON-GGUF
Download the vision projector:
hf download Goldlionren00/AEON-Qwen3.6-27B-Ultimate-Uncensored-Multimodal-NVFP4-MTP-XS-GGUF \
mmproj-BF16.gguf \
--local-dir ./AEON-GGUF
llama.cpp usage
Text-only
llama-server \
-m ./AEON-Qwen3.6-27B-Ultimate-Uncensored-Multimodal-NVFP4-MTP-XS.gguf \
--host 0.0.0.0 \
--port 10000 \
-ngl 999 \
-c 32768 \
--flash-attn on
Multimodal / image input
llama-server \
-m ./AEON-Qwen3.6-27B-Ultimate-Uncensored-Multimodal-NVFP4-MTP-XS.gguf \
--mmproj ./mmproj-BF16.gguf \
--host 0.0.0.0 \
--port 10000 \
-ngl 999 \
-c 32768 \
--flash-attn on
Windows example
.\llama-server.exe `
-m "F:\Models\gguf\AEON-Qwen3.6-27B-Ultimate-Uncensored-Multimodal-NVFP4-MTP-XS.gguf" `
--mmproj "F:\Models\gguf\mmproj-BF16.gguf" `
--host 0.0.0.0 `
--port 10000 `
-ngl 999 `
-c 32768 `
--flash-attn on
For text-only inference, remove the --mmproj line.
Suggested runtime settings
Start conservative:
-c 32768
If stable and there is enough VRAM headroom, try:
-c 65536
For very long context, monitor VRAM usage carefully. The maximum usable context depends on the runtime, KV-cache type, batch size, GPU memory, and whether multimodal input is enabled.
LM Studio
- Download both files.
- Load the main
.gguffile in LM Studio. - If using image input, manually select
mmproj-BF16.ggufas the multimodal projector if it is not detected automatically.
Compatibility notes
- Best suited for Blackwell GPUs with native FP4/NVFP4 support, such as RTX 5090.
- Older NVIDIA GPUs may run the model depending on runtime support, but may not benefit from native FP4 acceleration.
- If the model fails to load, update to a recent
llama.cppbuild with NVFP4 GGUF support. - Multimodal behavior depends on correct pairing of the main GGUF and
mmproj-BF16.gguf.
Conversion notes
This model was converted locally from the upstream AEON-7 NVFP4 MTP-XS model into GGUF format for llama.cpp-compatible runtimes.
The conversion was tested locally before upload.
Source / provenance
Upstream source model:
AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-Multimodal-NVFP4-MTP-XS
Base model lineage:
Qwen/Qwen3.6-27B
Please refer to the upstream AEON-7 model card for full details about the original quantization recipe, MTP-XS design, calibration, provenance, and deployment guidance.
License
Apache-2.0, following the upstream model license.
Responsibility
This is an uncensored model. Users are responsible for downstream safety controls, moderation, logging, access control, and compliance with applicable laws and policies.
Do not deploy this model in high-risk or production environments without appropriate safeguards.