← back to catalog · registered 2026-08-22 13:56

AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-Text-NVFP4-MTP

AEON-7 Qwen 9.4B second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/AEON-7%2FQwen3.6-27B-AEON-Ultimate-Uncensored-Text-NVFP4-MTP"
Response includes
  • classification m1
  • files 10
  • hub_downloads_all_time 14,134
  • author_summary 32 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
14K
7K last 30d - active
Likes
1
Model age
5mo ago
created 2026-04-28
Downloads over time
Now14.9K→from923↑1,515%
2245.6K10.9K16.3K923 on Apr 2914.9K on Oct 11AprMayJunJulAugSepOct
Apr 29 → Oct 11 · 64 snapshots · spans 165 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh multilingual
Tags
transformers safetensors qwen3_5 image-text-to-text abliterated uncensored qwen3 qwen3.6 nvfp4 modelopt mtp multi-token-prediction

Related

Total size
25.7 GB
Files
10
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-10-06 01:37

Files by quantization

Auxiliary files 10 files 25.7 GB
model.safetensors 25.7 GB 6d70fc41 download
tokenizer.json 10.6 MB 530dc3d0 download
cartridge.jpg 366 KB e613e97c download
README.md 14.2 KB 253bb997 download
chat_template.jinja 7.58 KB a8755d82 download
config.json 7.18 KB 84c28033 download
hf_quant_config.json 3.05 KB e14e3a08 download
.gitattributes 1.58 KB 0caea137 download
tokenizer_config.json 1.20 KB 7cd4b692 download
generation_config.json 213 B 303cc942 download

README current version from Hugging Face


license: apache-2.0
base_model: AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-BF16
language:

  • en
  • zh
  • multilingual
    library_name: transformers
    pipeline_tag: text-generation
    tags:
  • abliterated
  • uncensored
  • qwen3
  • qwen3.6
  • nvfp4
  • modelopt
  • mtp
  • multi-token-prediction
  • speculative-decoding
  • hybrid-attention
  • mamba
  • gated-deltanet
  • text-only
  • aeon
  • rtx-pro-6000
  • b100
  • b200
  • dedicated-vram-blackwell
  • sm_120
  • sm_100

Qwen3.6-27B-AEON-Ultimate-Uncensored-Text-NVFP4-MTP

AEON Qwen — Supreme Being of the Digital Cosmos

Deployment, operations & benchmarks → github.com/AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-DFlash

The GitHub repo is the source of truth for the production deployment guide, hardware-tuned docker-compose configs, full configuration reference, measured benchmarks, and AGENTS.md — an operator's manual that pre-empts common stale-documentation traps.

🙏 Reference recipe credit: The modelopt + MTP graft pipeline used to build this variant is based on sakamakismile's validated Qwen3.6-27B-NVFP4-MTP series (22K+ downloads). They worked out the modelopt config, the per-projection quantization choices, and the MTP-head graft technique on the un-abliterated base; we adapted the same recipe to AEON-Ultimate's abliterated weights. The reference benchmark numbers cited below are theirs. Full credit for the recipe → sakamakismile.

🆕 AEON vLLM Ultimate container (2026-06-04)

ghcr.io/aeon-7/aeon-vllm-ultimate:latest — vLLM 0.23.0 (= :2026-06-18-v0.23.0-dflashfix) + PR #44389 NVFP4 KV cache (~3× capacity) + DFlash + TurboQuant K8V4 + AEON sm_121a patches. Same recipe family as the -Multimodal-NVFP4-MTP-XS sibling which has been benchmarked end-to-end (production-style greedy + n_spec=15 by category: math/code peak ~45 tok/s, overall mean 34.7 tok/s; concurrent ×4 steady ~84 tok/s aggregate). This variant uses the same modelopt NVFP4 format, the same qwen3_5_mtp native head, and the same hybrid GDN+attention stack — it should serve identically with --quantization modelopt and either --speculative-config '{"method":"qwen3_5_mtp","num_speculative_tokens":3}' (native MTP) or a DFlash drafter (recommended on Spark — see container README Recipe A).

The v3 image (ghcr.io/aeon-7/vllm-aeon-ultimate-dflash:qwen36-v3) remains the stable production target if you need FP8 KV + DFlash; in the new image DFlash requires --kv-cache-dtype auto (BF16). Full setup + 4-config bench comparison: container README.

Variants

Format Size Use case
BF16 51 GB Full-precision reference weights (A100/H100 80 GB, RTX PRO 6000 96 GB, multi-GPU, fine-tuning)
NVFP4 (compressed-tensors + DFlash) 26 GB DGX Spark / GB10 — production validated with DFlash speculative decoding. Patched vllm-aeon-ultimate-dflash container.
Multimodal-NVFP4-MTP 27 GB High-bandwidth dedicated GPUs (RTX 5090, RTX PRO 6000, B100/B200) with MTP speculative decoding via the model's native mtp.* head. modelopt format, --quantization modelopt. Vision tower preserved.
Text-NVFP4-MTP (this repo) 20 GB Same recipe but with vision tower stripped. Smaller footprint for text-only deployments on tighter VRAM (RTX 5090 32 GB fits comfortably).

What this is

This is the modelopt-format NVFP4 variant with MTP speculative decoding, text-only (vision tower stripped), of AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-BF16 — the lossless abliteration of Qwen 3.6 27B (KL 0.000492 vs base, 0/100 refusals, multimodal preserved, hybrid GDN-aware quantization).

Specifically:

  • Body quantized to NVFP4 via nvidia-modelopt 0.43.0 with NVFP4_DEFAULT_CFG. This is the modelopt compressed-tensors format that vLLM serves through --quantization modelopt (different code path from the -NVFP4 sibling release which uses --quantization compressed-tensors).
  • Linear-attn / GatedDeltaNet layers preserved BF16 (432 keys across 48 GDN layers). NVFP4 quantization on Mamba/SSM state collapses the recurrence; modelopt's *linear_attn.conv1d* ignore plus our explicit *linear_attn* exclude keeps these intact.
  • Vision tower stripped (333 visual keys removed, ~0.92 GB). Text-only build — no image / video input. language_model_only: true set in config.json.
  • MTP head grafted from the base Qwen/Qwen3.6-27B checkpoint (15 tensors, BF16). The base contains MTP heads but Qwen3_5ForConditionalGeneration.from_pretrained drops them during loading; the lna-lab pipeline pattern (which this build follows) explicitly grafts them back into the quantized output, giving vLLM a working drafter for --speculative-config '{"method":"qwen3_5_mtp",...}'.

Why MTP — and where it actually wins

Multi-Token Prediction (MTP) lets the model predict multiple future tokens per forward pass via the trained mtp.* head, enabling speculative decoding without a separate drafter model. The acceptance rate is high because the drafter is the model itself — same architecture, same weights, same distribution.

Measured numbers on AEON-Ultimate (this MTP family)

Hardware Median tok/s Peak tok/s Spec-decode acceptance
RTX PRO 6000 Blackwell (96 GB dedicated VRAM) ~92 (regular) / 111.4 (XS sibling) 124.7 (XS sibling) 67.7 % regular / 69.2 % XS
DGX Spark / GB10 (unified memory) — MTP method 24.1 (XS sibling) 27.5 66.3 %
DGX Spark / GB10 — DFlash on the same XS body 🏆 38.5 tok/s thinking-on / 38.1 off 71.3 tok/s thinking-on / 68.4 off DFlash v2
RTX 5090, B100 / B200 not yet measured by us — community welcome

Reference numbers from sakamakismile's un-abliterated recipe (RTX 5090)

  • Single-stream short prompts at n=3: ~132 tok/s
  • Single-stream long-form: ~105 tok/s
  • 2-parallel aggregate (256K + KV FP8): ~189–207 tok/s
  • Mean MTP acceptance length: ~3.0–4.0 (vs DFlash chains ~2.0–2.3)

The hardware-routing punchline

On RTX PRO 6000 the XS sibling beats DFlash territory (~111 tok/s vs DFlash-class ~85 we'd expect there). On DGX Spark, DFlash beats MTP by 26 % median / 52 % peak — the unified-memory bandwidth caps how much MTP's high acceptance can translate to throughput. So: MTP is a dedicated-VRAM-Blackwell variant, not a universal upgrade. Full bench data: GitHub repo Performance section.

🎯 When to pick this variant — measured hardware routing

The right speculative-decode method depends on memory architecture:

Hardware tier Recommended variant Why
DGX Spark / GB10 (sm_121a, unified memory) -NVFP4 (DFlash) — not this MTP variant Bench on Spark: DFlash beats MTP by +26 % median, +52 % peak. Spark's unified-memory bandwidth doesn't reward MTP's high acceptance rate. Don't run MTP on Spark.
RTX PRO 6000 Blackwell (sm_120, 96 GB dedicated VRAM) This variant ✅ if text-only; Multimodal if you need vision MTP wins on dedicated VRAM. ~92 tok/s median measured (multimodal sibling, GDN BF16).
RTX 5090 (sm_120, 32 GB dedicated VRAM) Text-XS is the better fit (~20 GB), or this variant if you have headroom XS variant matches sakamakismile's reference footprint. 111.4 tok/s median measured on RTX PRO 6000; RTX 5090 should land near or above.
A100 / H100 (no native FP4) BF16 NVFP4 dequantizes to BF16 on Ampere/Hopper — no benefit.
B100 / B200 (sm_100, dedicated FP4) This variant or Multimodal Native FP4 + dedicated VRAM = MTP territory.

Full bench numbers: GitHub repo Performance section.

Usage

vLLM serve

# One-time: pull this repo locally
hf download AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored-Text-NVFP4-MTP \
  --local-dir ./aeon-ultimate-text-nvfp4-mtp

# Serve
export VLLM_NVFP4_GEMM_BACKEND=flashinfer-cutlass
export VLLM_USE_FLASHINFER_MOE_FP4=0
export VLLM_USE_FLASHINFER_SAMPLER=1

vllm serve ./aeon-ultimate-text-nvfp4-mtp \
&
  --mamba-cache-dtype float32 \
  --trust-remote-code \
  --max-model-len 262144 \
  --max-num-seqs 32 \
  --max-num-batched-tokens 32768 \
  --gpu-memory-utilization 0.94 \
  --enable-chunked-prefill \
  --enable-prefix-caching \
  --reasoning-parser qwen3 \
  --tool-call-parser qwen3_coder \
  --enable-auto-tool-choice \
  --speculative-config '{"method":"qwen3_5_mtp","num_speculative_tokens":3}'

num_speculative_tokens=3 is the canonical setting for qwen3_5_mtp. Higher values diverge the drafter further from the target distribution and acceptance falls.

Configuration notes

  • --quantization modelopt is required (not compressed-tensors — different format).
  • --speculative-config '{"method":"qwen3_5_mtp", ...}' activates the grafted MTP head as the spec-decode drafter. No external drafter download needed — the head is in the safetensors of this repo.
  • --gpu-memory-utilization 0.94 is the validated cap on RTX PRO 6000; 0.95 causes the FlashInfer NVFP4 GEMM autotuner to OOM on first boot. See the GitHub repo's RTX PRO 6000 page for the same OOM behavior under DFlash.

Quantization recipe

  • Tool: nvidia-modelopt 0.43.0 with NVFP4_DEFAULT_CFG
  • Loader: Qwen3_5ForConditionalGeneration.from_pretrained (multimodal-preserved class)
  • Calibration: neuralmagic/calibration LLM split, 20 samples × 8192 tokens
  • Excluded from quantization (kept BF16):
    • lm_head, proj_out.*, *router*, *mlp.gate.* (NVFP4_DEFAULT_CFG)
    • *linear_attn.conv1d*, *mixer.conv1d* (NVFP4_DEFAULT_CFG)
    • *linear_attn* (added — full GDN preservation)
    • *visual* (added — vision tower preservation)
    • *mtp* (added — MTP head preservation)
    • *output_layer*, output.*
  • Vision strip: post-export, model.visual.* keys (333 tensors, ~0.92 GB) removed; vision_config removed from config.json; language_model_only: true set; preprocessor configs cleaned
  • MTP graft: 15 tensors copied bf16 from Qwen/Qwen3.6-27B after modelopt export (AutoModelForCausalLM.from_pretrained drops them; explicit graft restores)
  • Pipeline: lna-lab/GGUF-to-NVFP4-SM120 reference recipe, adapted for AEON-Ultimate-BF16 input + separate MTP source

Provenance & credits

License + responsibility

Apache 2.0, inherited from Qwen/Qwen3.6-27B. This is an uncensored model. Read the full User Responsibility & Arbitration Clause on the BF16 source card before deploying. Summary: you implement downstream safety layers (input validation, output filtering, content moderation, audit logging, rate limiting, access controls, human-in-the-loop for high-risk workflows). The model has no opinions of its own — you supply the opinions, the judgment, and the ethics.


☕ Support the work

If this release has been useful, tips are deeply appreciated — they go directly toward more compute, more models, and more open releases.

₿ Bitcoin (BTC)
QR
bc1q09xmzn00q4z3c5raene0f3pzn9d9pvawfm0py4
Ξ Ethereum (ETH)
QR
0x1512667F6D61454ad531d2E45C0a5d1fd82D0500
◎ Solana (SOL)
QR
DgQsjHdAnT5PNLQTNpJdpLS3tYGpVcsHQCkpoiAKsw8t
ⓜ Monero (XMR)
QR
836XrSKw4R76vNi3QPJ5Fa9ugcyvE2cWmKSPv3AhpTNNKvqP8v5ba9JRL4Vh7UnFNjDz3E2GXZDVVenu3rkZaNdUFhjAvgd

Ethereum L2s (Base, Arbitrum, Optimism, Polygon, etc.) and EVM-compatible tokens can be sent to the same Ethereum address.

README history 16 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-10-06Add Patreon support section716c5af15.7 KB
    Loading...
  2. 2026-09-15docs: recommend Qwen3.8 MIXED + GH recipe card1ce4d8c15.1 KB
    Loading...
  3. 2026-09-12docs: redirect to Qwen3.8 NVFP4-MIXED successor74326f515 KB
    Loading...
  4. 2026-06-28add AEON Qwen cover art87f365f14.2 KB
    Loading...
  5. 2026-06-24docs: arch-level update — mamba-cache-dtype float32 + container vLLM 0.23.0 (...c3270f114.2 KB
    Loading...
  6. 2026-06-21tags: expand to maximally-searchable set (+42 tags, union with existing)7e8a04020.1 KB
    Loading...
  7. 2026-06-18docs(quickstart): unify on aeon-vllm-ultimate:latest + validated serve flags ...850734219.6 KB
    Loading...
  8. 2026-06-18docs: vLLM compatibility status on aeon-vllm-ultimate:latest9d7f81019 KB
    Loading...
  9. 2026-06-17Container assoc -> aeon-vllm-ultimate:latest + long-context SWA note57681b514.8 KB
    Loading...
  10. 2026-06-05docs: add AEON vLLM Ultimate (v0.22.1 + PR #44389 NVFP4 KV) container referen...afea46914.1 KB
    Loading...
  11. 2026-05-31Tip jar: single left-aligned QR column (fix narrow-viewport clipping)1a7a17112.9 KB
    Loading...
  12. 2026-05-01Add tip jar block (BTC/ETH/SOL/XMR with QR codes)8bf600313 KB
    Loading...
  13. 2026-04-29Add Spark+DFlash+v3 row to perf table (38.5/71.3 thinking-on)47eb42211.5 KB
    Loading...
  14. 2026-04-28Upload README.md with huggingface_hubd662d8111.3 KB
    Loading...
  15. 2026-04-28Upload README.md with huggingface_hub9f670d710.6 KB
    Loading...
  16. 2026-04-28Add model card11531749.2 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration