license: apache-2.0
base_model: AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16
base_model_relation: quantized
library_name: transformers
pipeline_tag: text-generation
language:
- en
- zh
- multilingual
tags: - quark
- mxfp4
- mxfp6
- ocp-mx
- rocm
- amd
- radeon
- instinct
- qwen3.8
- gated-deltanet
gated: manual
extra_gated_heading: "Request early access: AEON Ultimate MXFP4/MXFP6 for ROCm"
extra_gated_description: "Experimental early-access build; approval by hand."
extra_gated_button_content: "Request access"
extra_gated_fields:
Name (as shown on your Hugging Face profile): text
GPU model (for example Radeon AI PRO R9700 or Instinct MI355X): text
I understand this is an experimental build: checkbox
Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-MXFP4-MXFP6-ROCm
The AEON Ultimate uncensored 27B, quantized to OCP Microscaling (MX) formats for AMD GPUs.
The MLP runs in MXFP4, the Gated DeltaNet projections in MXFP6, and everything that carries the model's skills stays in BF16.
The checkpoint is 23.7 GB on disk and loads as about 22.8 GB of weights. Two 32 GB Radeon AI PRO R9700 cards hold it with the full 262,144-token context; one R9700 holds it with a shorter context. MXFP4 is also the format Instinct MI350 / MI355 execute natively.
[!WARNING]
Experimental: early access. Nobody has published runtime results for this build yet. The first Radeon test pass (2x R9700) is under way, and quality and speed numbers will appear in this card as they come in. Access is approved by hand.
What this is
- Base model: AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16 is AEON-7's uncensored build of Qwen/Qwen3.8-27B. It is a hybrid of 48 Gated DeltaNet (linear-attention) layers and 16 full-attention layers, with a vision tower, an MTP head and a native context of 262,144 tokens.
- Format: OCP Microscaling. Every block of 32 values shares one power-of-two (E8M0) scale. We used AMD Quark 0.12.post1 and exported real packed weights, not fake-quant.
- MLP (
gate_proj,up_proj,down_proj) is MXFP4, W4A4: FP4 E2M1 weights, with activations quantized to MXFP4 on the fly. - Gated DeltaNet projections (
in_proj_qkv,in_proj_z,out_proj) are MXFP6 E2M3, W6A6. - The rest stays BF16: full attention, embeddings,
lm_head, the vision tower, the MTP head, the GDN recurrence (conv1d,in_proj_a/b,A_log,dt_bias) and the norms.
- MLP (
- Unchanged from the BF16 master: the tokenizer, chat template and vision processor. Thinking is on by default.
Why MX on AMD? MX is the open block-scaled format that AMD's newest accelerators run in hardware. On Radeon the gain is footprint: the BF16 master is 55.6 GB, while this build is 23.7 GB. The precision is mixed on purpose. The MLP and GDN projections take 4-bit and 6-bit weights, and the attention path, the recurrence, the embeddings and the vision tower stay in full BF16.
Target hardware
Every AMD system this build can run on has its own labelled recipe in the Quickstart. In short:
- Native MXFP4: Instinct MI350X / MI355X (gfx950). The MLP runs on AITER's W4A4 GEMM. The MXFP6 layers are emulated, because vLLM has no MXFP6 kernel yet.
- Emulated: everything else, including Radeon RDNA 4 (R9700, RX 9070 XT), RDNA 3 (RX 7900 XTX, W7900, W7800), Strix Halo and Instinct MI300X / MI325X. vLLM keeps the packed MX weights in memory, dequantizes each layer to BF16 on every forward pass, quantizes and dequantizes the activations to the same MX grid, and runs the matmul in BF16. The numerics match W4A4 and W6A6 and memory stays near the packed size, but decoding is slower than a native kernel would be (vLLM 0.29 and 0.31 source:
QuarkOCP_MX,EmulationMxfp4LinearKernel,EmulationMxfp6LinearKernel; native MX needssupports_mx(), which is true only on gfx95x and gfx1250).
| Other platform | Runs? | Notes |
|---|---|---|
| NVIDIA Blackwell (B200, RTX 50-series, RTX PRO 6000, DGX Spark) | Should load in vLLM, not yet tested | For MXFP4 vLLM picks FlashInfer native W4A4 (compute capability ≥ 10.0, when installed), otherwise Marlin weight-only (activations stay BF16), otherwise emulation; MXFP6 is emulated. Emulation needs amd-quark in the image. On NVIDIA, use NVFP4-MIXED instead. |
| Apple Silicon | Not with this checkpoint | An MLX conversion is being tested; results will be added here. |
Quickstart
Deploying with an AI agent (Claude Code, Codex, Cursor)? Point it at
AGENTS.md. It has the preflight, the full configuration ladder with exact commands, verification steps and an exhaustive troubleshooting table.Testing before an event? Run
bash scripts/preflight_smoke.shthe day before. It checks the host, starts the recommended configuration, smoke-tests it in a few minutes and bundles the results to send back.
Step 0: Which setup do you have? (1 minute)
Run this on the Linux host. It reads the amdgpu driver directly, so it needs no ROCm tools, no Docker and no root:
# one line per AMD GPU: gfx target, VRAM, render node, PCI name
for p in /sys/class/kfd/kfd/topology/nodes/*/properties; do
v=$(awk '$1=="gfx_target_version"{print $2}' "$p"); [ "${v:-0}" -gt 0 ] || continue
r=$(awk '$1=="drm_render_minor"{print $2}' "$p"); d=/sys/class/drm/renderD$r/device
printf 'gfx%d%x%x %5.1f GiB /dev/dri/renderD%s %s\n' $((v/10000)) $((v/100%100)) $((v%100)) \
"$(awk '{print $1/2^30}' "$d/mem_info_vram_total")" "$r" \
"$(lspci -s "$(basename "$(readlink -f "$d")")" 2>/dev/null | cut -d' ' -f2-)"
done
With ROCm tools on the host, amd-smi static --asic --vram or rocm-smi --showproductname --showmeminfo vram show the same thing. After the download, python3 scripts/detect_setup.py does this lookup and prints the matching recipe and launch script, and bash scripts/preflight.sh checks the rest of the host (driver, groups, PCIe slots, IOMMU, Resizable BAR, disk, ports).
Match the output to the table below. Count only the discrete GPUs.
- Ignore an integrated GPU (gfx1036, gfx1035, gfx1103, gfx1150 and similar, with a few GiB of VRAM), unless it gets passed into the container. vLLM reads the GPU architecture from the first GPU that
amdsmilists, andHIP_VISIBLE_DEVICESdoesn't change that. If the iGPU comes first, the RDNA 4 paths switch off. - To hide it, either disable the iGPU in the BIOS, or pass only the discrete GPUs' nodes into the container: replace
--device /dev/driwith--device /dev/dri/renderD129 --device /dev/dri/card1 ...for each discrete GPU. The launch scripts do this withAEON_DEVICES=auto, and they refuse to start if vLLM would see the wrong architecture.
Pick your system
Status: Verified = run with this checkpoint (none yet). Expected = the vLLM code path for this GPU architecture and GPU count is confirmed from source, and published reports run other Qwen3.x checkpoints in vLLM on the same GPU family (RDNA 4). Untested = the code path should work, but nobody has run anything like it on that hardware. Not supported = it doesn't fit.
| Your GPUs | VRAM | gfx | MX path | Status | Recipe |
|---|---|---|---|---|---|
| 2x Radeon AI PRO R9700 | 2 × 32 GB | gfx1201 | emulated | Expected | Ladder: A (recommended) → A+ (fast) → A-Max (max); fallback Z; variant A+DF |
| 1x Radeon AI PRO R9700 | 32 GB | gfx1201 | emulated | Expected | B |
| 2x Radeon RX 9070 XT / 9070 | 2 × 16 GB | gfx1201 | emulated | Expected (tight) | C |
| 4x Radeon RX 9070 XT / 9070 | 4 × 16 GB | gfx1201 | emulated | Untested | D |
| 2x Radeon RX 7900 XTX | 2 × 24 GB | gfx1100 | emulated | Untested | E |
| Radeon PRO W7900, or W7800 48 GB | 48 GB | gfx1100 | emulated | Untested | F |
| Radeon PRO W7800 | 32 GB | gfx1100 | emulated | Untested | G |
| Ryzen AI Max+ 395 (Strix Halo) | 64 or 128 GB unified | gfx1151 | emulated | Untested | H |
| Instinct MI300X / MI325X | 192 / 256 GB | gfx942 | emulated | Untested | I |
| Instinct MI350X / MI355X | 288 GB | gfx950 | MXFP4 native, MXFP6 emulated | Untested | J |
| 1x RX 7900 XTX, 1x RX 9070 XT / 9070, any single GPU under 32 GB | Not supported | why |
Common setup (every recipe)
You need Linux on bare metal (not a VM with GPU passthrough: RCCL is reported to hang there), a recent amdgpu driver for your GPU, Docker, and about 80 GB of free disk.
Every recipe uses the official vLLM ROCm image vllm/vllm-openai-rocm:v0.31.0 (ROCm 7.2.3; built for gfx90a, gfx942, gfx950, gfx1100, gfx1101, gfx1150, gfx1151, gfx1200 and gfx1201; includes amd-quark).
- Don't use v0.28, v0.29 or v0.30. They crash while loading this checkpoint with
AttributeError: 'dict' object has no attribute 'endswith', because their Quark loader can't handle the list-valuedalgo_configinconfig.json. v0.31.0 fixes it. - Don't use the
nightly/rocm100images on Radeon. They ship no gfx1201 kernels (ROCm/aiter#5229). - After the first pull, pin the digest:
docker image inspect vllm/vllm-openai-rocm:v0.31.0 --format '{{index .RepoDigests 0}}', then useAEON_IMAGE=vllm/vllm-openai-rocm@sha256:....
# 1) After your access request is approved (log in with your own account and token)
hf auth login
hf download AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-MXFP4-MXFP6-ROCm \
--local-dir ~/models/aeon-mxfp4-rocm
# early-access downloads include dflash2/ (3.85 GB, only Recipe A+DF needs it); skip it with --exclude "dflash2/*"
# 2) Variables the commands below use (re-run in every new terminal)
export AEON_DIR="$HOME/models/aeon-mxfp4-rocm"
export AEON_CACHE="$HOME/.cache/aeon-mxfp4-rocm" # compiled kernels persist here
export RENDER_GID="$(getent group render | cut -d: -f3)"; [ -n "$RENDER_GID" ] || RENDER_GID=video
mkdir -p "$AEON_CACHE"
# 3) Check the host
bash "$AEON_DIR/scripts/preflight.sh"
Each recipe below has a one-line launch script (in scripts/, shipped with the model) and the plain docker run it executes. The scripts also check which GPUs and which vLLM version the container sees before starting, refuse known-bad combinations, wait for the server, and save the startup log. Every value can be changed through an environment variable, for example AEON_MAX_MODEL_LEN=262144 AEON_MAX_NUM_SEQS=4 bash "$AEON_DIR/scripts/serve_2x_r9700.sh"; AEON_DRY_RUN=1 prints the command without starting anything. The header of each script lists its target system, status and memory estimate.
Then talk to it:
curl -s http://127.0.0.1:8000/v1/chat/completions -H "Content-Type: application/json" -d '{
"model": "aeon",
"messages": [{"role": "user", "content": "Explain MX block scaling in two sentences."}],
"max_tokens": 200,
"chat_template_kwargs": {"enable_thinking": false}
}'
Applies to every recipe:
- Don't pass
--quantization. vLLM reads the Quark config fromconfig.json. - The first start is slow (allow 10–30 minutes). vLLM loads the weights, then Triton compiles and autotunes the Gated DeltaNet and attention kernels, and Quark JIT-compiles its MX emulation kernel with hipcc. The cache mount keeps all of that for later starts. Setting
PYTORCH_ROCM_ARCHto your GPU's target makes Quark compile for that one target instead of the nine the image lists. - BF16 KV cache everywhere (
--kv-cache-dtype auto). The checkpoint ships no calibrated KV scales, so FP8 KV would run with a scale of 1.0 and change the outputs. FP8 KV is an opt-in (AEON_KV_DTYPE=fp8) for 16 GB cards only. VLLM_ROCM_USE_AITER=0on every recipe. On Radeon, AITER's unified attention overflows RDNA 4's 64 KB LDS. On MI350/MI355 the native MXFP4 GEMM is picked without it.--attention-backend TRITON_ATTNeverywhere, and inside--speculative-configtoo. The draft model doesn't inherit the target's backend; vLLM's automatic pick on RDNA 4 (ROCM_ATTN) collapses speculative acceptance under concurrency.--override-generation-configis required. The shippedgeneration_config.jsonsays temperature 1.0; the override sets the model card's 0.6 / 0.95 / 20.- Multi-GPU Radeon uses
NCCL_PROTO=SimpleandNCCL_P2P_DISABLE=1(the configuration users report working for TP=2 and TP=4 on R9700). - Ready when the log prints
Application startup complete.(docker logs -f aeon-mxfp4). Stop withdocker rm -f aeon-mxfp4.
The 2x R9700 configuration ladder
Try the rungs in order: A (recommended) → A+ (fast) → A-Max (max). If a TP=2 rung fails twice, use Z (never fails). A+DF is an A/B variant of A+ for short-context work. Every rung serves this checkpoint with its exact numerics (W4A4 / W6A6 emulation, BF16 elsewhere, BF16 KV). None of them has run on RDNA 4 hardware yet; confidence comes from reading the vLLM v0.31.0 source and from public reports.
| Rung | Script | Confidence | Context / sequences | Speed (estimate) |
|---|---|---|---|---|
| A (recommended) | serve_2x_r9700.sh |
medium-high | 131,072 / 16 | 4–9 tok/s per stream; aggregate grows almost linearly to ~16 streams |
| A+ (fast, MTP) | serve_2x_r9700_mtp.sh |
medium | 131,072 / 12 | ~2–2.5× A per stream |
| A+DF (DFlash2 variant) | serve_2x_r9700_dflash.sh |
medium-low | 65,536 / 8, no prefix caching | best on short single-shot work |
| A-Max (max) | serve_2x_r9700_max.sh |
low-medium | 65,536 / 4 | ~1.4–1.8× A+ per step |
| Z (never fails) | serve_2x_r9700_replicas.sh |
high (relative) | 32,768 / 4 per card | 2–5 tok/s per stream, two servers |
Recipe A: 2x Radeon AI PRO R9700 (recommended)
Target: 2x AMD Radeon AI PRO R9700 · VRAM: 2 × 32 GB · Arch: gfx1201 (RDNA 4) · MX path: emulated · Status: Expected · Script:
scripts/serve_2x_r9700.sh
bash "$AEON_DIR/scripts/serve_2x_r9700.sh"
Plain docker run for Recipe A
docker run -d --name aeon-mxfp4 \
--device /dev/kfd --device /dev/dri \
--group-add video --group-add "$RENDER_GID" \
--cap-add SYS_PTRACE --security-opt seccomp=unconfined --ipc=host \
-p 127.0.0.1:8000:8000 \
-v "$AEON_DIR":/model:ro -v "$AEON_CACHE":/root/.cache \
-e TRITON_CACHE_DIR=/root/.cache/triton -e TORCH_EXTENSIONS_DIR=/root/.cache/torch_extensions \
-e PYTORCH_ROCM_ARCH=gfx1201 -e VLLM_ROCM_USE_AITER=0 -e NCCL_PROTO=Simple -e NCCL_P2P_DISABLE=1 -e VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=1800 \
vllm/vllm-openai-rocm:v0.31.0 \
/model --served-model-name aeon --dtype bfloat16 \
--tensor-parallel-size 2 --attention-backend TRITON_ATTN --kv-cache-dtype auto \
--max-model-len 131072 --max-num-seqs 16 --max-num-batched-tokens 8192 \
--gpu-memory-utilization 0.90 --language-model-only \
--reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
--override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'
Why these settings:
--tensor-parallel-size 2splits every layer across both cards. Each card holds about 10.5 GiB of weights and has to dequantize only half of them on every forward pass. That is the main speed gain under emulation. All shard boundaries land on whole 32-element MX blocks (checked against the real shapes), and both cards carry 12 attention heads, 2 KV heads, 8 GDN key heads and 24 value heads.NCCL_PROTO=Simple,NCCL_P2P_DISABLE=1and--attention-backend TRITON_ATTNcome from the 2x R9700 TP=2 setups reported working (vllm#40980). vLLM's custom all-reduce is MI300/MI350-only, so Radeon cards talk through RCCL over PCIe.- No speculation, prefix caching on (vLLM's default for this model; it switches the Gated DeltaNet cache to "align" mode). This rung avoids every speculative-decoding issue, so it is the one most likely to work on the first try.
--max-model-len 131072 --max-num-seqs 16: estimated about 15–16 GiB per card left for cache. BF16 KV costs 32 KiB per token per card, so that is roughly 450K tokens in total: three full 131,072-token contexts, or many 20–60K agent contexts. Each running sequence also holds about 0.22 GiB of Gated DeltaNet state per card. For the full 262,144-token context useAEON_MAX_MODEL_LEN=262144 AEON_MAX_NUM_SEQS=4. The startup lineGPU KV cache size: N tokensis the real number.--language-model-onlyskips the vision tower (it fails to load on RDNA 4 today).--gpu-memory-utilization 0.90: use 0.85–0.88 if a card drives a display.
Recipe A+ Fast: 2x R9700 with MTP speculative decoding
Target: 2x AMD Radeon AI PRO R9700 · VRAM: 2 × 32 GB · Arch: gfx1201 (RDNA 4) · MX path: emulated; the MTP drafter runs in BF16 · Status: Untested · Script:
scripts/serve_2x_r9700_mtp.sh
bash "$AEON_DIR/scripts/serve_2x_r9700_mtp.sh"
Plain docker run for Recipe A+
docker run -d --name aeon-mxfp4 \
--device /dev/kfd --device /dev/dri \
--group-add video --group-add "$RENDER_GID" \
--cap-add SYS_PTRACE --security-opt seccomp=unconfined --ipc=host \
-p 127.0.0.1:8000:8000 \
-v "$AEON_DIR":/model:ro -v "$AEON_CACHE":/root/.cache \
-e TRITON_CACHE_DIR=/root/.cache/triton -e TORCH_EXTENSIONS_DIR=/root/.cache/torch_extensions \
-e PYTORCH_ROCM_ARCH=gfx1201 -e VLLM_ROCM_USE_AITER=0 -e NCCL_PROTO=Simple -e NCCL_P2P_DISABLE=1 -e VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=1800 -e GPU_MAX_HW_QUEUES=1 \
vllm/vllm-openai-rocm:v0.31.0 \
/model --served-model-name aeon --dtype bfloat16 \
--tensor-parallel-size 2 --attention-backend TRITON_ATTN --kv-cache-dtype auto \
--max-model-len 131072 --max-num-seqs 12 --max-num-batched-tokens 8192 \
--gpu-memory-utilization 0.90 --language-model-only \
--speculative-config '{"method":"mtp","num_speculative_tokens":3,"attention_backend":"TRITON_ATTN"}' \
--reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
--override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'
- Why it should help a lot here. Under emulation, every forward pass dequantizes all the MX weights no matter how many tokens it processes. MTP drafts 3 tokens with the model's own BF16 prediction head, and the main model checks all of them in one pass. Speculative decoding with the standard rejection sampler is lossless. MTP also keeps prefix caching, which matters most for multi-turn agents.
- Why it's untested.
config.jsonlists the MTP head's layers (mtp.fc,mtp.layers.0.*) as excluded from quantization, so vLLM builds them in BF16 to match the weights on disk. Reading the vLLM 0.31 source says the exclusion reaches the drafter, but nobody has run it yet, and the acceptance rate on this fine-tune is unmeasured (public Qwen3.x MTP-3 runs reach a mean acceptance length of about 2.2–2.8). - Pass check. The log shows
Detected MTP model. Sharing target model embedding weights with the draft model., and under loadSpecDecoding metrics: Mean acceptance length: ...of about 2 or more. A value near 1 under concurrency means the draft is not on TRITON_ATTN. - If it fails to load, or acceptance stays below about 1.5, go back to Recipe A and send us the log.
GPU_MAX_HW_QUEUES=1mitigates the R9700 "half speed for the whole process" behaviour (ROCm#6347). Remove it (AEON_HW_QUEUES=) if it measures slower.- Memory: each running sequence holds about 0.45 GiB of Gated DeltaNet state per card (one extra state slot per draft token), hence 12 sequences instead of 16.
- Tuning:
AEON_MTP=2orAEON_MTP=4to compare.
Recipe A+DF: 2x R9700 with the DFlash2 drafter (A/B variant)
Target: 2x AMD Radeon AI PRO R9700 · MX path: emulated; the DFlash2 drafter runs in BF16 · Status: Untested · Script:
scripts/serve_2x_r9700_dflash.sh· Needs: thedflash2/drafter folder (AEON-DFlash2-Qwen3.8-27B-BF16)
bash "$AEON_DIR/scripts/serve_2x_r9700_dflash.sh" # AEON_DFLASH=<dir> if the drafter lives elsewhere
Plain docker run for Recipe A+DF
docker run -d --name aeon-mxfp4 \
--device /dev/kfd --device /dev/dri \
--group-add video --group-add "$RENDER_GID" \
--cap-add SYS_PTRACE --security-opt seccomp=unconfined --ipc=host \
-p 127.0.0.1:8000:8000 \
-v "$AEON_DIR":/model:ro -v "$AEON_CACHE":/root/.cache \
-v "$AEON_DIR/dflash2":/draft:ro \
-e TRITON_CACHE_DIR=/root/.cache/triton -e TORCH_EXTENSIONS_DIR=/root/.cache/torch_extensions \
-e PYTORCH_ROCM_ARCH=gfx1201 -e VLLM_ROCM_USE_AITER=0 -e NCCL_PROTO=Simple -e NCCL_P2P_DISABLE=1 -e VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=1800 -e VLLM_USE_V2_MODEL_RUNNER=1 \
vllm/vllm-openai-rocm:v0.31.0 \
/model --served-model-name aeon --dtype bfloat16 \
--tensor-parallel-size 2 --attention-backend TRITON_ATTN --kv-cache-dtype auto \
--max-model-len 65536 --max-num-seqs 8 --max-num-batched-tokens 8192 \
--gpu-memory-utilization 0.90 --language-model-only \
--no-enable-prefix-caching \
--speculative-config '{"method":"dflash","model":"/draft","num_speculative_tokens":9,"attention_backend":"TRITON_ATTN","draft_sample_method":"greedy","rejection_sample_method":"standard"}' \
--reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
--override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'
- For short-context, single-shot work (long code or report generation). On our CUDA reference system this drafter reaches a mean acceptance length of about 3.3–3.6 at K=9 (lossless), higher than MTP is likely to.
- Not for long multi-turn tool loops. With this hybrid model, vLLM 0.29–0.31 needs prefix caching off for DFlash2 (otherwise it crashes or drops to 0% acceptance after a cache hit: vllm#55601, vllm#58894). Every turn then re-prefills the whole conversation, which is slow under emulation. Recipe A+ keeps prefix caching.
- Required settings:
VLLM_USE_V2_MODEL_RUNNER=1(otherwise it quietly runs as DFlash 1),"attention_backend":"TRITON_ATTN"in the speculative config, K=9 (the drafter was trained with block size 10),rejection_sample_method: "standard"(never"synthetic", which is not lossless). - Memory: about 1.8 GiB of drafter per card plus about 0.72 GiB of Gated DeltaNet state per sequence per card.
Recipe A-Max: 2x R9700, BF16-resident MLP (experimental)
Target: 2x AMD Radeon AI PRO R9700 · VRAM: 2 × 32 GB · Arch: gfx1201 (RDNA 4) · MX path: MLP dequantized once at load; MXFP6 emulated; MTP in BF16 · Status: Untested (experimental) · Script:
scripts/serve_2x_r9700_max.sh
vLLM 0.31 added VLLM_MXFP4_EMULATION_DEQUANT_AT_LOAD=1. It dequantizes the MXFP4 MLP weights to BF16 once, at load time, using the same function vLLM otherwise runs on every pass. The activations still go through MXFP4 quantize-dequantize, so the output is bit-identical to Recipe A+. That removes the largest per-pass cost of emulation (estimate: about 1.4–1.8× A+ per step). The cost is memory: the MLP becomes about 15.9 GiB of BF16 per card (about 22.6 GiB of weights per card), leaving about 4–5 GiB per card for cache.
bash "$AEON_DIR/scripts/serve_2x_r9700_max.sh"
Plain docker run for Recipe A-Max
docker run -d --name aeon-mxfp4 \
--device /dev/kfd --device /dev/dri \
--group-add video --group-add "$RENDER_GID" \
--cap-add SYS_PTRACE --security-opt seccomp=unconfined --ipc=host \
-p 127.0.0.1:8000:8000 \
-v "$AEON_DIR":/model:ro -v "$AEON_CACHE":/root/.cache \
-e TRITON_CACHE_DIR=/root/.cache/triton -e TORCH_EXTENSIONS_DIR=/root/.cache/torch_extensions \
-e PYTORCH_ROCM_ARCH=gfx1201 -e VLLM_ROCM_USE_AITER=0 -e NCCL_PROTO=Simple -e NCCL_P2P_DISABLE=1 -e VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=1800 -e GPU_MAX_HW_QUEUES=1 -e VLLM_MXFP4_EMULATION_DEQUANT_AT_LOAD=1 \
vllm/vllm-openai-rocm:v0.31.0 \
/model --served-model-name aeon --dtype bfloat16 \
--tensor-parallel-size 2 --attention-backend TRITON_ATTN --kv-cache-dtype auto \
--max-model-len 65536 --max-num-seqs 4 --max-num-batched-tokens 4096 \
--gpu-memory-utilization 0.92 --language-model-only \
--speculative-config '{"method":"mtp","num_speculative_tokens":3,"attention_backend":"TRITON_ATTN"}' \
--reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
--override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'
If it fails with an out-of-memory error or ... larger than the available KV cache memory, step the context down (AEON_MAX_MODEL_LEN=49152, then 32768); if it still fails, fall back to Recipe A+. Never set VLLM_MXFP4_EMULATION_DEQUANT_AT_LOAD on a single 32 GB card: the BF16 MLP alone is about 32 GiB.
Recipe Z: 2x R9700 as two independent servers (never-fails fallback)
Target: 2x AMD Radeon AI PRO R9700 · MX path: emulated · Status: Expected (fewest moving parts) · Script:
scripts/serve_2x_r9700_replicas.sh(stop with... replicas.sh stop)
Use this when a TP=2 recipe fails twice (RCCL hang, hipIpcGetMemHandle error, crash at graph capture). It runs one complete copy of the model per card, with no inter-GPU communication, no speculative decoding, no graph capture (--enforce-eager) and no AITER: aeon-mxfp4-0 on port 8000, aeon-mxfp4-1 on port 8001, and a small least-connections proxy (scripts/aeon_lb.py, or scripts/nginx_aeon_replicas.conf) on port 8080.
bash "$AEON_DIR/scripts/serve_2x_r9700_replicas.sh"
Plain docker run for Recipe Z
for i in 0 1; do
docker run -d --name aeon-mxfp4-$i \
--device /dev/kfd --device /dev/dri \
--group-add video --group-add "$RENDER_GID" \
--cap-add SYS_PTRACE --security-opt seccomp=unconfined --ipc=host \
-p 127.0.0.1:$((8000+i)):8000 \
-v "$AEON_DIR":/model:ro -v "$AEON_CACHE":/root/.cache \
-e TRITON_CACHE_DIR=/root/.cache/triton -e TORCH_EXTENSIONS_DIR=/root/.cache/torch_extensions \
-e PYTORCH_ROCM_ARCH=gfx1201 -e VLLM_ROCM_USE_AITER=0 -e HIP_VISIBLE_DEVICES=$i \
vllm/vllm-openai-rocm:v0.31.0 \
/model --served-model-name aeon --dtype bfloat16 \
--tensor-parallel-size 1 --attention-backend TRITON_ATTN --kv-cache-dtype auto --enforce-eager \
--max-model-len 32768 --max-num-seqs 4 --max-num-batched-tokens 4096 \
--gpu-memory-utilization 0.88 --language-model-only \
--reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
--override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'
[ "$i" = 0 ] && until docker logs aeon-mxfp4-0 2>&1 | grep -q "Application startup complete"; do sleep 15; done # GPU 1 reuses the kernel cache GPU 0 built
done
python3 "$AEON_DIR/scripts/aeon_lb.py" --listen 127.0.0.1:8080 127.0.0.1:8000 127.0.0.1:8001 &
- Cost (estimate): each card holds about 20.4 GiB of weights and about 6 GiB of cache: 32,768-token context, about 4 sequences per card, about 2–5 tok/s per stream.
- The script starts the second server only after the first is ready, so the one-time kernel build isn't run twice at the same time. Each container gets only its own GPU's device nodes.
- Once it works, drop
--enforce-eagerfirst (AEON_EAGER=0): it is the cheapest speed gain. - Don't use
--data-parallel-sizefor this: it still sets up inter-GPU process groups.
Recipe B: 1x Radeon AI PRO R9700
Target: 1x AMD Radeon AI PRO R9700 · VRAM: 32 GB · Arch: gfx1201 (RDNA 4) · MX path: emulated · Status: Expected · Script:
scripts/serve_1x_r9700.sh
bash "$AEON_DIR/scripts/serve_1x_r9700.sh"
Plain docker run for Recipe B
docker run -d --name aeon-mxfp4 \
--device /dev/kfd --device /dev/dri \
--group-add video --group-add "$RENDER_GID" \
--cap-add SYS_PTRACE --security-opt seccomp=unconfined --ipc=host \
-p 127.0.0.1:8000:8000 \
-v "$AEON_DIR":/model:ro -v "$AEON_CACHE":/root/.cache \
-e TRITON_CACHE_DIR=/root/.cache/triton -e TORCH_EXTENSIONS_DIR=/root/.cache/torch_extensions \
-e PYTORCH_ROCM_ARCH=gfx1201 -e VLLM_ROCM_USE_AITER=0 \
vllm/vllm-openai-rocm:v0.31.0 \
/model --served-model-name aeon --dtype bfloat16 \
--tensor-parallel-size 1 --attention-backend TRITON_ATTN --kv-cache-dtype auto \
--max-model-len 32768 --max-num-seqs 4 --max-num-batched-tokens 4096 \
--gpu-memory-utilization 0.90 --language-model-only \
--reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
--override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'
- Memory (estimate): about 20.4 GiB of weights leaves about 6 GiB for cache: BF16 KV at 64 KiB per token plus about 0.45 GiB of Gated DeltaNet state per running sequence (prefix caching on).
- Longer context:
AEON_MAX_MODEL_LEN=65536 AEON_MAX_NUM_SEQS=2. - MTP (untested):
AEON_MTP=3 AEON_MAX_NUM_SEQS=2. - Most robust:
AEON_EAGER=1 AEON_MAX_MODEL_LEN=16384 AEON_MAX_NUM_SEQS=2 AEON_GPU_UTIL=0.88. - Dequant-at-load and DFlash2 don't fit on one 32 GB card.
Recipe C: 2x Radeon RX 9070 XT
Target: 2x AMD Radeon RX 9070 XT or RX 9070 · VRAM: 2 × 16 GB · Arch: gfx1201 (RDNA 4) · MX path: emulated · Status: Expected, tight memory · Script:
scripts/serve_2x_rx9070xt.sh
bash "$AEON_DIR/scripts/serve_2x_rx9070xt.sh"
Plain docker run for Recipe C
docker run -d --name aeon-mxfp4 \
--device /dev/kfd --device /dev/dri \
--group-add video --group-add "$RENDER_GID" \
--cap-add SYS_PTRACE --security-opt seccomp=unconfined --ipc=host \
-p 127.0.0.1:8000:8000 \
-v "$AEON_DIR":/model:ro -v "$AEON_CACHE":/root/.cache \
-e TRITON_CACHE_DIR=/root/.cache/triton -e TORCH_EXTENSIONS_DIR=/root/.cache/torch_extensions \
-e PYTORCH_ROCM_ARCH=gfx1201 -e VLLM_ROCM_USE_AITER=0 -e NCCL_PROTO=Simple -e NCCL_P2P_DISABLE=1 -e VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=1800 \
vllm/vllm-openai-rocm:v0.31.0 \
/model --served-model-name aeon --dtype bfloat16 \
--tensor-parallel-size 2 --attention-backend TRITON_ATTN --kv-cache-dtype auto \
--max-model-len 16384 --max-num-seqs 2 --max-num-batched-tokens 2048 \
--gpu-memory-utilization 0.92 --language-model-only \
--reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
--override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'
Same code path as Recipe A with half the memory: about 10.5 GiB of weights per card leaves about 2 GiB per card for cache. If a card drives a display, use --gpu-memory-utilization 0.88. MTP (AEON_MTP=2) only if the GPU KV cache size line shows room. FP8 KV (AEON_KV_DTYPE=fp8) doubles the context but changes outputs (uncalibrated scales); it is opt-in. There is no single-card fallback: the model doesn't fit on 16 GB.
Recipe D: 4x Radeon RX 9070 XT
Target: 4x AMD Radeon RX 9070 XT or RX 9070 · VRAM: 4 × 16 GB · Arch: gfx1201 (RDNA 4) · MX path: emulated · Status: Untested · Script:
scripts/serve_4x_rx9070xt.sh
bash "$AEON_DIR/scripts/serve_4x_rx9070xt.sh"
Plain docker run for Recipe D
docker run -d --name aeon-mxfp4 \
--device /dev/kfd --device /dev/dri \
--group-add video --group-add "$RENDER_GID" \
--cap-add SYS_PTRACE --security-opt seccomp=unconfined --ipc=host \
-p 127.0.0.1:8000:8000 \
-v "$AEON_DIR":/model:ro -v "$AEON_CACHE":/root/.cache \
-e TRITON_CACHE_DIR=/root/.cache/triton -e TORCH_EXTENSIONS_DIR=/root/.cache/torch_extensions \
-e PYTORCH_ROCM_ARCH=gfx1201 -e VLLM_ROCM_USE_AITER=0 -e NCCL_PROTO=Simple -e NCCL_P2P_DISABLE=1 -e VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=1800 \
vllm/vllm-openai-rocm:v0.31.0 \
/model --served-model-name aeon --dtype bfloat16 \
--tensor-parallel-size 4 --attention-backend TRITON_ATTN --kv-cache-dtype auto \
--max-model-len 131072 --max-num-seqs 8 --max-num-batched-tokens 8192 \
--gpu-memory-utilization 0.90 --language-model-only \
--reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
--override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'
About 5.3 GiB of weights per card leaves about 8 GiB per card; BF16 KV costs 16 KiB per token per card. TP=4 shards check out (6 attention heads and 1 KV head per card; every MX shard on a whole block), and TP=4 with NCCL_PROTO=Simple is reported working on 4x R9700. Speed will likely be limited by all-reduce over consumer PCIe. Optional: MTP (AEON_MTP=3). If startup hangs at graph capture, AEON_EAGER=1. Don't use 8 cards with TP=8: 4 KV heads can't be split 8 ways.
Recipe E: 2x Radeon RX 7900 XTX
Target: 2x AMD Radeon RX 7900 XTX · VRAM: 2 × 24 GB · Arch: gfx1100 (RDNA 3) · MX path: emulated · Status: Untested · Script:
scripts/serve_2x_rx7900xtx.sh
bash "$AEON_DIR/scripts/serve_2x_rx7900xtx.sh"
Plain docker run for Recipe E
docker run -d --name aeon-mxfp4 \
--device /dev/kfd --device /dev/dri \
--group-add video --group-add "$RENDER_GID" \
--cap-add SYS_PTRACE --security-opt seccomp=unconfined --ipc=host \
-p 127.0.0.1:8000:8000 \
-v "$AEON_DIR":/model:ro -v "$AEON_CACHE":/root/.cache \
-e TRITON_CACHE_DIR=/root/.cache/triton -e TORCH_EXTENSIONS_DIR=/root/.cache/torch_extensions \
-e PYTORCH_ROCM_ARCH=gfx1100 -e VLLM_ROCM_USE_AITER=0 -e NCCL_PROTO=Simple -e NCCL_P2P_DISABLE=1 -e VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=1800 \
vllm/vllm-openai-rocm:v0.31.0 \
/model --served-model-name aeon --dtype bfloat16 \
--tensor-parallel-size 2 --attention-backend TRITON_ATTN --kv-cache-dtype auto \
--max-model-len 65536 --max-num-seqs 8 --max-num-batched-tokens 8192 \
--gpu-memory-utilization 0.90 --language-model-only \
--reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
--override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'
RDNA 3 has no FP8 in vLLM, so the KV cache stays BF16 at 32 KiB per token per card. About 10 GiB per card is left: about 300K tokens in total, enough for 65,536 × 8, or 131,072 × 2 (AEON_MAX_MODEL_LEN=131072 AEON_MAX_NUM_SEQS=2). MTP: AEON_MTP=3. For 2x RX 7900 XT (20 GB) use AEON_MAX_MODEL_LEN=32768 AEON_MAX_NUM_SEQS=4.
Recipe F: Radeon PRO W7900 (48 GB)
Target: 1x AMD Radeon PRO W7900, or W7800 48 GB · VRAM: 48 GB · Arch: gfx1100 (RDNA 3) · MX path: emulated · Status: Untested · Script:
scripts/serve_w7900.sh
bash "$AEON_DIR/scripts/serve_w7900.sh"
Plain docker run for Recipe F
docker run -d --name aeon-mxfp4 \
--device /dev/kfd --device /dev/dri \
--group-add video --group-add "$RENDER_GID" \
--cap-add SYS_PTRACE --security-opt seccomp=unconfined --ipc=host \
-p 127.0.0.1:8000:8000 \
-v "$AEON_DIR":/model:ro -v "$AEON_CACHE":/root/.cache \
-e TRITON_CACHE_DIR=/root/.cache/triton -e TORCH_EXTENSIONS_DIR=/root/.cache/torch_extensions \
-e PYTORCH_ROCM_ARCH=gfx1100 -e VLLM_ROCM_USE_AITER=0 \
vllm/vllm-openai-rocm:v0.31.0 \
/model --served-model-name aeon --dtype bfloat16 \
--tensor-parallel-size 1 --attention-backend TRITON_ATTN --kv-cache-dtype auto \
--max-model-len 131072 --max-num-seqs 4 --max-num-batched-tokens 8192 \
--gpu-memory-utilization 0.90 --language-model-only \
--reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
--override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'
BF16 KV at 64 KiB per token; about 20 GiB is left for cache. A single 262,144-token stream fits (AEON_MAX_MODEL_LEN=262144 AEON_MAX_NUM_SEQS=1). MTP (AEON_MTP=3), the DFlash2 variant (AEON_DFLASH="$AEON_DIR/dflash2" AEON_MAX_MODEL_LEN=65536) and vision (AEON_VISION=1) all fit.
Recipe G: Radeon PRO W7800 (32 GB)
Target: 1x AMD Radeon PRO W7800 32 GB · VRAM: 32 GB · Arch: gfx1100 (RDNA 3) · MX path: emulated · Status: Untested · Script:
scripts/serve_w7800_32gb.sh
bash "$AEON_DIR/scripts/serve_w7800_32gb.sh"
Plain docker run for Recipe G
docker run -d --name aeon-mxfp4 \
--device /dev/kfd --device /dev/dri \
--group-add video --group-add "$RENDER_GID" \
--cap-add SYS_PTRACE --security-opt seccomp=unconfined --ipc=host \
-p 127.0.0.1:8000:8000 \
-v "$AEON_DIR":/model:ro -v "$AEON_CACHE":/root/.cache \
-e TRITON_CACHE_DIR=/root/.cache/triton -e TORCH_EXTENSIONS_DIR=/root/.cache/torch_extensions \
-e PYTORCH_ROCM_ARCH=gfx1100 -e VLLM_ROCM_USE_AITER=0 \
vllm/vllm-openai-rocm:v0.31.0 \
/model --served-model-name aeon --dtype bfloat16 \
--tensor-parallel-size 1 --attention-backend TRITON_ATTN --kv-cache-dtype auto \
--max-model-len 32768 --max-num-seqs 4 --max-num-batched-tokens 4096 \
--gpu-memory-utilization 0.90 --language-model-only \
--reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
--override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'
Like Recipe B on RDNA 3: about 6 GiB left for BF16 KV and Gated DeltaNet state. MTP: AEON_MTP=3 AEON_MAX_NUM_SEQS=2. Text-only.
Recipe H: Strix Halo (Ryzen AI Max+ 395)
Target: AMD Ryzen AI Max / Max+ 395 APU (Radeon 8060S) · Memory: 64 or 128 GB unified · Arch: gfx1151 (RDNA 3.5) · MX path: emulated · Status: Untested · Script:
scripts/serve_strix_halo.sh
Before you start, the GPU must be able to allocate at least about 40 GB. vLLM sizes its memory from what HIP reports, so raise the TTM/GTT limit the way AMD documents it for Strix Halo (amd-ttm, or the ttm pages_limit module parameter; amdgpu.gttsize is deprecated), or raise the UMA frame buffer in the BIOS. scripts/detect_setup.py prints both values.
bash "$AEON_DIR/scripts/serve_strix_halo.sh"
Plain docker run for Recipe H
docker run -d --name aeon-mxfp4 \
--device /dev/kfd --device /dev/dri \
--group-add video --group-add "$RENDER_GID" \
--cap-add SYS_PTRACE --security-opt seccomp=unconfined --ipc=host \
-p 127.0.0.1:8000:8000 \
-v "$AEON_DIR":/model:ro -v "$AEON_CACHE":/root/.cache \
-e TRITON_CACHE_DIR=/root/.cache/triton -e TORCH_EXTENSIONS_DIR=/root/.cache/torch_extensions \
-e PYTORCH_ROCM_ARCH=gfx1151 -e VLLM_ROCM_USE_AITER=0 \
vllm/vllm-openai-rocm:v0.31.0 \
/model --served-model-name aeon --dtype bfloat16 \
--tensor-parallel-size 1 --attention-backend TRITON_ATTN --kv-cache-dtype auto \
--max-model-len 65536 --max-num-seqs 2 --max-num-batched-tokens 4096 \
--gpu-memory-utilization 0.85 --language-model-only \
--reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
--override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'
BF16 KV (no FP8 on gfx1151). About 256 GB/s of memory bandwidth makes emulated decode slow, so speculation (AEON_MTP=3) is the main lever; on 128 GB systems AEON_DEQUANT_AT_LOAD=1 AEON_MTP=3 also fits and roughly halves MLP traffic (bit-identical). Most robust: AEON_EAGER=1 AEON_MAX_MODEL_LEN=32768 AEON_MAX_NUM_SEQS=1.
Recipe I: Instinct MI300X / MI325X
Target: AMD Instinct MI300X or MI325X · VRAM: 192 / 256 GB · Arch: gfx942 (CDNA 3) · MX path: emulated (no MX hardware) · Status: Untested · Script:
scripts/serve_mi300x.sh
bash "$AEON_DIR/scripts/serve_mi300x.sh"
Plain docker run for Recipe I
docker run -d --name aeon-mxfp4 \
--device /dev/kfd --device /dev/dri \
--group-add video --group-add "$RENDER_GID" \
--cap-add SYS_PTRACE --security-opt seccomp=unconfined --ipc=host \
-p 127.0.0.1:8000:8000 \
-v "$AEON_DIR":/model:ro -v "$AEON_CACHE":/root/.cache \
-e TRITON_CACHE_DIR=/root/.cache/triton -e TORCH_EXTENSIONS_DIR=/root/.cache/torch_extensions \
-e PYTORCH_ROCM_ARCH=gfx942 -e VLLM_ROCM_USE_AITER=0 \
vllm/vllm-openai-rocm:v0.31.0 \
/model --served-model-name aeon --dtype bfloat16 \
--tensor-parallel-size 1 --attention-backend TRITON_ATTN --kv-cache-dtype auto \
--max-model-len 262144 --max-num-seqs 32 \
--gpu-memory-utilization 0.90 --language-model-only \
--reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
--override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'
- Memory is not a constraint, so use one GPU per copy and scale out with data parallelism (
--data-parallel-size 8, orAEON_DP=8) instead of tensor parallelism. - Faster:
AEON_DEQUANT_AT_LOAD=1 AEON_MTP=3keeps the MLP in BF16 after loading (+23.7 GiB, bit-identical) and adds MTP. For short-context throughput,AEON_DEQUANT_AT_LOAD=1 AEON_DFLASH="$AEON_DIR/dflash2"(prefix caching off); keep MTP for multi-turn agents. - Images:
AEON_VISION=1(4 images per prompt). - An FP8 or BF16 build is the better everyday choice on MI300, since memory isn't the limit there.
Recipe J: Instinct MI350X / MI355X
Target: AMD Instinct MI350X or MI355X · VRAM: 288 GB · Arch: gfx950 (CDNA 4) · MX path: MXFP4 native (AITER), MXFP6 emulated · Status: Untested · Script:
scripts/serve_mi355x.sh
bash "$AEON_DIR/scripts/serve_mi355x.sh"
Plain docker run for Recipe J
docker run -d --name aeon-mxfp4 \
--device /dev/kfd --device /dev/dri \
--group-add video --group-add "$RENDER_GID" \
--cap-add SYS_PTRACE --security-opt seccomp=unconfined --ipc=host \
-p 127.0.0.1:8000:8000 \
-v "$AEON_DIR":/model:ro -v "$AEON_CACHE":/root/.cache \
-e TRITON_CACHE_DIR=/root/.cache/triton -e TORCH_EXTENSIONS_DIR=/root/.cache/torch_extensions \
-e PYTORCH_ROCM_ARCH=gfx950 -e VLLM_ROCM_USE_AITER=0 \
vllm/vllm-openai-rocm:v0.31.0 \
/model --served-model-name aeon --dtype bfloat16 \
--tensor-parallel-size 1 --attention-backend TRITON_ATTN --kv-cache-dtype auto \
--max-model-len 262144 --max-num-seqs 32 \
--gpu-memory-utilization 0.90 --language-model-only \
--reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
--override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'
- Native MXFP4 needs no flag. vLLM picks
AiterMxfp4LinearKernel(AITER W4A4 GEMM) on its own when the GPU supports MX; it doesn't depend onVLLM_ROCM_USE_AITER. - Check the log for
Using AiterMxfp4LinearKernel for MXFP4 GEMMandUsing EmulationMxfp6LinearKernel for MXFP6 GEMM. The MXFP6 warning is expected. - Native is not bit-identical to emulation. For outputs that match the other recipes bit for bit,
AEON_MXFP4_EMULATE=1(setsVLLM_DISABLED_KERNELS=AiterMxfp4LinearKernel). Use the same switch if loading fails with aModuleNotFoundErrorforaiter.ops.triton.gemm_afp4wfp4(newer AITER builds moved that module). - Faster: DFlash2 (
AEON_DFLASH="$AEON_DIR/dflash2") or MTP (AEON_MTP=3). Scaling out:AEON_DP=N.
Not supported
- 1x RX 7900 XTX (24 GB), 1x RX 9070 XT / 9070 (16 GB), and any single GPU under 32 GB: about 20.4 GiB of text-only weights leaves no usable room for KV cache, or doesn't fit at all. Use two or four cards (Recipes C, D and E).
- 8-GPU tensor parallelism: the model has 4 KV heads, so TP=8 can't split them evenly. Use TP=1 or TP=2 replicas with data parallelism.
- Virtual machines with GPU passthrough on Radeon: RCCL initialization has been reported to hang. Use bare metal, or Recipe Z.
- llama.cpp, Ollama, LM Studio: this is not a GGUF.
Vision (Instinct and RDNA 3 only)
Every recipe starts text-only, which leaves the most memory for context. On RDNA 4 the vision tower currently fails to load in vLLM (vllm#49851), so keep --language-model-only there. Elsewhere, to serve images:
- Set
AEON_VISION=1with the script, or edit thedocker run:- remove
--language-model-only; - add
--limit-mm-per-prompt '{"image":1,"video":0}' --mm-processor-kwargs '{"max_pixels":1048576}'; - on Radeon / Strix Halo, add
-e FLASH_ATTENTION_TRITON_AMD_ENABLE=TRUEbefore the image name.
- remove
- The vision tower adds about 0.9 GB, split across the cards.
The image-size cap matters. Without it, vLLM sizes the encoder's startup memory test for this processor's maximum of about 16.7 megapixels (about 16K tokens per image), which can run out of memory. With the cap it plans for about 1,024 tokens per image. Vision is untested with this checkpoint.
Troubleshooting
The full symptom → cause → fix table is in AGENTS.md. The most common ones:
| Symptom | Likely cause | Fix |
|---|---|---|
AttributeError: 'dict' object has no attribute 'endswith' in quark.py |
vLLM image v0.28–v0.30 | Use vllm/vllm-openai-rocm:v0.31.0 |
The launch script stops with vLLM detects 'gfx1036' (or another iGPU target) |
An integrated GPU is listed first, so vLLM takes its architecture | AEON_DEVICES=auto, or pass only the discrete /dev/dri/renderD* and card* nodes, or disable the iGPU in the BIOS. HIP_VISIBLE_DEVICES alone doesn't help |
| Startup or the first request hangs with both GPUs at 100%, or responses are empty (TP ≥ 2) | RCCL over PCIe on gfx12 (vllm#40980) | Check NCCL_PROTO=Simple and NCCL_P2P_DISABLE=1 are set; both cards in equal-width CPU slots; IOMMU on with iommu=pt; bare metal. Still stuck: Recipe Z |
hipIpcGetMemHandle ... invalid argument |
P2P IPC between the cards | NCCL_P2P_DISABLE=1, then Recipe Z |
| Stuck in compilation for over 30 minutes, one CPU core busy | Triton autotuning the Gated DeltaNet kernels on first start | Wait, then retry with AEON_EXTRA_ENV="VLLM_TRITON_FORCE_FIRST_CONFIG=1" |
out of resource: shared memory |
AITER attention on RDNA | VLLM_ROCM_USE_AITER=0 (the scripts set it) |
... larger than the available KV cache memory or No available memory for the cache blocks |
Not enough VRAM left for KV | Lower --max-model-len, then --max-num-seqs; keep --language-model-only; close other GPU apps |
| Speculative acceptance near 1 with several requests in flight | Draft model on the ROCM_ATTN backend | "attention_backend":"TRITON_ATTN" inside --speculative-config |
| DFlash2 crashes or acceptance drops to 0 after a repeated prompt | Prefix caching with this hybrid model | --no-enable-prefix-caching (Recipe A+DF sets it) |
| Endless thinking or repetition | Client sends temperature: 0 |
Send 0.6 or omit it; cap max_tokens; enable_thinking: false for tool loops |
| Decode speed about half of a previous run | R9700 speed set at process start (ROCm#6347) | GPU_MAX_HW_QUEUES=1, restart and measure again |
| Tool calls come back as plain text | Tool parser not active | --enable-auto-tool-choice --tool-call-parser qwen3_coder; the client must send tools |
Expected performance (estimates, not measurements)
No one has measured this checkpoint on AMD hardware yet. The table below gives upper bounds worked out from memory bandwidth, so you know the order of magnitude and why the recipes look the way they do. Real numbers will be lower: our working estimate for Recipe A is 4–9 tok/s at one request. Measured results will replace this section.
| Recipe | Bytes moved per forward pass, per GPU | Bandwidth per GPU (spec) | c=1 decode ceiling | Ceiling with MTP k=3 |
|---|---|---|---|---|
| A: 2x R9700 | 55–77 GB | 640 GB/s | 8–12 tok/s | 20–32 tok/s |
| A-Max: 2x R9700, BF16 MLP | 33–56 GB | 640 GB/s | 12–19 tok/s | 27–49 tok/s |
| B: 1x R9700 | 110–154 GB | 640 GB/s | 4–6 tok/s | 10–16 tok/s |
| C: 2x RX 9070 XT | 55–77 GB | 640 GB/s | 8–12 tok/s | (MTP not advised) |
| D: 4x RX 9070 XT | 28–39 GB | 640 GB/s | 17–23 tok/s | 39–63 tok/s |
| E: 2x RX 7900 XTX | 55–77 GB | 960 GB/s | 12–17 tok/s | 30–47 tok/s |
| F: W7900 | 110–154 GB | 864 GB/s | 6–8 tok/s | 13–21 tok/s |
| G: W7800 32 GB | 110–154 GB | 576 GB/s | 4–5 tok/s | 9–14 tok/s |
| H: Strix Halo | 110–154 GB | 256 GB/s | about 2 tok/s | 4–6 tok/s |
How these numbers are worked out:
- MXFP4 MLP, 17.11B weights. Every forward pass reads the packed weights (0.53 B per weight including scales), writes a BF16 copy (2 B) and reads it back for the matmul (2 B): about 77.5 GB.
- MXFP6 GDN projections, 5.54B weights. The same pattern gives about 26.5 GB. The MXFP6 unpack is plain PyTorch with int32 and FP32 temporaries; up to 8 more bytes per weight (about 44 GB) is the top of the range.
- BF16 attention and
lm_head: about 6 GB. - Total: 110–154 GB per pass at TP=1, divided by the number of GPUs. The ceiling is bandwidth ÷ bytes per GPU, so it ignores all-reduce latency (more than 128 per pass over PCIe), kernel-launch gaps and compute.
- MTP. With k draft tokens and a per-token acceptance rate α, each pass yields (1 − α^(k+1)) / (1 − α) tokens: 2.53 at α = 0.7 and 2.95 at α = 0.8 for k = 3. Each draft token also reads the MTP layer and
lm_head(3.4 GB ÷ TP). The acceptance rate for this quantized, uncensored model is unknown until someone runs it. - A-Max. The MLP is read once in BF16 (34.2 GB) instead of 77.5 GB.
- Concurrency. The weight traffic is per pass, not per sequence, so aggregate throughput at 4, 8 or 16 concurrent requests should rise close to linearly until compute runs out. This is an expectation, not a measurement.
- Instinct. At Instinct bandwidths (5.3 TB/s for MI300X, 8 TB/s for MI355X) launch overhead, not bandwidth, likely sets the limit, so no ceiling is given.
For scale only: Puget Systems reports 15.9 tok/s stock and 62.8 tok/s tuned with MTP at concurrency 1 on 2x R9700 with Qwen3.6-27B in FP8 (article). That is a different model and a different format with no emulation.
Testers: the step-by-step tester kit ships with the files, in TESTER_GUIDE.md, tests/run_radeon_eval.py and scripts/ (start with scripts/preflight_smoke.sh). Agents: AGENTS.md.
Expected behavior and limits
- Speed. Under emulation, decode speed is limited by how many bytes each forward pass moves, not by the 23.7 GB of packed weights. See Expected performance. Tensor parallelism, MTP and concurrent requests all spread that cost. Prefill is affected much less. On MI350/MI355 the MLP runs natively and only the MXFP6 projections are emulated.
- Memory.
- About 22.8 GB of weights load with the vision tower, or about 21.9 GB (20.4 GiB) with
--language-model-only. The MTP head (0.85 GB) loads only when you turn on MTP. - The KV cache costs 64 KiB per token at BF16 (16 full-attention layers × 4 KV heads × 256 × 2 × 2 bytes), divided by the TP size. Every recipe uses BF16 KV, the reference numerics; FP8 KV (32 KiB) is an opt-in that changes outputs.
- Each running sequence also holds Gated DeltaNet state: about 147 MiB per state slot, FP32 as the model config sets it, divided by the TP size. With prefix caching on (vLLM's default for this model) a sequence can hold two slots, plus one more per MTP draft token.
- The
GPU KV cache sizeline in the startup log is the real capacity.
- About 22.8 GB of weights load with the vision tower, or about 21.9 GB (20.4 GiB) with
- Speculative decoding is untested on AMD. MTP: see Recipe A+; use
{"method":"mtp",...}(the older nameqwen3_5_mtpstill works but logs a deprecation warning) and always add"attention_backend":"TRITON_ATTN". DFlash2: see Recipe A+DF. - Thinking.
- Thinking is on by default, and the chat template's default reasoning effort is
xhigh. - For quick answers, send
"chat_template_kwargs": {"enable_thinking": false}. - For short thinking, send
{"enable_thinking": true, "reasoning_effort": "low"}. - Set both through
chat_template_kwargs, not through a top-level request field.
- Thinking is on by default, and the chat template's default reasoning effort is
- Inherited behavior. The BF16 master is an Early Access Draft, and very long answers can fall into loops. 4-bit MLP weights, including the residual-writing
down_proj, may make that more likely. The tester kit checks for loops explicitly. - Uncensored. This model writes what the base model refuses. Read User responsibility.
- Runtimes.
- vLLM is the only supported path.
- This is not a GGUF, so llama.cpp, Ollama and LM Studio can't load it.
- Transformers with
amd-quarkmay load it in emulation, but that is untested. - Apple needs an MLX conversion.
- Known ROCm issues.
- Decode speed on the R9700 can differ by about 20% from one process start to the next (ROCm#6347).
- A TP=2 hang on dual R9700 is still open (vllm#40980); Recipe A uses the settings reported to avoid it (
NCCL_PROTO=Simple,NCCL_P2P_DISABLE=1), and Recipe Z avoids inter-GPU communication entirely. - vLLM v0.28–v0.30 can't load this checkpoint (
algo_configcrash in the Quark loader); use v0.31.0. - With this hybrid model, DFlash2 speculative decoding needs prefix caching off on vLLM 0.29–0.31 (vllm#55601, vllm#58894).
- The first-start Gated DeltaNet compile hang on RDNA 4 (vllm#45929) is fixed in the images these recipes use.
Quality and validation status
| Check | Where | Status |
|---|---|---|
| Export integrity: 336 MX layers written; all 15 MTP tensors and 333 vision tensors carried over in BF16 | Quantization run | Done |
vLLM code path: Quark OCP MX loader, native or emulated kernel selection on each platform, TP=2/TP=4 MX sharding, MTP exclusion, algo_config handling (crashes on ≤ 0.30, fixed in 0.31) |
vLLM 0.29.0 and 0.31.0 source review | Done (a code review, not a runtime test) |
DFlash2 drafter (dflash2/): mean acceptance length at K=9 |
CUDA reference system | about 3.3–3.6 (not yet measured on AMD) |
| Apple Silicon, via an MLX conversion | Mac | Pending |
| W4A4 / W6A6 quality against the BF16 master, using emulated MX numerics | GPU, vLLM emulation | Pending |
| 2x Radeon AI PRO R9700, Recipes A, A+, A+DF, A-Max and Z: smoke tests, GSM8K-50, IFEval-20, 10 tool calls, throughput at c=1/4/8, TTFT, VRAM per GPU, speculative acceptance | ROCm 7.2.3, vLLM 0.31.0 | Pending (first tester run in progress) |
| Other Radeon systems (Recipes B–H) | n/a | Not yet tested |
| Instinct MI300X / MI325X (emulated) and MI350X / MI355X (native MXFP4) | n/a | Not yet tested |
Quantization recipe
| Module group | Linear layers | Format | Weights | Activations | Approx. size |
|---|---|---|---|---|---|
MLP gate_proj / up_proj / down_proj |
192 (64 layers × 3) | MXFP4 | FP4 E2M1, 32-element blocks, E8M0 scale | MXFP4, dynamic per 32-element block | 9.1 GB |
GDN in_proj_qkv / in_proj_z / out_proj |
144 (48 layers × 3) | MXFP6 E2M3 | FP6 E2M3, 32-element blocks, E8M0 scale | MXFP6 E2M3, dynamic | 4.3 GB |
Full attention q/k/v/o_proj, plus q_norm/k_norm |
64 (16 layers × 4) | BF16 | n/a | n/a | 3.4 GB |
GDN recurrence: in_proj_a/b, conv1d, A_log, dt_bias, norm |
96 linears + params | BF16 | n/a | n/a | 0.05 GB |
embed_tokens and lm_head (untied) |
n/a | BF16 | n/a | n/a | 5.1 GB |
| Vision tower (27 blocks + merger) | n/a | BF16 | n/a | n/a | 0.9 GB |
MTP head (1 layer + fc) |
n/a | BF16 | n/a | n/a | 0.85 GB |
- Tool: AMD Quark 0.12.post1 in eager mode, with Transformers 5.14.1 and PyTorch 2.13. The export is
real_quantized: packeduint8weights plusuint8E8M0 scales. The Quark MX settings arescale_calculation_mode: evenandround_method: half_even. - Algorithm: AutoSmoothQuant with an MSE scale search, applied to the MLP only, along the edges
post_attention_layernorm → gate/upandup_proj → down_proj. The smoothing factors are folded into those norms and weights. The GDN projections have no safe smoothing edge, so they get none. - Calibration: 1,024 chat, code and math samples (UltraChat, Open-Platypus, CodeAlpaca, GSM8K), each truncated to 64 tokens, for 65,536 tokens in total. To fit in memory, the AutoSmoothQuant scale search used a 64-token subsample per layer. MX weight scales come from the weights themselves and activations are quantized at runtime, so calibration data only affects the MLP smoothing factors.
- Exact config: see
quantization_configinconfig.jsonandBAKE_STAMP_Q2.json.
Model family
| Variant | Format | Size | Best hardware | Status | Link |
|---|---|---|---|---|---|
| BF16 master | BF16, with vision and MTP | 55.6 GB | H200, multi-GPU, RTX PRO 6000 | Public | AEON-7/…-BF16 |
| NVFP4-MIXED | NVFP4 MLP (layers 0–55); FP8 attention, GDN writers and MLP layers 56–63; BF16 for the rest (ModelOpt) | 24.7 GB | DGX Spark (GB10), RTX 5090, RTX PRO 6000 | Public | AEON-7/…-NVFP4-MIXED |
| MXFP4-MXFP6-ROCm (this repo) | MXFP4 MLP and MXFP6 GDN projections; BF16 for the rest (AMD Quark) | 23.7 GB | Instinct MI350/MI355 (native MXFP4); Radeon AI PRO R9700 32 GB (emulated) | Early Release · Gated Access | AEON-7/…-MXFP4-MXFP6-ROCm |
| DFlash 2 drafter for NVFP4-MIXED | Speculative-decoding drafter | TBA | DGX Spark, RTX 50-series | Early access, coming soon | n/a |
| NVFP4-GDNFP8 | NVFP4 MLP with FP8 GDN projections (LLM Compressor) | TBA | NVIDIA Blackwell | Early access (Patreon); not public | n/a |
| FP8-MIXED | FP8 dynamic MLP; BF16 for the rest (LLM Compressor) | 38.5 GB | FP8-capable GPUs with 48 GB or more | Internal: built, not yet validated | n/a |
| MLX (Apple Silicon) | MLX conversion | TBA | Apple Silicon Macs | Planned | n/a |
User responsibility
By accessing, downloading or running this model, you agree to the following:
- You are responsible for its use. You alone are responsible for every prompt, every response, every downstream action, and any harm that results.
- No warranty. The model is provided "AS IS", without warranty of any kind.
- Follow the law. You must comply with all applicable laws and policies in every jurisdiction you operate in.
- Add safety layers in production. Use input validation, output filtering, access controls, and human review for high-risk workflows.
- The duty of care is yours. An uncensored model doesn't refuse on your behalf, so if you are unsure about a request, don't make it.
- No endorsement. The authors do not endorse any particular output.
- Arbitration. Disputes go to binding individual arbitration (under the AAA Consumer Rules if no other body is agreed), waiving jury trials and class actions.
- Indemnification. You indemnify the authors, contributors and publishers against claims arising from your use.
- Severability. If a provision is invalid, the closest enforceable equivalent replaces it.
- Acceptance. Using the model means you accept these terms. If you don't accept them, don't use the model.
License and attribution
- License: Apache-2.0, inherited from Qwen/Qwen3.8-27B.
- Qwen team (Alibaba): the Qwen3.8-27B base model.
- AEON-7: the uncensored fine-tune (the BF16 master) and this quantization. The upstream tools behind the fine-tune are credited on the BF16 card.
- AMD Quark: the quantization toolkit used to produce the OCP MX export.