Qwen3.8 Flash Next - Uncensored NVFP4 (calibrated, ours)
Calibrated W4A4 NVFP4: AWQ+GPTQ on routed experts, FP8 attention/shared/lm_head,
BF16 keep-list (incl. PLE table), static FP8 KV scheme (dropped at serve time:
vLLM v0.29 QSA does not support KV quantization).
Status: serving under debug (NaN logits on vLLM v0.29 TP8; benchmark vs BF16
and reference NVFP4 in progress).