base_model: huihui-ai/Huihui-Qwen3.8-27B-abliterated
library_name: transformers
tags:
- qwen3.8
- nvfp4
- modelopt
- sglang
Usage with DSpark for SGLang
sglang serve \
--trust-remote-code \
--model-path Qwen/Qwen3.8-27B-FP8 \
--tp-size 1 \
--speculative-algorithm DSPARK \
--speculative-draft-model-path RadixArk/Qwen3.8-27B-DSpark \
--speculative-dspark-block-size 7 \
--speculative-draft-model-quantization unquant \
--mamba-scheduler-strategy extra_buffer \
--attention-backend fa3
Usage with DSpark for vLLM
vllm serve /path/to/Qwen3.8-27B-GPTQ-Int4 \
--dtype bfloat16 \
--max-model-len 16384 \
--max-num-seqs 16 \
--gpu-memory-utilization 0.90 \
--kv-cache-dtype fp8 \
--block-size 64 \
--trust-remote-code \
--speculative-config '{
"method": "dspark",
"model": "/path/to/Qwen3.8-27B-DSpark-vLLM",
"num_speculative_tokens": 7,
"draft_sample_method": "probabilistic"
}'
Credit: