base_model: huihui-ai/Huihui-Qwen3.8-Flash-Next-abliterated
base_model_relation: quantized
license: other
license_name: qwen-community-1.0
license_link: https://huggingface.co/huihui-ai/Huihui-Qwen3.8-Flash-Next-abliterated/blob/298f94632b784e26a7fe576114f82066689d5baa/LICENSE
pipeline_tag: image-text-to-text
language:
- zh
- en
tags: - orinfer
- jetson
- e8p
- int8
- mixed-precision
- moe
- vision-language
- safetensors
- abliterated
- uncensored
- huihui
Huihui Qwen3.8-Flash-Next · Orinfer E8P Q2A8
A quantized deployment of huihui-ai/Huihui-Qwen3.8-Flash-Next-abliterated for Jetson AGX Orin 64GB, using Orinfer. The source model uses abliteration to reduce refusals. Supports text, images, multiple images, thinking, prefix caching and native MTP.
Model
| Item | Configuration |
|---|---|
| Routed experts | Near-2-bit E8P encoding with rotation; INT8 computation |
| Dense projections, shared experts, output head | W8; higher precision for sensitive paths |
| Vision encoder | FP16 |
| KV cache | Group-64 INT8 with FP16 scales |
| Context capacity | 262,144 tokens, including prompts, images and output |
| Model directory | Approximately 51.2 GB |
| GPU read-only weights | Approximately 34.52 GiB |
E8P stores each eight-weight vector in a 16-bit code. A fixed signed Hadamard rotation over 128 input channels spreads outlier energy before encoding; scales and codebook metadata add a small overhead above 2 bits per expert weight. Weight-only quantization uses seed 20261002.
Source revision: 298f946.
Deploy
Build Orinfer using its quick start, then download the model and its matching execution package:
hf download immiq/Huihui-Qwen3.8-Flash-Next-abliterated-Orinfer-E8P-Q2A8 \
--local-dir ./huihui-flash-next-orinfer
curl -fL \
https://github.com/iMMIQ/orinfer/releases/download/execution-flash-next-sm87-5bcf64efe5d8/flash_next-int8_quality-sm87-5bcf64efe5d81703a02fba41753101eaaaaf6d39c1500580de32c1d747f2b1d8.tar.gz \
-o ./flash-next-execution.tar.gz
PYTHONPATH=. .venv/bin/python tools/model/package.py install \
./flash-next-execution.tar.gz \
~/.cache/orinfer/packages
./target/release/orinfer serve ./huihui-flash-next-orinfer \
--listen 0.0.0.0:8088 --model huihui-flash-next \
--max-active-requests 1 --prefix-cache-mib 512 \
--cuda-graph full --mtp-drafts 7
The OpenAI-compatible Chat Completions endpoint is http://localhost:8088/v1; use model ID huihui-flash-next. Set enable_thinking: true to enable thinking and provide images through image_url. MTP supports up to 7 drafts. See the API examples.
Performance
Measured on Jetson AGX Orin 64GB, MAXN, GPU 1.30 GHz, CUDA 12.6.
Text prefill
| Input tokens | Tokens/s | Time to first token |
|---|---|---|
| 512 | 589 | 0.87 s |
| 2,048 | 794 | 2.58 s |
| 8,192 | 809 | 10.13 s |
Single stream, full CUDA Graph, MTP with up to 7 drafts, thinking off, prefix cache disabled. Median of three warmed requests per length; throughput is input tokens divided by end-to-end time to first token.
Decode with MTP
| Code generation task | Tokens/s |
|---|---|
| Python merge sort | 53.84 |
| Rust LRU cache | 41.19 |
| TypeScript async map | 30.25 |
Single stream, decode_only CUDA Graph, MTP with up to 7 drafts, greedy decoding, seed 20261002, thinking off. Measurements use warmed native replay with 96 generated tokens; throughput counts committed output tokens and excludes the first token and prefill time.
License
Model weights: Qwen Community License 1.0. Orinfer and execution-package code: LGPL-3.0-or-later. Credits to huihui-ai and Qwen.