← back to catalog · registered 2026-10-09 01:58

immiq/Huihui-Qwen3.8-Flash-Next-abliterated-Orinfer-E8P-Q2A8

immiq MoE multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/immiq%2FHuihui-Qwen3.8-Flash-Next-abliterated-Orinfer-E8P-Q2A8"
Response includes
  • classification m1
  • files 12
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-10-09

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
other
Languages
zh en
Tags
qwen4_exp orinfer jetson e8p int8 mixed-precision moe vision-language safetensors abliterated uncensored huihui

Related

Total size
0 B
Files
12
Quantizations
1
Registered
2026-10-09 01:58
Last updated on HF
2026-10-09 02:18

Files by quantization

Auxiliary files 12 files 9.65 MB
vocab.json 6.41 MB 0aa0ce06 download
merges.txt 3.20 MB a494e019 download
tokenizer_config.json 17.5 KB 5de744b3 download
chat_template.jinja 8.74 KB c0c686f9 download
config.json 4.36 KB 0eaf9719 download
README.md 4.22 KB 77f65119 download
README_zh.md 3.63 KB b9f08bc1 download
MODEL_LICENSE 3.16 KB 9557a896 download
.gitattributes 1.48 KB a6344aac download
source.json 779 B e8845eaf download
preprocessor_config.json 390 B 2ea84a43 download
generation_config.json 202 B 023756cf download

README current version from Hugging Face


base_model: huihui-ai/Huihui-Qwen3.8-Flash-Next-abliterated
base_model_relation: quantized
license: other
license_name: qwen-community-1.0
license_link: https://huggingface.co/huihui-ai/Huihui-Qwen3.8-Flash-Next-abliterated/blob/298f94632b784e26a7fe576114f82066689d5baa/LICENSE
pipeline_tag: image-text-to-text
language:

  • zh
  • en
    tags:
  • orinfer
  • jetson
  • e8p
  • int8
  • mixed-precision
  • moe
  • vision-language
  • safetensors
  • abliterated
  • uncensored
  • huihui

Huihui Qwen3.8-Flash-Next · Orinfer E8P Q2A8

中文

A quantized deployment of huihui-ai/Huihui-Qwen3.8-Flash-Next-abliterated for Jetson AGX Orin 64GB, using Orinfer. The source model uses abliteration to reduce refusals. Supports text, images, multiple images, thinking, prefix caching and native MTP.

Model

Item Configuration
Routed experts Near-2-bit E8P encoding with rotation; INT8 computation
Dense projections, shared experts, output head W8; higher precision for sensitive paths
Vision encoder FP16
KV cache Group-64 INT8 with FP16 scales
Context capacity 262,144 tokens, including prompts, images and output
Model directory Approximately 51.2 GB
GPU read-only weights Approximately 34.52 GiB

E8P stores each eight-weight vector in a 16-bit code. A fixed signed Hadamard rotation over 128 input channels spreads outlier energy before encoding; scales and codebook metadata add a small overhead above 2 bits per expert weight. Weight-only quantization uses seed 20261002.

Source revision: 298f946.

Deploy

Build Orinfer using its quick start, then download the model and its matching execution package:

hf download immiq/Huihui-Qwen3.8-Flash-Next-abliterated-Orinfer-E8P-Q2A8 \
  --local-dir ./huihui-flash-next-orinfer

curl -fL \
  https://github.com/iMMIQ/orinfer/releases/download/execution-flash-next-sm87-5bcf64efe5d8/flash_next-int8_quality-sm87-5bcf64efe5d81703a02fba41753101eaaaaf6d39c1500580de32c1d747f2b1d8.tar.gz \
  -o ./flash-next-execution.tar.gz

PYTHONPATH=. .venv/bin/python tools/model/package.py install \
  ./flash-next-execution.tar.gz \
  ~/.cache/orinfer/packages

./target/release/orinfer serve ./huihui-flash-next-orinfer \
  --listen 0.0.0.0:8088 --model huihui-flash-next \
  --max-active-requests 1 --prefix-cache-mib 512 \
  --cuda-graph full --mtp-drafts 7

The OpenAI-compatible Chat Completions endpoint is http://localhost:8088/v1; use model ID huihui-flash-next. Set enable_thinking: true to enable thinking and provide images through image_url. MTP supports up to 7 drafts. See the API examples.

Performance

Measured on Jetson AGX Orin 64GB, MAXN, GPU 1.30 GHz, CUDA 12.6.

Text prefill

Input tokens Tokens/s Time to first token
512 589 0.87 s
2,048 794 2.58 s
8,192 809 10.13 s

Single stream, full CUDA Graph, MTP with up to 7 drafts, thinking off, prefix cache disabled. Median of three warmed requests per length; throughput is input tokens divided by end-to-end time to first token.

Decode with MTP

Code generation task Tokens/s
Python merge sort 53.84
Rust LRU cache 41.19
TypeScript async map 30.25

Single stream, decode_only CUDA Graph, MTP with up to 7 drafts, greedy decoding, seed 20261002, thinking off. Measurements use warmed native replay with 96 generated tokens; throughput counts committed output tokens and excludes the first token and prefill time.

License

Model weights: Qwen Community License 1.0. Orinfer and execution-package code: LGPL-3.0-or-later. Credits to huihui-ai and Qwen.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration