license: other
license_name: qwen-community-1.0
license_link: LICENSE
base_model:
- huihui-ai/Huihui-Qwen3.8-Flash-Next-abliterated
base_model_relation: quantized
pipeline_tag: text-generation
language: - en
- zh
tags: - nvfp4
- gptq
- strata
- moe
- qwen3.8
- flash-next
- abliterated
- uncensored
- mtp
Huihui-Qwen3.8-Flash-Next-abliterated — NVFP4 (GPTQ) for Strata
huihui-ai's abliterated Qwen3.8-Flash-Next
with every routed expert in NVFP4, quantized from the BF16 checkpoint by GPTQ (not by rounding each weight to the
nearest value), packed for Strata NVFP4: the 125B hybrid MoE on one RTX
20-50 card (12 GB of VRAM or more) and 64 GB of RAM or more.
Quick start (Windows): from the Strata NVFP4 release, runSTART-HERE.bat --family huihui-nvfp4: it checks the PC, downloads this repository and starts the model. That works
once the fork's installer lists this repository (its next release); until then, see Run it.
- Chain:
Qwen/Qwen3.8-Flash-Next→huihui-ai/Huihui-Qwen3.8-Flash-Next-abliterated
(abliteration, BF16) → this quantization. - 4.5 bits a weight for the experts, 63.3 GiB: the size and the decode cost of a ModelOpt NVFP4 checkpoint, about
34% of plain rounding's error on held-out tokens (below). No layer is kept in 8 bits. - Only for Strata NVFP4. The files are Strata's own runtime formats (an expert pack, GGUFs the fork reads, the MTP
draft head). They do not load in transformers, vLLM or llama.cpp.
The model does not refuse. It is an abliterated model: the refusal direction was removed from the weights
(see the base model's card). It will
follow harmful requests. What it is used for is on whoever runs it.
Files
| file | size | Strata flag | what it is |
|---|---|---|---|
pack/experts.bin |
63.28 GiB | --pack pack |
the routed experts, 48 layers x 512, NVFP4 by GPTQ (this repo's point) |
pack/native_experts.txt |
1.6 KB | the experts' layout | |
pack/dense.bin, pack/index.txt |
1.43 GiB | weights the engine keeps in its own form: hyper-connections, routers, norms, PLE key/value, indexer, DeltaNet gates | |
pack/tokenizer/ |
10.1 MB (9.6 MiB) | the server's tokenizer and chat template | |
huihui-nvfp4-dense.gguf |
5.58 GiB | --native, --native-dense-gguf |
the model's metadata and dense projections (Q8_0 / BF16) and head, without the routed experts |
ple-fp8.gguf |
47.68 GiB | --ple-gguf |
the per-layer n-gram (PLE) table in FP8 E4M3 as Qwen ships it, read from the SSD |
token-embd-bf16.gguf |
1.18 GiB | --embd-gguf |
the token embedding in BF16 as Qwen ships it |
mtp/ |
0.77 GiB | --mtp mtp |
Qwen's MTP draft head (Q2_0 experts), for speculative decoding |
LICENSE, SHA256SUMS, SIZES.txt |
the license, the files' SHA-256 and sizes (every file but this card) |
What huihui's abliteration changed, compared tensor by tensor with the BF16 checkpoints: the weights that write into
the residual stream, in every layer - attention o_proj (12 layers), DeltaNet out_proj (36), the shared expert'sdown_proj and every routed expert's down_proj (48). Which parts depend on it:
- Model-specific: the experts (
down_projedited in every layer;gate_up_projis Qwen's, but GPTQ fits the
codes to this model's activations), the dense GGUF andpack/dense.bin(huihui'so_proj,out_projand shareddown_proj, Qwen's everything else). - The same as Qwen's original: the token embedding and the MTP draft head (every tensor byte-identical to
Qwen/Qwen3.8-Flash-Next's, compared in full), the n-gram table (huihui's and OrcaRouter's are byte-identical, andple-fp8.ggufis the very file the OrcaRouter NVFP4 repository ships: the same SHA-256), the experts'gate_up_proj, the routers andlm_head(compared on samples), the tokenizer, the chat template,config.json
and the license file. These are here so the repository is complete on its own.
The dense GGUF is llama.cpp's converter run on the BF16 checkpoint with the fork's tools/nvfp4_convert.py type
policy (the projections Q8_0, the small ones BF16, the PLE convolution F16, the n-gram table and the MTP head left to
their own files), with the routed experts left out: next to pack/experts.bin the engine never reads them. Its
tensors are those of the OrcaRouter repository's orca-nvfp4-dense.gguf less the experts' per-matrix scales (the
pack carries its own), of the same types and shapes, and where neither abliteration edited a tensor the bytes are
the same.
How it was quantized
NVFP4 stores a weight as a 4-bit E2M1 code times its 16-value block's FP8 (E4M3) scale times one FP32 scale per
expert matrix. ModelOpt's NVFP4 scales each block to its largest value and rounds every weight to the nearest code
(RTN). Here each expert was quantized from the BF16 checkpoint with tools/requant.py --method gptq from the fork:
- Calibration: the engine's own MoE inputs at every layer (
STRATA_DUMP_MOE_INPUT,STRATA_DUMP_MOE_LAYER=all)
for 57.5K tokens: Ukrainian and English Wikipedia, llama.cpp and Rust sources, and 16 chat answers
(data/requant_calibin the fork), with the routing of each token, from a run of this model with its experts
rounded to NVFP4 by ModelOpt's recipe (requant.py pack --method rtn). - Hessians: for each expert,
H = sum w^2 x x^Tover the tokens routed to it (wits routing weight: its
output enters the residual timesw), mixed with the layer's average input as 256 pseudo-tokens so rarely used
experts get a sensible one.down's Hessian is built fromsilu(gate x) * up xof the already quantized
gate/up, so down corrects the error gate/up left. - GPTQ: columns left to right in blocks of 128, 1% damping; each column's rounding error is spread over the
columns not yet quantized through the inverse Hessian. - Block scales: each 16-column group's FP8 scale is chosen on the weights as GPTQ has updated them, among the
codes for amax -> 6 and amax -> 4 ("four over six") and two neighbours of each, by the least squared error
weighted by diag(H). The FP32 scale per expert matrix leaves 25% headroom so the updates do not saturate E4M3. - Every routed expert is NVFP4 (no 8-bit layers): the pack is as large as a ModelOpt one and a decode step costs
the same. The shared experts, attention and DeltaNet projections are Q8_0 in the GGUF.
Error against BF16
requant.py analyze: per layer, on held-out tokens (every 4th calibration token, not used for the Hessians), the
routed-weighted squared error of the experts' whole output down(silu(gate x) * up x) against BF16, relative to the
output's energy; summed over the 48 layers weighted by each layer's energy:
48 layers, 57551 tokens per layer (14387 held out)
| method | summed error (rel. MSE x energy) | energy-weighted rel. MSE | of RTN |
|---|---|---|---|
| RTN (ModelOpt's rounding) | 13607.41 | 0.01776 | 100.0% |
| GPTQ + block-scale search (this pack) | 4561.85 | 0.00596 | 33.5% |
| GPTQ gate/up + Q8_0 down (reference, not used) | 1905.40 | 0.00249 | 14.0% |
| Q8_0 everything (reference) | 46.80 | 0.00006 | 0.3% |
GPTQ below RTN in 48 of 48 layers; ratio GPTQ/RTN per layer: min 0.20, median 0.39, max 0.46
Per layer
| layer | RTN | GPTQ | GPTQ / RTN |
|---|---|---|---|
| 0 | 0.00959 | 0.00204 | 0.21 |
| 1 | 0.01229 | 0.00380 | 0.31 |
| 2 | 0.01087 | 0.00342 | 0.31 |
| 3 | 0.01354 | 0.00447 | 0.33 |
| 4 | 0.01050 | 0.00219 | 0.21 |
| 5 | 0.01424 | 0.00498 | 0.35 |
| 6 | 0.01427 | 0.00514 | 0.36 |
| 7 | 0.01528 | 0.00572 | 0.37 |
| 8 | 0.01549 | 0.00605 | 0.39 |
| 9 | 0.01555 | 0.00588 | 0.38 |
| 10 | 0.01673 | 0.00691 | 0.41 |
| 11 | 0.01735 | 0.00733 | 0.42 |
| 12 | 0.01766 | 0.00761 | 0.43 |
| 13 | 0.01726 | 0.00722 | 0.42 |
| 14 | 0.01575 | 0.00621 | 0.39 |
| 15 | 0.01634 | 0.00484 | 0.30 |
| 16 | 0.01762 | 0.00611 | 0.35 |
| 17 | 0.01927 | 0.00802 | 0.42 |
| 18 | 0.01893 | 0.00814 | 0.43 |
| 19 | 0.02056 | 0.00902 | 0.44 |
| 20 | 0.02046 | 0.00893 | 0.44 |
| 21 | 0.01992 | 0.00913 | 0.46 |
| 22 | 0.02020 | 0.00930 | 0.46 |
| 23 | 0.02150 | 0.00916 | 0.43 |
| 24 | 0.01855 | 0.00763 | 0.41 |
| 25 | 0.01792 | 0.00725 | 0.40 |
| 26 | 0.01971 | 0.00837 | 0.42 |
| 27 | 0.02010 | 0.00867 | 0.43 |
| 28 | 0.02113 | 0.00928 | 0.44 |
| 29 | 0.01930 | 0.00820 | 0.42 |
| 30 | 0.01667 | 0.00612 | 0.37 |
| 31 | 0.01923 | 0.00636 | 0.33 |
| 32 | 0.01997 | 0.00761 | 0.38 |
| 33 | 0.02056 | 0.00872 | 0.42 |
| 34 | 0.01838 | 0.00727 | 0.40 |
| 35 | 0.02034 | 0.00845 | 0.42 |
| 36 | 0.01847 | 0.00682 | 0.37 |
| 37 | 0.01833 | 0.00657 | 0.36 |
| 38 | 0.01768 | 0.00705 | 0.40 |
| 39 | 0.01746 | 0.00585 | 0.33 |
| 40 | 0.01786 | 0.00627 | 0.35 |
| 41 | 0.01910 | 0.00766 | 0.40 |
| 42 | 0.01937 | 0.00726 | 0.37 |
| 43 | 0.01951 | 0.00798 | 0.41 |
| 44 | 0.01908 | 0.00706 | 0.37 |
| 45 | 0.01986 | 0.00697 | 0.35 |
| 46 | 0.02026 | 0.00659 | 0.33 |
| 47 | 0.01598 | 0.00314 | 0.20 |
Checks
Strata NVFP4 0.1.40.3-nvfp4.1 on an RTX 5090 (32 GB) with 128 GB of RAM, the arguments under Run it
(--spec 6 --spec-min-p 0.7).
It answers. "Explain in detail how a hash table works ... at least 300 words" (no thinking): a correct
476-token answer that ends by itself (EOS), 182.8 tok/s. A 2,100-token prompt (a CUDA source file, cut off
mid-line): prefill 2,810 tok/s; the model finishes the line (... q2[l0 / 2] >> 9);) and closes the turn, and
in a 200-token greedy run it goes on to describe, in its thinking, what the file implements.A decode round costs the same as with plain rounding (RTN) of the same experts. Three interleaved pairs on
the hash-table question, each answer to its end (--stop-eos), expert cache auto (8,343 slots in every run),
onlyexperts.bindifferent:experts ms per round tokens per round tok/s RTN (ModelOpt's recipe) 12.60 ± 0.28 2.30 ± 0.05 182.4 ± 3.1 this GPTQ pack 12.60 ± 0.07 2.38 ± 0.06 188.6 ± 5.7 The paired difference per round is -0.003 ± 0.32 ms. Tokens per round depend on how many MTP drafts each
answer's text lets through (64-79% accepted here); three pairs cannot tell their difference from the answers' own
spread.
Run it
Build or download Strata NVFP4 (Windows, NVIDIA driver 580+; see its README
for the requirements: RAM, pagefile, large pages). Then:
hf download Maximilian228/Huihui-Qwen3.8-Flash-Next-abliterated-NVFP4-GPTQ-Strata --local-dir models\huihui-gptq
One-shot (prompt.txt holds token ids):
strata.exe --pack models\huihui-gptq\pack ^
--native models\huihui-gptq\huihui-nvfp4-dense.gguf --native-dense-gguf models\huihui-gptq\huihui-nvfp4-dense.gguf ^
--ple-gguf models\huihui-gptq\ple-fp8.gguf --embd-gguf models\huihui-gptq\token-embd-bf16.gguf ^
--mtp models\huihui-gptq\mtp --spec 6 --spec-min-p 0.7 --prefill auto ^
--expert-profile data\expert-profile.bin --expert-cache auto ^
--max-context 262144 --kv int8 --tokens-file prompt.txt --max-new 512 --stop-eos
--max-context 262144 fits a 32 GB card; with less VRAM lower it (the fork's README and setup use these):
| card's VRAM | --max-context |
|---|---|
| 32 GB | 262144 |
| 24 GB | 131072 |
| 16 GB | 65536 |
| 12 GB | 32768 |
data\expert-profile.bin ships with the engine. As a server (OpenAI and Anthropic APIs), put the same arguments in
the release's config\strata-nvfp4.json ("args") and point "tokenizer" at models/huihui-gptq/pack/tokenizer.
Pictures need the image encoder (mmproj-f32.gguf, not in this repo: the release's prepare-model.cmd builds
it from OrcaRouter's checkpoint, whose vision tower is byte-identical to huihui's, or convert_hf_to_gguf.py --mmproj
on the BF16 checkpoint); point the config's "vision" entry's "model" at huihui-nvfp4-dense.gguf (the encoder
reads only its vocabulary). Without the encoder, drop --vision from the arguments.
Already have the OrcaRouter NVFP4 repository? Its ple-fp8.gguf is this one, byte for byte; skip it with--exclude ple-fp8.gguf and point --ple-gguf at yours. Its token-embd-bf16.gguf and mtp/ are OrcaRouter's
edited ones and do not belong to this model.
Reproduce
With the fork's tools and the BF16 checkpoint (hf download huihui-ai/Huihui-Qwen3.8-Flash-Next-abliterated, 361 GB):
- The dense GGUF: llama.cpp's
convert_hf_to_gguf.pywithtools/nvfp4_convert.py's type policy, the routed
experts left out;pack/dense.bin,pack/index.txtandpack/tokenizer/from it withtools/iq_pack.py'sindex_standaloneandtools/strata_tokenizer.py(every float tensor in the engine's form; the script's CLI
wants a GGUF with experts);ple-fp8.ggufwithtools/ple_fp8_pack.py,token-embd-bf16.ggufwithtools/embd_bf16_pack.py,mtp/withtools/mtp_extract.py->mtp_pack.py --experts q2_0->mtp_rt.py(plus the engine'sdata/draft_vocab.bin). - A calibration pack by plain rounding, then the dump and GPTQ:
echo {"default": {"gu": "nvfp4", "d": "nvfp4"}} > calib\plan_nvfp4.json
.venv\Scripts\python tools\requant.py pack --bf16 models\huihui-bf16 --plan calib\plan_nvfp4.json --method rtn ^
--base packs\huihui-dense --out packs\huihui-nvfp4-rtn
set STRATA_DUMP_MOE_INPUT=calib\d
set STRATA_DUMP_MOE_LAYER=all
strata.exe <the arguments above, --pack packs\huihui-nvfp4-rtn> --tokens-file data\requant_calib\calib_uk.txt --max-new 1 (and en, code, chat)
.venv\Scripts\python tools\requant.py analyze --bf16 models\huihui-bf16 --calib calib\d --methods rtn,gptq --out calib\errors.jsonl
.venv\Scripts\python tools\requant.py pack --bf16 models\huihui-bf16 --calib calib\d --plan calib\plan_nvfp4.json ^
--method gptq --base packs\huihui-nvfp4-rtn --out packs\huihui-nvfp4-gptq
(packs\huihui-dense is the folder of step 1: dense.bin, index.txt, tokenizer/.) The packs here were built
layer group by layer group with the same requant.pack_layer calls, so the GPU was free between groups.
License
The Qwen Community License 1.0, as Qwen/Qwen3.8-Flash-Next andhuihui-ai/Huihui-Qwen3.8-Flash-Next-abliterated ship it (the same LICENSE file). Strata NVFP4 itself is MIT.