← back to catalog · registered 2026-10-09 04:58

Maximilian228/Huihui-Qwen3.8-Flash-Next-abliterated-NVFP4-GPTQ-Strata

Maximilian228 GGUF MoE second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/Maximilian228%2FHuihui-Qwen3.8-Flash-Next-abliterated-NVFP4-GPTQ-Strata"
Response includes
  • classification m-uncensored
  • files 8
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-10-09

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
other
Languages
en zh
Quantizations
BF16 FP8
Tags
gguf nvfp4 gptq strata moe qwen3.8 flash-next abliterated uncensored mtp text-generation en

Related

Total size
54.4 GB
Files
8
Quantizations
3
Registered
2026-10-09 04:58
Last updated on HF
2026-10-09 04:21

Files by quantization

FP8 1 file 47.7 GB
ple-fp8.gguf 47.7 GB 40f95a62 download
BF16 1 file 1.18 GB
token-embd-bf16.gguf 1.18 GB 2533a5c4 download
Auxiliary files 6 files 5.58 GB
huihui-nvfp4-dense.gguf 5.58 GB 4a61a287 download
README.md 14.4 KB 4a109981 download
LICENSE 3.16 KB 9557a896 download
.gitattributes 1.65 KB 2bcc967f download
SHA256SUMS 1.44 KB bd82230a download
SIZES.txt 942 B 43c33976 download

README current version from Hugging Face


license: other
license_name: qwen-community-1.0
license_link: LICENSE
base_model:

  • huihui-ai/Huihui-Qwen3.8-Flash-Next-abliterated
    base_model_relation: quantized
    pipeline_tag: text-generation
    language:
  • en
  • zh
    tags:
  • nvfp4
  • gptq
  • strata
  • moe
  • qwen3.8
  • flash-next
  • abliterated
  • uncensored
  • mtp

Huihui-Qwen3.8-Flash-Next-abliterated — NVFP4 (GPTQ) for Strata

huihui-ai's abliterated Qwen3.8-Flash-Next
with every routed expert in NVFP4, quantized from the BF16 checkpoint by GPTQ (not by rounding each weight to the
nearest value), packed for Strata NVFP4: the 125B hybrid MoE on one RTX
20-50 card (12 GB of VRAM or more) and 64 GB of RAM or more.

Quick start (Windows): from the Strata NVFP4 release, run
START-HERE.bat --family huihui-nvfp4: it checks the PC, downloads this repository and starts the model. That works
once the fork's installer lists this repository (its next release); until then, see Run it.

  • Chain: Qwen/Qwen3.8-Flash-Next →
    huihui-ai/Huihui-Qwen3.8-Flash-Next-abliterated
    (abliteration, BF16) → this quantization.
  • 4.5 bits a weight for the experts, 63.3 GiB: the size and the decode cost of a ModelOpt NVFP4 checkpoint, about
    34% of plain rounding's error on held-out tokens (below). No layer is kept in 8 bits.
  • Only for Strata NVFP4. The files are Strata's own runtime formats (an expert pack, GGUFs the fork reads, the MTP
    draft head). They do not load in transformers, vLLM or llama.cpp.

The model does not refuse. It is an abliterated model: the refusal direction was removed from the weights
(see the base model's card). It will
follow harmful requests. What it is used for is on whoever runs it.

Files

file size Strata flag what it is
pack/experts.bin 63.28 GiB --pack pack the routed experts, 48 layers x 512, NVFP4 by GPTQ (this repo's point)
pack/native_experts.txt 1.6 KB the experts' layout
pack/dense.bin, pack/index.txt 1.43 GiB weights the engine keeps in its own form: hyper-connections, routers, norms, PLE key/value, indexer, DeltaNet gates
pack/tokenizer/ 10.1 MB (9.6 MiB) the server's tokenizer and chat template
huihui-nvfp4-dense.gguf 5.58 GiB --native, --native-dense-gguf the model's metadata and dense projections (Q8_0 / BF16) and head, without the routed experts
ple-fp8.gguf 47.68 GiB --ple-gguf the per-layer n-gram (PLE) table in FP8 E4M3 as Qwen ships it, read from the SSD
token-embd-bf16.gguf 1.18 GiB --embd-gguf the token embedding in BF16 as Qwen ships it
mtp/ 0.77 GiB --mtp mtp Qwen's MTP draft head (Q2_0 experts), for speculative decoding
LICENSE, SHA256SUMS, SIZES.txt the license, the files' SHA-256 and sizes (every file but this card)

What huihui's abliteration changed, compared tensor by tensor with the BF16 checkpoints: the weights that write into
the residual stream, in every layer - attention o_proj (12 layers), DeltaNet out_proj (36), the shared expert's
down_proj and every routed expert's down_proj (48). Which parts depend on it:

  • Model-specific: the experts (down_proj edited in every layer; gate_up_proj is Qwen's, but GPTQ fits the
    codes to this model's activations), the dense GGUF and pack/dense.bin (huihui's o_proj, out_proj and shared
    down_proj, Qwen's everything else).
  • The same as Qwen's original: the token embedding and the MTP draft head (every tensor byte-identical to
    Qwen/Qwen3.8-Flash-Next's, compared in full), the n-gram table (huihui's and OrcaRouter's are byte-identical, and
    ple-fp8.gguf is the very file the OrcaRouter NVFP4 repository ships: the same SHA-256), the experts'
    gate_up_proj, the routers and lm_head (compared on samples), the tokenizer, the chat template, config.json
    and the license file. These are here so the repository is complete on its own.

The dense GGUF is llama.cpp's converter run on the BF16 checkpoint with the fork's tools/nvfp4_convert.py type
policy (the projections Q8_0, the small ones BF16, the PLE convolution F16, the n-gram table and the MTP head left to
their own files), with the routed experts left out: next to pack/experts.bin the engine never reads them. Its
tensors are those of the OrcaRouter repository's orca-nvfp4-dense.gguf less the experts' per-matrix scales (the
pack carries its own), of the same types and shapes, and where neither abliteration edited a tensor the bytes are
the same.

How it was quantized

NVFP4 stores a weight as a 4-bit E2M1 code times its 16-value block's FP8 (E4M3) scale times one FP32 scale per
expert matrix. ModelOpt's NVFP4 scales each block to its largest value and rounds every weight to the nearest code
(RTN). Here each expert was quantized from the BF16 checkpoint with tools/requant.py --method gptq from the fork:

  • Calibration: the engine's own MoE inputs at every layer (STRATA_DUMP_MOE_INPUT, STRATA_DUMP_MOE_LAYER=all)
    for 57.5K tokens: Ukrainian and English Wikipedia, llama.cpp and Rust sources, and 16 chat answers
    (data/requant_calib in the fork), with the routing of each token, from a run of this model with its experts
    rounded to NVFP4 by ModelOpt's recipe (requant.py pack --method rtn).
  • Hessians: for each expert, H = sum w^2 x x^T over the tokens routed to it (w its routing weight: its
    output enters the residual times w), mixed with the layer's average input as 256 pseudo-tokens so rarely used
    experts get a sensible one. down's Hessian is built from silu(gate x) * up x of the already quantized
    gate/up, so down corrects the error gate/up left.
  • GPTQ: columns left to right in blocks of 128, 1% damping; each column's rounding error is spread over the
    columns not yet quantized through the inverse Hessian.
  • Block scales: each 16-column group's FP8 scale is chosen on the weights as GPTQ has updated them, among the
    codes for amax -> 6 and amax -> 4 ("four over six") and two neighbours of each, by the least squared error
    weighted by diag(H). The FP32 scale per expert matrix leaves 25% headroom so the updates do not saturate E4M3.
  • Every routed expert is NVFP4 (no 8-bit layers): the pack is as large as a ModelOpt one and a decode step costs
    the same. The shared experts, attention and DeltaNet projections are Q8_0 in the GGUF.

Error against BF16

requant.py analyze: per layer, on held-out tokens (every 4th calibration token, not used for the Hessians), the
routed-weighted squared error of the experts' whole output down(silu(gate x) * up x) against BF16, relative to the
output's energy; summed over the 48 layers weighted by each layer's energy:

48 layers, 57551 tokens per layer (14387 held out)

method summed error (rel. MSE x energy) energy-weighted rel. MSE of RTN
RTN (ModelOpt's rounding) 13607.41 0.01776 100.0%
GPTQ + block-scale search (this pack) 4561.85 0.00596 33.5%
GPTQ gate/up + Q8_0 down (reference, not used) 1905.40 0.00249 14.0%
Q8_0 everything (reference) 46.80 0.00006 0.3%

GPTQ below RTN in 48 of 48 layers; ratio GPTQ/RTN per layer: min 0.20, median 0.39, max 0.46

Per layer
layer RTN GPTQ GPTQ / RTN
0 0.00959 0.00204 0.21
1 0.01229 0.00380 0.31
2 0.01087 0.00342 0.31
3 0.01354 0.00447 0.33
4 0.01050 0.00219 0.21
5 0.01424 0.00498 0.35
6 0.01427 0.00514 0.36
7 0.01528 0.00572 0.37
8 0.01549 0.00605 0.39
9 0.01555 0.00588 0.38
10 0.01673 0.00691 0.41
11 0.01735 0.00733 0.42
12 0.01766 0.00761 0.43
13 0.01726 0.00722 0.42
14 0.01575 0.00621 0.39
15 0.01634 0.00484 0.30
16 0.01762 0.00611 0.35
17 0.01927 0.00802 0.42
18 0.01893 0.00814 0.43
19 0.02056 0.00902 0.44
20 0.02046 0.00893 0.44
21 0.01992 0.00913 0.46
22 0.02020 0.00930 0.46
23 0.02150 0.00916 0.43
24 0.01855 0.00763 0.41
25 0.01792 0.00725 0.40
26 0.01971 0.00837 0.42
27 0.02010 0.00867 0.43
28 0.02113 0.00928 0.44
29 0.01930 0.00820 0.42
30 0.01667 0.00612 0.37
31 0.01923 0.00636 0.33
32 0.01997 0.00761 0.38
33 0.02056 0.00872 0.42
34 0.01838 0.00727 0.40
35 0.02034 0.00845 0.42
36 0.01847 0.00682 0.37
37 0.01833 0.00657 0.36
38 0.01768 0.00705 0.40
39 0.01746 0.00585 0.33
40 0.01786 0.00627 0.35
41 0.01910 0.00766 0.40
42 0.01937 0.00726 0.37
43 0.01951 0.00798 0.41
44 0.01908 0.00706 0.37
45 0.01986 0.00697 0.35
46 0.02026 0.00659 0.33
47 0.01598 0.00314 0.20

Checks

Strata NVFP4 0.1.40.3-nvfp4.1 on an RTX 5090 (32 GB) with 128 GB of RAM, the arguments under Run it
(--spec 6 --spec-min-p 0.7).

  • It answers. "Explain in detail how a hash table works ... at least 300 words" (no thinking): a correct
    476-token answer that ends by itself (EOS), 182.8 tok/s. A 2,100-token prompt (a CUDA source file, cut off
    mid-line): prefill 2,810 tok/s; the model finishes the line (... q2[l0 / 2] >> 9);) and closes the turn, and
    in a 200-token greedy run it goes on to describe, in its thinking, what the file implements.

  • A decode round costs the same as with plain rounding (RTN) of the same experts. Three interleaved pairs on
    the hash-table question, each answer to its end (--stop-eos), expert cache auto (8,343 slots in every run),
    only experts.bin different:

    experts ms per round tokens per round tok/s
    RTN (ModelOpt's recipe) 12.60 ± 0.28 2.30 ± 0.05 182.4 ± 3.1
    this GPTQ pack 12.60 ± 0.07 2.38 ± 0.06 188.6 ± 5.7

    The paired difference per round is -0.003 ± 0.32 ms. Tokens per round depend on how many MTP drafts each
    answer's text lets through (64-79% accepted here); three pairs cannot tell their difference from the answers' own
    spread.

Run it

Build or download Strata NVFP4 (Windows, NVIDIA driver 580+; see its README
for the requirements: RAM, pagefile, large pages). Then:

hf download Maximilian228/Huihui-Qwen3.8-Flash-Next-abliterated-NVFP4-GPTQ-Strata --local-dir models\huihui-gptq

One-shot (prompt.txt holds token ids):

strata.exe --pack models\huihui-gptq\pack ^
  --native models\huihui-gptq\huihui-nvfp4-dense.gguf --native-dense-gguf models\huihui-gptq\huihui-nvfp4-dense.gguf ^
  --ple-gguf models\huihui-gptq\ple-fp8.gguf --embd-gguf models\huihui-gptq\token-embd-bf16.gguf ^
  --mtp models\huihui-gptq\mtp --spec 6 --spec-min-p 0.7 --prefill auto ^
  --expert-profile data\expert-profile.bin --expert-cache auto ^
  --max-context 262144 --kv int8 --tokens-file prompt.txt --max-new 512 --stop-eos

--max-context 262144 fits a 32 GB card; with less VRAM lower it (the fork's README and setup use these):

card's VRAM --max-context
32 GB 262144
24 GB 131072
16 GB 65536
12 GB 32768

data\expert-profile.bin ships with the engine. As a server (OpenAI and Anthropic APIs), put the same arguments in
the release's config\strata-nvfp4.json ("args") and point "tokenizer" at models/huihui-gptq/pack/tokenizer.
Pictures need the image encoder (mmproj-f32.gguf, not in this repo: the release's prepare-model.cmd builds
it from OrcaRouter's checkpoint, whose vision tower is byte-identical to huihui's, or convert_hf_to_gguf.py --mmproj
on the BF16 checkpoint); point the config's "vision" entry's "model" at huihui-nvfp4-dense.gguf (the encoder
reads only its vocabulary). Without the encoder, drop --vision from the arguments.

Already have the OrcaRouter NVFP4 repository? Its ple-fp8.gguf is this one, byte for byte; skip it with
--exclude ple-fp8.gguf and point --ple-gguf at yours. Its token-embd-bf16.gguf and mtp/ are OrcaRouter's
edited ones and do not belong to this model.

Reproduce

With the fork's tools and the BF16 checkpoint (hf download huihui-ai/Huihui-Qwen3.8-Flash-Next-abliterated, 361 GB):

  1. The dense GGUF: llama.cpp's convert_hf_to_gguf.py with tools/nvfp4_convert.py's type policy, the routed
    experts left out; pack/dense.bin, pack/index.txt and pack/tokenizer/ from it with tools/iq_pack.py's
    index_standalone and tools/strata_tokenizer.py (every float tensor in the engine's form; the script's CLI
    wants a GGUF with experts); ple-fp8.gguf with tools/ple_fp8_pack.py,
    token-embd-bf16.gguf with tools/embd_bf16_pack.py, mtp/ with tools/mtp_extract.py -> mtp_pack.py --experts q2_0 -> mtp_rt.py (plus the engine's data/draft_vocab.bin).
  2. A calibration pack by plain rounding, then the dump and GPTQ:
echo {"default": {"gu": "nvfp4", "d": "nvfp4"}} > calib\plan_nvfp4.json
.venv\Scripts\python tools\requant.py pack --bf16 models\huihui-bf16 --plan calib\plan_nvfp4.json --method rtn ^
  --base packs\huihui-dense --out packs\huihui-nvfp4-rtn
set STRATA_DUMP_MOE_INPUT=calib\d
set STRATA_DUMP_MOE_LAYER=all
strata.exe <the arguments above, --pack packs\huihui-nvfp4-rtn> --tokens-file data\requant_calib\calib_uk.txt --max-new 1   (and en, code, chat)
.venv\Scripts\python tools\requant.py analyze --bf16 models\huihui-bf16 --calib calib\d --methods rtn,gptq --out calib\errors.jsonl
.venv\Scripts\python tools\requant.py pack --bf16 models\huihui-bf16 --calib calib\d --plan calib\plan_nvfp4.json ^
  --method gptq --base packs\huihui-nvfp4-rtn --out packs\huihui-nvfp4-gptq

(packs\huihui-dense is the folder of step 1: dense.bin, index.txt, tokenizer/.) The packs here were built
layer group by layer group with the same requant.pack_layer calls, so the GPU was free between groups.

License

The Qwen Community License 1.0, as Qwen/Qwen3.8-Flash-Next and
huihui-ai/Huihui-Qwen3.8-Flash-Next-abliterated ship it (the same LICENSE file). Strata NVFP4 itself is MIT.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-10-09Upload checksums and model card10c65c914.4 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration