← back to catalog · registered 2026-10-03 08:58

alesha-pro/Huihui-Qwen3.8-Flash-Next-abliterated-exl3-4bit-hq_h6_ng6

alesha-pro multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/alesha-pro%2FHuihui-Qwen3.8-Flash-Next-abliterated-exl3-4bit-hq_h6_ng6"
Response includes
  • classification unknown
  • files 20
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
?
Primary method

Unclassified

No clear signals of an abliteration technique in this model.
Confidence
UNKNOWN
Why this label 1 signal
No classification signals present. This may not be an abliterated model at all - it could be a repackaging, a merge with unrelated goals, or unrelated content that mentions the term.
  • no classification signals present (no abliterated, uncensored, or known producer/method markers)
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-10-03

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
other
Tags
safetensors qwen4_exp exl3 quantized abliterated image-text-to-text conversational base_model:huihui-ai/Huihui-Qwen3.8-Flash-Next-abliterated base_model:quantized:huihui-ai/Huihui-Qwen3.8-Flash-Next-abliterated license:other 4-bit region:us

Related

Total size
30.0 GB
Files
20
Quantizations
1
Registered
2026-10-03 08:58
Last updated on HF
2026-10-03 09:20

Files by quantization

Auxiliary files 20 files 30.2 GB
model-00006-of-00009.safetensors 7.51 GB 1b84d4fe download
model-00002-of-00009.safetensors 7.51 GB 8ee5f735 download
model-00007-of-00009.safetensors 7.51 GB 8c07b238 download
model-00001-of-00009.safetensors 7.51 GB 29da1499 download
quantization_config.json 92.2 MB e6fac7ae download
model.safetensors.index.json 31.2 MB 9b05bf5d download
tokenizer.json 12.2 MB 0997f410 download
vocab.json 6.41 MB 0aa0ce06 download
merges.txt 3.20 MB a494e019 download
tokenizer_config.json 17.5 KB 5de744b3 download
strength1-patch-provenance.json 10.4 KB 31f88e58 download
chat_template.jinja 8.74 KB c0c686f9 download
README.md 5.99 KB d6bf2333 download
config.json 4.96 KB a56c3d49 download
LICENSE 3.16 KB 9557a896 download
SHA256SUMS 2.21 KB e6198b7e download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
.gitattributes 227 B 239d2999 download
generation_config.json 202 B 023756cf download

README current version from Hugging Face


license: other
license_name: qwen-community-license-1.0
license_link: LICENSE
base_model: huihui-ai/Huihui-Qwen3.8-Flash-Next-abliterated
pipeline_tag: image-text-to-text
tags:

  • exl3
  • quantized
  • abliterated

Huihui-Qwen3.8-Flash-Next-abliterated-exl3-4bit-hq_h6_ng6

This is a 4.05 bpw EXL3 HQ quantization of a corrected Huihui Qwen3.8 Flash Next checkpoint. It includes MTP, vision and the n-gram embedding table.

I rebuilt the BF16 source with ablation strength1.0 before quantization. These weights contain the correction and need no runtime ablation hook. The public repo name omits strength1.0, but the correction is part of the model you download.

Source correction

The source is huihui-ai/Huihui-Qwen3.8-Flash-Next-abliterated at revision 298f94632b784e26a7fe576114f82066689d5baa. I applied the following projection to 144 residual-writing tensors across its 48 trunk layers, using the fitted unit direction r:

W_new = W_Huihui - r (r^T W_Huihui)

The calculation runs in FP32 and writes back BF16. BF16 rounding leaves a small residual component. MTP, vision and n-gram source tensors were unchanged. strength1-patch-provenance.json includes the direction SHA256, affected tensor names and measured residual after rounding.

Quantization

EXL 1.5.3 produced this checkpoint with fresh calibration of 250 rows of 2048 tokens. The n-gram table reuses my unchanged K6 table from the corrected Q4 conversion.

Component Bits
Trunk target 4
Head 6
MTP 4
Vision 6
N-gram 6

The converter reported 4.05 bpw. The converter output totaled 107,463,164,330 bytes, including the 39,040,214,168-byte n-gram table. With ngram_ram: true, that table loads into system RAM. Download size is therefore larger than GPU weight memory. SHA256SUMS covers the packaged files.

Validation

On four RTX 3090s, I manually reviewed reasoning, a tool call followed by a real tool reply, and vision answers with MTP off and MTP2. Both modes also completed two turns of an OMP coding task at xhigh. The agent edited the code, ran the four original tests after each turn and produced coherent final answers. Independent checks covered the implementation and CLI.

An additional Range task passed 31 agent tests with MTP off and 38 with MTP2. The same 18 original tests passed independently in both modes. The MTP-off agent created and removed /tmp/orig_check despite a project-only instruction. Functional tests passed. That containment instruction was violated. These results cover the tested tasks only.

The speed run used eight prompt lengths from 4096 to 261632 tokens, C1 and C2, and two repeats in each MTP mode. Each request generated 512 actual tokens with zero reported cached prefix tokens. At the longest point, every request reached 261632 + 512 = 262144 total tokens. The two C2 requests had independent prompts and overlapping decode.

Full-context retrieval returned all three exact secrets in each C1/C2 request. Those answers used 57 or 58 output tokens. Full-context vision also reached 262144 total tokens per request in two C1/C2 repeats, with the shapes, colors and positions checked manually. The same image was reused, so its image-embedding cache may have been warm. The 512-token continuations were reviewed as excerpts. They do not establish complete long essays or factual accuracy of every claim.

Selected measurements on 4× RTX 3090 at 300 W per GPU. Decode is the median tokens/s per request. TTFT is seconds until the first token. C2 means two concurrent requests.

Prompt tokens C Decode off Decode MTP2 TTFT off TTFT MTP2
4096 1 58.09 86.26 2.65 2.74
4096 2 35.67 52.91 5.33 5.57
32768 1 58.22 90.96 18.88 19.74
32768 2 31.13 48.60 37.62 39.47
131072 1 57.21 87.45 76.44 80.10
131072 2 24.88 41.90 152.69 160.15
261632 1 57.21 87.98 156.17 163.81
261632 2 20.70 35.24 311.64 327.33

At 261632 prompt tokens in C1, median end-to-end time including prefill was 165.23 s with MTP off and 169.71 s with MTP2. Faster decode did not reduce total latency for that test.

The speed curve uses greedy nonthinking generation. Thinking quality tests use the sampler below. C2 decode rates are per-request rates, and changing MTP also changes the automatic layer placement. These measurements describe the tested serving configurations.

Sampling and serving

The tested thinking sampler uses these OpenAI-compatible request fields:

{
  "temperature": 1.0,
  "top_k": 20,
  "top_p": 0.95,
  "min_p": 0.0,
  "presence_penalty": 0.0,
  "repetition_penalty": 1.0,
  "reasoning_effort": "medium",
  "chat_template_kwargs": {"enable_thinking": true}
}

I checked low, medium and xhigh on the actual OMP requests. Serve with a compatible EXL3 runtime and the included chat template, vision enabled, ngram_ram: true and output_chunking: false.

The tested server used a 262144-token sequence limit, a 540672-token FP16 KV-cache, a maximum batch size of two and automatic layer placement across the four GPUs. CPU MoE expert offload was disabled. Sampled mode-wide GPU memory peaked at 82620 MiB with MTP off and 85670 MiB with MTP2, summed across all four GPUs.

The recorded runtime was EXL 1.5.3+cu128.torch2.8.0 at commit d3739fd393337b1ff4d6c2a342b12f0c87a9592f, TabbyAPI at be74bf0a00bcb3a518e6feb7606f150c189be637, Torch 2.8.0+cu128 and Transformers 5.13.1.

Other sizes and license

The separate Q3, Q4 and Q5 repos use the same corrected BF16 source.

These weights retain the Qwen Community License 1.0 shipped with the source checkpoint. Read that file for its terms.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration