---
license: other
license_name: qwen-community-1.0
license_link: https://huggingface.co/Dimachaerus/Qwen3.8-Flash-Next-Uncensored-QF-107GiB-GGUF/blob/main/LICENSE
base_model:
- orcarouter/Qwen3.8-Flash-Next-Uncensored
base_model_relation: quantized
pipeline_tag: image-text-to-text
library_name: gguf
quantized_by: Dimachaerus
language: - en
- zh
tags: - qwen
- qwen3.8
- qwen4-exp
- flash-next
- gguf
- llama.cpp
- uncensored
- abliterated
- mixture-of-experts
- vision-language
- multimodal
- imatrix
- mixed-quant
- 128gb
- dgx-spark
- quality-focused
- research
- not-for-all-audiences
extra_gated_prompt: >-
This is an abliterated research model with substantially reduced safety alignment.
It may produce harmful, illegal, offensive, biased, or otherwise unsafe content that
an aligned model would refuse. Access is provided for legitimate research, controlled
evaluation, and local inference. You are responsible for lawful use, appropriate
downstream safeguards, and compliance with the Qwen Community License 1.0.
extra_gated_fields:
I have read and accept the included Qwen Community License: checkbox
I understand that safety alignment has been substantially reduced: checkbox
I will use this release lawfully and take responsibility for appropriate safeguards: checkbox
Qwen3.8-Flash-Next-Uncensored QF-107GiB GGUF
Quality-focused mixed GGUF for ~128 GiB-class local inference systems.
This release is a tensor-specific quantization oforcarouter/Qwen3.8-Flash-Next-Uncensored,
derived from the Qwen3.8-Flash-Next family and optimized to spend a ~107 GiB language-model
budget where additional precision produced the greatest measured fidelity benefit.
The design goal is not to minimize file size. It is to use the memory envelope of
128-GiB-class systems intelligently: keep the very large PLE / n-gram table at Q5_1,
then spend substantially more precision on the routed MoE expert weights.
[!CAUTION]
Research artifact with substantially reduced safety alignment.
This model is derived from an abliterated checkpoint and should not be treated as having
dependable built-in safety guardrails. Do not expose it to untrusted end users or production
traffic without independent moderation, access controls, logging, abuse prevention, and any
other safeguards appropriate to the deployment. Users are responsible for lawful use.
What is different about this quant?
The allocation is intentionally architecture-aware rather than a uniform Q4/Q5/Q6 recipe.
| Component | Quantization |
|---|---|
| PLE / per-layer token embedding table | Q5_1 |
| Expert gate/up, blocks 0-3 | Q6_K |
| Expert gate/up, blocks 4-39 | IQ4_XS |
| Expert gate/up, blocks 40-47 | Q6_K |
| Expert down | IQ4_NL |
| Other quantizable tensors | Q8_0 |
blk.1.ple_conv1d.weight |
F16 |
| Vision projector | BF16 |
The quantizer dry run reported approximately 109,785.93 MiB / 107.21 GiB for the
language GGUF at an average of approximately 5.20 BPW, plus the separate ~908 MB BF16
vision projector.
Why this allocation?
AtomicChat's Qwen3.8-Flash-Next work showed that increasing the precision of the enormous
PLE table produced very little additional fidelity compared with its storage cost.
This build therefore leaves PLE at Q5_1 and redirects the additional memory budget toward
routed experts, with Q6_K reserved for edge blocks and IQ4_XS used for the remaining
gate/up expert tensors.
The result is a larger build than AD-4.27, intended for systems where roughly 128 GiB of
aggregate/unified AI memory is available and model fidelity is preferred over minimum size.
Evaluation
Both this release and Navin's AD-4.27 uncensored GGUF were evaluated locally against the
same high-quality reference quant, on the same held-out corpus, with the same
llama.cpp evaluator and settings.
The high-quality reference was itself derived from the same BF16 source and used:
- PLE: Q5_1
- expert gate/up: Q6_K
- expert down: Q8_0
- other quantizable tensors: Q8_0
- odd PLE convolution tensor: F16
The held-out set was AtomicChat/calib-corpora/eval/neutral/eval_neutral.txt,
87 chunks at 4096 context.
HQ-reference-relative results
| Metric | QF-107GiB (this release) | AD-4.27 | Relative result |
|---|---|---|---|
| Reference PPL | 4.048445 | 4.048445 | same reference |
| Candidate PPL | 4.055184 | 4.158098 | lower is better |
| PPL degradation vs reference | +0.1665% | +2.7085% | 93.9% less degradation |
| Mean KLD | 0.024226 | 0.082145 | 70.5% lower |
| Median KLD | 0.005256 | 0.018782 | 72.0% lower |
| 99% KLD | 0.275788 | 0.918131 | 70.0% lower |
| 99.9% KLD | 0.852830 | 2.510427 | 66.0% lower |
| RMS probability error | 4.840% | 8.819% | 45.1% lower |
| Same top token | 94.257% | 89.620% | +4.637 points |
Top-token disagreement falls from 10.380% to 5.743%, a 44.7% reduction in
top-token disagreements relative to the common HQ reference.
A separate PPL-only run on the same 87-chunk corpus produced:
| Model | PPL |
|---|---|
| QF-107GiB | 4.0572 ± 0.02168 |
| AD-4.27 | 4.1569 ± 0.02235 |
Important evaluation limitation
These KLD figures are not BF16-relative KLD values. The reference is a much
higher-precision local Q6/Q8 proxy derived from the BF16 source, not the literal BF16 model.
The numbers are therefore suitable for a controlled A/B comparison between these two
quants, but should not be compared numerically with published BF16-reference KLD values
from another evaluation.
See EVALUATION.md for the protocol and exact commands.
Source, calibration, and toolchain
- Upstream family:
Qwen/Qwen3.8-Flash-Next - Uncensored BF16 source:
orcarouter/Qwen3.8-Flash-Next-Uncensored - Reproduction source pin:
8336e613ea508b13c2159bd0f68965d97a606b95 - Importance matrix:
AtomicChat/Qwen3.8-Flash-Next-GGUF/imatrix.gguf - Importance-matrix SHA-256:
5591ce3dc3bf0b73d3c074bc588c90b6c4f7b3c273de6b10e50d111b75f05487 - AtomicChat calibration: 4,967,044 tokens across 4,000 chunks
- Vision projector: AtomicChat
mmproj-Qwen3.8-Flash-Next-BF16.gguf - llama.cpp conversion/quantization revision:
05af0d2b1
See REPRODUCIBILITY.md and ATTRIBUTION.md.
Hardware target
This release is aimed at ~128 GiB-class systems, including configurations such as:
- 96 GiB + 32 GiB multi-GPU workstations
- 4x32 GiB systems
- 2x64 GiB systems
- 128 GiB unified-memory systems such as GB10 / DGX Spark-class machines
- larger systems where preserving context/KV headroom is desirable
Actual usable context length depends on runtime, KV precision, projector placement, GPU
topology, and other allocations. "Fits in 128 GiB" should not be interpreted as a guarantee
that every 128-GB system can run every context configuration without tuning.
Vision
The release includes the BF16 vision projector:
mmproj-Qwen3.8-Flash-Next-BF16.gguf
Keep the projector with the model shards and select it in a multimodal-capable llama.cpp /
LM Studio runtime.
llama.cpp example
Use the first shard of the model and the BF16 projector:
llama-mtmd-cli \
-m <first-model-shard>.gguf \
--mmproj mmproj-Qwen3.8-Flash-Next-BF16.gguf \
-ngl 999 \
-c 32768 \
--jinja
Hardware-specific tensor splits, context size, KV precision, and offload settings should be
adapted to the host.
LM Studio
This build was created for and tested with a modern llama.cpp-based LM Studio runtime.
Download all model shards plus the BF16 mmproj, keep the shard filenames unchanged,
and load the first shard.
Intended use
This artifact is published primarily for:
- local inference research
- quantization and rate-distortion research
- controlled capability evaluation
- multimodal/long-context experiments
- study of refusal-removal / alignment behavior
- red-team and robustness research in controlled environments
It is not a safety-tuned deployment model.
Limitations
Quantization changes model behavior. Even with the measured fidelity results above, output
can differ from the BF16 source. Long-context behavior, tool use, multimodal quality, rare
token behavior, and generation trajectories are not fully characterized by PPL/KLD alone.
The source is abliterated / uncensored. It may respond to requests that the aligned upstream
model would refuse and may emit offensive, biased, unsafe, or factually incorrect content.
License
The upstream Qwen3.8-Flash-Next model is distributed under Qwen Community License 1.0.
The OrcaRouter source checkpoint used for this derivative includes the same license text.
This repository therefore includes and follows Qwen Community License 1.0 rather than
relying on any more permissive metadata label that may appear on a derivative repository.
Review the Qwen Community License 1.0 before using or redistributing the model.
Attribution
This release depends on work from Qwen / Alibaba, OrcaRouter, AtomicChat, and llama.cpp.
See ATTRIBUTION.md for detailed provenance and acknowledgements.
Safety and access
This repository is released as a gated model and is taggednot-for-all-audiences. Gating is an access-management measure, not a technical safety
control. Anyone who receives the weights may redistribute copies subject to the applicable
license.
See SAFETY.md.