license: other
license_name: qwen-community-1.0
license_link: https://huggingface.co/vmlinux/Qwen3.8-Flash-Next-Uncensored-Gufo-Q4Mix-GGUF/blob/main/LICENSE
base_model: orcarouter/Qwen3.8-Flash-Next-Uncensored
base_model_relation: quantized
library_name: gguf
pipeline_tag: image-text-to-text
language:
- en
- zh
tags: - gguf
- gufo
- qwen3.8
- flash-next
- strix-halo
- gfx1151
- rocm
- mtp
- vision
- uncensored
- q4_k
- q5_1
- iq4_nl
- q8_0
model_name: Qwen3.8 Flash-Next Uncensored for Gufo - Q4 Mix
Qwen3.8 Flash-Next Uncensored for Gufo — Q4 Mix
A Gufo-compatible GGUF conversion of orcarouter/Qwen3.8-Flash-Next-Uncensored. This is a quantization of that checkpoint, not a new fine-tune. It was built and tested on AMD Strix Halo (gfx1151) with Gufo commit eb91584.
Q4 Mix is this release's shorthand for its mixed tensor recipe. It is not a stock Q4_K_M preset: routed expert gate/up weights are Q4_K, expert down weights are Q5_1, the large per-layer embedding table is IQ4_NL, and the token embeddings, output head, and most dense projections are Q8_0. The quality-sensitive output head remains at higher precision. The exact recipe and build notes are included.
Files
| File | Purpose | Size | SHA-256 |
|---|---|---|---|
Qwen3.8-Flash-Next-Uncensored-Gufo-Q4Mix.gguf |
Text target with embedded PLE weights | 109,776,678,624 bytes (102.23 GiB) | c7f2331227aaefd1947b15fd6a31adeb29a4609a4dec5713670743b61ec44b7a |
mmproj-BF16.gguf |
Matching BF16 vision projector | 907,543,360 bytes | 0c2760d6f3e6686500989ad6ccc28aab69abfb272a3419ac33d4e26e7708ce8a |
The optional Q8_0 MTP predictor is not included. The tested sidecar is mtp/Qwen3.8-Flash-CIRU-STRIX-Orca-MTP-Q8_0.gguf from jcbtc's CIRU Orca release. Gufo also runs this target without MTP.
Run with Gufo
Build or install Gufo for Strix Halo, then download this repository's GGUF files. The source checkpoint is gated; this derivative is gated for the same reason. Accept access on the model page and log in with hf auth login before downloading.
hf download vmlinux/Qwen3.8-Flash-Next-Uncensored-Gufo-Q4Mix-GGUF \
--include '*.gguf' --local-dir ./orca-gufo
gufo serve --host 127.0.0.1 --port 8080 --sessions 2 llm \
--model ./orca-gufo/Qwen3.8-Flash-Next-Uncensored-Gufo-Q4Mix.gguf \
--mmproj ./orca-gufo/mmproj-BF16.gguf \
--served-model-name orca-gufo --context 262144
For speculative MTP, download the tested sidecar and add the following options to the gufo serve ... llm command:
hf download jcbtc/Qwen3.8-Flash-CIRU-STRIX-Orca \
--include 'mtp/Qwen3.8-Flash-CIRU-STRIX-Orca-MTP-Q8_0.gguf' \
--local-dir ./ciru-sidecar
# Add to gufo serve ... llm:
--speculative mtp \
--mtp-model ./ciru-sidecar/mtp/Qwen3.8-Flash-CIRU-STRIX-Orca-MTP-Q8_0.gguf \
--draft-tokens 6
The API is OpenAI-compatible at http://127.0.0.1:8080/v1; send model ID orca-gufo. Gufo's Qwen formatter is compiled into the runtime and does not load an external Jinja template. Thinking and effort can be controlled per request with reasoning_effort; --think off is a server default if desired. See Gufo's Flash-Next guide and server API documentation.
The repository contains a BF16 vision projector, but image input was tested only as a simple controlled smoke check. Generic GGUF runners and other GPUs have not been qualified for this particular package.
What was tested
The target loaded with Gufo on a Ryzen AI MAX+ 395 / Radeon 8060S (gfx1151). Text, a required function tool call, and a simple image question passed. The target also loaded with the linked MTP sidecar. A 262,144-token context configuration was tested with one session and short text/tool/image requests. The operator subsequently ran two concurrent 262,144-token text sessions on bare-bones Linux and reported roughly 40 tokens/s during long-context development-harness use. A live last-request /metrics gauge read 40.31 tokens/s, but it did not record prompt depth or aggregate two-user throughput. These are informal observations, not a controlled two-user benchmark or a test of two fully filled contexts.
For a controlled single-session comparison, the same prompts were served by Gufo and the existing jcbtc CIRU Orca package on the same host at 32K context, MTP enabled, greedy sampling, and thinking off. Each case had one warm-up and three measured streaming requests with cache_n=0:
| Case | Prompt / output tokens | Gufo median wall | CIRU median wall | Gufo speedup |
|---|---|---|---|---|
| Mixed short | 371 / 128 | 5.65 s | 6.82 s | 1.21× |
| Mixed medium | 2,678 / 128 | 6.52 s | 13.48 s | 2.07× |
| Mixed long | 10,583 / 128 | 12.20 s | 35.28 s | 2.89× |
| Repetitive short | 369 / 109 Gufo (110 CIRU) | 2.49 s | 3.62 s | 1.45× |
Prompt token counts matched across runtimes. The repetitive visible output matched, but CIRU counted it as 110 completion tokens. Gufo MTP reproduced Gufo autoregressive visible output on all 12 measured prompts. These results compare complete serving setups: GGUF quantization and runtime kernels both differ, so they do not isolate a format effect. They are not a broad quality evaluation. Raw JSON and the benchmark client are included.
Provenance and license
- BF16 source: orcarouter/Qwen3.8-Flash-Next-Uncensored, revision
8336e613. - Original base: Qwen/Qwen3.8-Flash-Next.
- Converter and quantizer: llama.cpp, local build at commit
d3e63dbfcfdac54bc7925a8ec5f44057035ac127. - Inference engine: Gufo, tested at commit
eb915840ffb62a8ec4b5c1adb41b04b5c1c75892. - Optional MTP sidecar: jcbtc/Qwen3.8-Flash-CIRU-STRIX-Orca; credit and download remain with its publisher.
The source and original base repositories contain identical copies of the Qwen Community License 1.0, reproduced here with the weights. The source model card currently labels its metadata apache-2.0; this release follows the license text distributed with the source weights and original base. Review LICENSE for its terms. This conversion is unofficial and is not endorsed by Qwen, OrcaRouter, jcbtc, or the Gufo project.