orcarouter/Qwen3.8-Flash-Next-Uncensored-GGUF
? Why do I need an app?
The browser can't see your hardware
For privacy reasons a browser tells us at most "≥ 8 GB RAM, 8 cores" - same reading whether you have 8 GB or 128 GB. It has no idea how much RAM is free right now, which apps are open, or whether you have a GPU.
The Abliteration app is integrated with your machine
It reads your exact RAM, GPU model and VRAM, free memory right now, and picks the sharpest quant that still fits. Every model page lights up precisely for your rig.
And you can chat with any model, right now
The app is a full local runtime - no API keys, no subscription, everything runs on your machine. Click any model on this site and start a conversation in seconds.
Get the free app →This is a rough estimate. Install the free app - we'll show exact numbers.
Reading real hardware from your app right now. Numbers below are exact.
Below is the per-quantization compatibility for this model.
curl -H "Authorization: Bearer $ABL_KEY" \
"https://abliteration.org/api/v1/models/orcarouter%2FQwen3.8-Flash-Next-Uncensored-GGUF" - classification m8
- files 6
- hub_downloads_all_time 249,994
- author_summary 26 models
- readme_text full
Repackaging (quantization)
Why this label 3 signals
- 'abliterated' in name/tags
- is_gguf=1
- assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.
What is a refusal direction? →Genealogy
Full fork graph →This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.
Variants by this author
The same weights this author released in different packaging. Pick the format that matches your runtime.
- orcarouter/Qwen3.8-Flash-Next-Uncensored-GGUF (this)
- orcarouter/Qwen3.8-Flash-Next-Uncensored-FP8
- orcarouter/Qwen3.8-Flash-Next-Uncensored-NVFP4
- orcarouter/Qwen3.8-Flash-Next-Uncensored
- orcarouter/Qwen3.8-Flash-Next-Uncensored-MLX
Metadata
Related
Files by quantization
Q4_K 3 files 111 GB - ★ recommended for you
| Qwen3.8-Flash-Next-Uncensored-Q4_K_M-00002-of-00003.gguf | 41.6 GB | ******** | download |
| Qwen3.8-Flash-Next-Uncensored-Q4_K_M-00001-of-00003.gguf | 41.5 GB | ******** | download |
| Qwen3.8-Flash-Next-Uncensored-Q4_K_M-00003-of-00003.gguf | 27.8 GB | ******** | download |
F16 1 file 866 MB - ★ recommended for you
| mmproj-Qwen3.8-Flash-Next-Uncensored-F16.gguf | 866 MB | ******** | download |
README
Discussions
2026-10-09Can you please update the MTP draft model to be compatible with current llama.c…
Loading...2026-09-19Qwen3.8-Flash-Next iQ3_M GGUF: MTP achieves ~60% acceptance but is slower than …
Loading...2026-09-18Qwen3.8-Flash-Next Q6_K on 3× Radeon Instinct MI50
Loading...2026-09-15PRtest
Loading...2026-09-14KLD/PPL Calculations Request
Loading...2026-09-09Severe performance regression vs. unsloth's GGUF on identical hardware/config —…
Loading...2026-09-07Unable to download?
Loading...2026-09-07Is it possible to release mtp-head with Q4_K_M and Q6_K quantization?
Loading...2026-09-07Q8 has a tensor count issue.
Loading...2026-09-01Possibile Q8 + MTP
Loading...2026-08-31infinite loop
Loading...2026-08-31Dedicated n-gram GGUF file
Loading...2026-08-30Apple Silicon note: LM Studio can't load it, and the GGUF drops the MTP head (m…
Loading...2026-08-30Is it possible to have the Q6_K model in RAM with full precision (BF16?) PLE on…
Loading...2026-08-29i have a question
Loading...2026-08-29Q6_K looks reachable: the PLE tensor is 51.2B, not ~66B
Loading...2026-08-27NVFP4?
Loading...