RVN Qwen3.8-Flash-Next Abliterated — native Halogen (.hgn)
Unofficial community build: a native Halogen checkpoint repacked from0bserverx/RVN-Qwen3.8-Flash-Next-Abliterated-Uncensored-GGUF (IQ4_XS, 8 shards)
with halogen-flash-server:0.15.3. Serves directly on halogen-flash-server >= 0.15.3 —
no GGUF, no sidecar files needed.
File
rvn-q38fn-abliterated.hgn— 101,995,332,160 bytes- SHA256:
c46b0da5f8c661c5a086e03711416aa3cd866467fc4d441bc5639e8b46b5c130 - HGN1 container, 1198 tensors. FP8 PLE ngram table and MTP draft head are folded in.
Provenance
- Base model:
Qwen/Qwen3.8-Flash-Next, used under its original license terms;
this repository follows those terms. - Abliteration weights:
0bserverx/RVN-...-GGUF(IQ4_XS) — created independently by
the community author credited in the source repo. No retraining here. - This repo contains only repair + container conversion of the source weights
(no requantization; weights pass through unchanged).
Indexer repair
The conversion tooling that produced this GGUF exhibits systematic defects in the
sparse-attention (QSA) indexer: misaligned index_qk_proj values, zeroed or shifted
layernorms, and in some builds F16 stored where BF16 is expected. These defects are
silent (the model loads and converses normally) but degrade sparse-layer retrieval.
All indexer values below were repaired from the pristine base model
(Qwen/Qwen3.8-Flash-Next):
- All 12
index_qk_projtensors set to base-model BF16 (bit-exact). - F16 -> BF16 dtype correction on the indexer tensors.
- Indexer layernorms validated against the canonical GGUF convention (stored gamma = 1 + w).
- Container conversion with
flash_serve --repack(0.15.3), folding in the official
MTP head and the engine's FP8 PLE ngram table.
Verification receipts
halogen-tools verify(0.15.3): PASS — 1198 tensors, 94.99 GiB, v2,
model_id qwen3.8-flash-next-gguf; header, table, layout, payload sizes, checksums,
codebooks and scales OK; geometry checked on 1197 tensors (1 with no rule).- Teacher-forced perplexity, 6,105-token mixed corpus (prose/code/dialog), chunk 1024,
identical tokenizer and corpus for every run:- this file: PPL 3.511 (mean NLL 1.2558)
- official
qwen38-flash-next-v2.hgnanchor: PPL 3.307 (mean NLL 1.1960)
Uniform small deltas across all length bands are consistent with quantization-level
difference rather than structural damage.
- Live serve on 0.15.3: chat completions OK, instruction-following exact,
~33 tok/s decode with MTP drafting, ~107 tok/s prefill (Strix Halo, gfx1151). - SHA256 above re-verified against the file that was served.
Usage
Point HALOGEN_CHECKPOINT at the file with halogen-flash-server >= 0.15.3
(tokenizer mounted at /tokenizer). The ngram table and draft head are inside the file.
Access & terms
This repository is gated: please state your intended use when requesting access.
- Unofficial research build. Not affiliated with the Qwen team.
- Behaviour (including "abliterated/uncensored" outputs) is inherited unchanged from
the source GGUF and its license. Evaluate at your own discretion; no warranty.
Notes
- Repack is deterministic: an independent second repack of the same inputs is byte-identical.