Orca Qwen3.8-Flash-Next Uncensored — native Halogen (.hgn)
Unofficial community build: a native Halogen checkpoint repacked fromorcarouter/Qwen3.8-Flash-Next-Uncensored-GGUF (IQ4_XS, 3 shards) withhalogen-flash-server:0.15.3. Serves directly on halogen-flash-server >= 0.15.3 —
no GGUF, no sidecar files needed.
File
orca-q38fn-uncensored.hgn— 101,995,332,160 bytes- SHA256:
35d97bba09e2f6a1c71bd8e8af6acdf582baca829d640e68306fd366073e2e6b - HGN1 container, 1198 tensors. FP8 PLE ngram table and MTP draft head are folded in.
Provenance
- Base model:
Qwen/Qwen3.8-Flash-Next, used under its original license terms;
this repository follows those terms. - Abliteration weights:
orcarouter/Qwen3.8-Flash-Next-Uncensored-GGUF— created
independently by the community author credited in the source repo. No retraining here. - This repo contains only repair + container conversion of the source weights
(no requantization; weights pass through unchanged).
Indexer repair
The conversion tooling that produced this GGUF exhibits systematic defects in the
sparse-attention (QSA) indexer: misaligned index_qk_proj values and zeroed or
shifted layernorms (23/24 indexer tensors affected in this build). These defects are
silent (the model loads and converses normally) but degrade sparse-layer retrieval.
All indexer values below were repaired from the pristine base model
(Qwen/Qwen3.8-Flash-Next):
- All 12
index_qk_projtensors set to base-model BF16 (bit-exact;
this build already stored BF16, no dtype correction needed). - Indexer layernorms validated against the canonical GGUF convention (stored gamma = 1 + w).
- Container conversion with
flash_serve --repack(0.15.3), folding in the official
MTP head and the engine's FP8 PLE ngram table.
Verification receipts
halogen-tools verify(0.15.3): PASS — 1198 tensors, 94.99 GiB, v2,
model_id qwen3.8-flash-next-gguf; header, table, layout, payload sizes, checksums,
codebooks and scales OK; geometry checked on 1197 tensors (1 with no rule).- Teacher-forced perplexity, 6,105-token mixed corpus (prose/code/dialog), chunk 1024,
identical tokenizer and corpus for every run:- this file: PPL 3.527 (mean NLL 1.2603)
- official
qwen38-flash-next-v2.hgnanchor: PPL 3.307 (mean NLL 1.1960)
Uniform small deltas across all length bands are consistent with quantization-level
difference rather than structural damage.
- Engine-load + decode validated on the sibling build (identical container format,
same 1198-tensor layout): chat completions OK, ~33 tok/s with MTP drafting. - SHA256 above re-verified against the on-disk file.
Usage
Point HALOGEN_CHECKPOINT at the file with halogen-flash-server >= 0.15.3
(tokenizer mounted at /tokenizer). The ngram table and draft head are inside the file.
Access & terms
This repository is gated: please state your intended use when requesting access.
- Unofficial research build. Not affiliated with the Qwen team.
- Behaviour (including "uncensored" outputs) is inherited unchanged from
the source GGUF and its license. Evaluate at your own discretion; no warranty.
Notes
- Repack is deterministic: an independent second repack of the same inputs is byte-identical.