← back to catalog · registered 2026-10-02 22:58

Jon-Nielsen/orca-q38fn-uncensored-hgn

Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/Jon-Nielsen%2Forca-q38fn-uncensored-hgn"
Response includes
  • classification m-uncensored
  • files 3
  • author_summary 3 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-10-02

Metadata

Tags
region:us

Related

Total size
0 B
Files
3
Quantizations
1
Registered
2026-10-02 22:58
Last updated on HF
2026-10-02 22:48

Files by quantization

Auxiliary files 3 files 95.0 GB
orca-q38fn-uncensored.hgn 95.0 GB ******** download
README.md 3.19 KB 57222c27 download
.gitattributes 1.54 KB 28dd4482 download

README current version from Hugging Face

Orca Qwen3.8-Flash-Next Uncensored — native Halogen (.hgn)

Unofficial community build: a native Halogen checkpoint repacked from
orcarouter/Qwen3.8-Flash-Next-Uncensored-GGUF (IQ4_XS, 3 shards) with
halogen-flash-server:0.15.3. Serves directly on halogen-flash-server >= 0.15.3 —
no GGUF, no sidecar files needed.

File

  • orca-q38fn-uncensored.hgn — 101,995,332,160 bytes
  • SHA256: 35d97bba09e2f6a1c71bd8e8af6acdf582baca829d640e68306fd366073e2e6b
  • HGN1 container, 1198 tensors. FP8 PLE ngram table and MTP draft head are folded in.

Provenance

  • Base model: Qwen/Qwen3.8-Flash-Next, used under its original license terms;
    this repository follows those terms.
  • Abliteration weights: orcarouter/Qwen3.8-Flash-Next-Uncensored-GGUF — created
    independently by the community author credited in the source repo. No retraining here.
  • This repo contains only repair + container conversion of the source weights
    (no requantization; weights pass through unchanged).

Indexer repair

The conversion tooling that produced this GGUF exhibits systematic defects in the
sparse-attention (QSA) indexer: misaligned index_qk_proj values and zeroed or
shifted layernorms (23/24 indexer tensors affected in this build). These defects are
silent (the model loads and converses normally) but degrade sparse-layer retrieval.

All indexer values below were repaired from the pristine base model
(Qwen/Qwen3.8-Flash-Next):

  1. All 12 index_qk_proj tensors set to base-model BF16 (bit-exact;
    this build already stored BF16, no dtype correction needed).
  2. Indexer layernorms validated against the canonical GGUF convention (stored gamma = 1 + w).
  3. Container conversion with flash_serve --repack (0.15.3), folding in the official
    MTP head and the engine's FP8 PLE ngram table.

Verification receipts

  • halogen-tools verify (0.15.3): PASS — 1198 tensors, 94.99 GiB, v2,
    model_id qwen3.8-flash-next-gguf; header, table, layout, payload sizes, checksums,
    codebooks and scales OK; geometry checked on 1197 tensors (1 with no rule).
  • Teacher-forced perplexity, 6,105-token mixed corpus (prose/code/dialog), chunk 1024,
    identical tokenizer and corpus for every run:
    • this file: PPL 3.527 (mean NLL 1.2603)
    • official qwen38-flash-next-v2.hgn anchor: PPL 3.307 (mean NLL 1.1960)
      Uniform small deltas across all length bands are consistent with quantization-level
      difference rather than structural damage.
  • Engine-load + decode validated on the sibling build (identical container format,
    same 1198-tensor layout): chat completions OK, ~33 tok/s with MTP drafting.
  • SHA256 above re-verified against the on-disk file.

Usage

Point HALOGEN_CHECKPOINT at the file with halogen-flash-server >= 0.15.3
(tokenizer mounted at /tokenizer). The ngram table and draft head are inside the file.

Access & terms

This repository is gated: please state your intended use when requesting access.

  • Unofficial research build. Not affiliated with the Qwen team.
  • Behaviour (including "uncensored" outputs) is inherited unchanged from
    the source GGUF and its license. Evaluate at your own discretion; no warranty.

Notes

  • Repack is deterministic: an independent second repack of the same inputs is byte-identical.
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration