← back to catalog · registered 2026-09-23 20:57

klee100/Qwen3.8-Flash-Next-Uncensored-AutoRound-3bpw-MTP

klee100 multimodal second-order
curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/klee100%2FQwen3.8-Flash-Next-Uncensored-AutoRound-3bpw-MTP"
Response includes
  • classification m-uncensored
  • files 53
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-09-23

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
other
Languages
en zh
Tags
safetensors qwen4_exp qwen3.8 auto-round signroundv2 mixed-precision agentic-coding mtp vision-language en zh base_model:orcarouter/Qwen3.8-Flash-Next-Uncensored

Related

Total size
146 GB
Files
53
Quantizations
1
Registered
2026-09-23 20:57
Last updated on HF
2026-09-23 20:52

Files by quantization

Auxiliary files 53 files 146 GB
model-00026-of-00031.safetensors 5.00 GB f2341499 download
model-00027-of-00031.safetensors 5.00 GB df417dd6 download
model-00024-of-00031.safetensors 5.00 GB d86ba7d6 download
model-00028-of-00031.safetensors 5.00 GB dced362f download
model-00029-of-00031.safetensors 5.00 GB f74bc42e download
model-00023-of-00031.safetensors 5.00 GB c7f4ecb8 download
model-00025-of-00031.safetensors 5.00 GB b513e31b download
model-00018-of-00031.safetensors 4.47 GB 43a8aad9 download
model-00019-of-00031.safetensors 4.47 GB 1bc3357e download
model-00020-of-00031.safetensors 4.47 GB b02d2fab download
model-00021-of-00031.safetensors 4.47 GB 55cc10c9 download
model-00003-of-00031.safetensors 4.47 GB a014abd2 download
model-00004-of-00031.safetensors 4.47 GB 1a3f7b71 download
model-00005-of-00031.safetensors 4.47 GB 05c20a44 download
model-00006-of-00031.safetensors 4.47 GB a939fcd3 download
model-00007-of-00031.safetensors 4.47 GB e4fd7f48 download
model-00008-of-00031.safetensors 4.47 GB 15e7a1ae download
model-00009-of-00031.safetensors 4.47 GB 60bcc4c3 download
model-00010-of-00031.safetensors 4.47 GB d7ed7c4f download
model-00011-of-00031.safetensors 4.47 GB 6ea7deba download
model-00012-of-00031.safetensors 4.47 GB 8bcdd91b download
model-00013-of-00031.safetensors 4.47 GB 96a95ac9 download
model-00014-of-00031.safetensors 4.47 GB 1b18fc67 download
model-00015-of-00031.safetensors 4.47 GB a4b61cd7 download
model-00016-of-00031.safetensors 4.47 GB 42f038ab download
model-00017-of-00031.safetensors 4.47 GB 8d78c93d download
model-00002-of-00031.safetensors 4.47 GB 7b3aca2f download
model-00030-of-00031.safetensors 4.38 GB 4a39e74b download
mtp-packed-00001-of-00002.safetensors 3.13 GB 7ef73e87 download
ple-dedicated-00002-of-00002.safetensors 2.98 GB 7a2ffeec download
ple-dedicated-00001-of-00002.safetensors 2.98 GB 5cf40a85 download
model-00031-of-00031.safetensors 2.38 GB 804deeb2 download
regular-repacked-00002-of-00002.safetensors 2.02 GB a073ef38 download
regular-repacked-00001-of-00002.safetensors 2.00 GB 437fa775 download
mtp-packed-00002-of-00002.safetensors 1.73 GB 6907400d download
model.safetensors.index.json 22.8 MB 15466197 download
tokenizer.json 12.2 MB 0997f410 download
vocab.json 6.41 MB 0aa0ce06 download
config.json 4.51 MB bffbdeab download
quantization_config.json 4.06 MB 1092cad6 download
merges.txt 3.20 MB a494e019 download
tokenizer_config.json 17.5 KB 5de744b3 download
chat_template.jinja 8.74 KB c0c686f9 download
RECIPE.md 8.25 KB b396607f download
LICENSE 3.16 KB 9557a896 download
README.md 3.11 KB b2edcaca download
.gitattributes 1.60 KB aa7aacd0 download
processor_config.json 1.19 KB 43c4343e download
recipe.json 1.13 KB 182d88f3 download
calibration_manifest.json 978 B 04d3618a download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
generation_config.json 202 B 023756cf download

README current version from Hugging Face


language:


Qwen3.8-Flash-Next-Uncensored AutoRound ~3 bpw

This is a coding-calibrated mixed-bit quantization of the pinned
OrcaRouter BF16 source
revision 8336e613ea508b13c2159bd0f68965d97a606b95. It retains the
original vision tower, native BF16 MTP head, tokenizer, chat template, and
BF16 PLE table. ~3 bpw applies to the routed experts, not the entire
checkpoint
; the PLE table alone is approximately 102.4 GB.

Quantization

Intel AutoRound SignRoundV2 tuned all 48 backbone MoE blocks using 64
repository-distinct coding/tool windows of 2,048 tokens and 50 iterations.
Blocks 0–11 and 36–47 use W3A16G128; blocks 12–35 use W2A16G64. Dense
attention and DeltaNet projections use W8A16G128; routers, shared experts,
hyperconnections, vision, MTP, and PLE remain BF16. The calibration source
is NVIDIA Open-SWE-Traces
at f8fb5b3d2c787f85f8a00f5fe04fe3f1a11088ef (CC BY 4.0). See
RECIPE.md and recipe.json for reproducibility details.

Validation

  • Audit: 35 indexed weight shards, 223,082 tensors, 156,770,420,728 tensor
    bytes; all 31 native BF16 MTP tensors were bit-identical to the source.
  • Patched-vLLM generation and native MTP speculative decoding were tested.
    A 131,072-token, one-token generation smoke test peaked at 64,927 MiB
    (~63.4 GiB) GPU memory on an RTX PRO 6000 Blackwell with a 0.61
    utilization cap.
  • Paired held-out coding-trace BF16→quantized fidelity, one window and 128
    scored positions per length: 8K context KL 0.03718 nats, top-1
    agreement 96.09%; 32K context KL 0.13202 nats, top-1 agreement
    92.19%.
  • Four-task mini-SWE-agent/SWE-bench Verified check: 4/4 resolved
    by the official grader. Testing
    used thinking mode with temperature 1.0, top-p 0.95, top-k 20, min-p 0,
    presence penalty 0, and repetition penalty 1. The per-turn max_tokens
    limit was 8,192; observed completions stayed below it. All task
    repositories were absent from the calibration shard.

Serving

Use the pinned patched vLLM runtime and guide,
including its PLE SSD-offload and mixed-bit patch
against vLLM a5a30471ff2bb7f0824f2da10e358af98d304472. Set
QWEN_MODEL_DIR to this checkpoint and follow the guide's build steps; do
not download the previous checkpoint. The PLE table stays on SSD. For the
native BF16 MTP draft on the tested Blackwell host, use Triton MoE backend.
Qwen's recommended thinking-mode sampling was used for the agent validation.

This derivative retains the Qwen Community License 1.0 notice.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in app" button that hands off directly to a local runtime of your choice - Abliteration, LM Studio, or Ollama. No API keys, no subscription, no prompt leakage.