← back to catalog · registered 2026-10-10 16:58

beginning-ai/Qwen3.8-Flash-Next-Uncensored-NVFP4

beginning-ai multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/beginning-ai%2FQwen3.8-Flash-Next-Uncensored-NVFP4"
Response includes
  • classification m-uncensored
  • files 223
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-10-10

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
other
Tags
safetensors qwen4_exp nvfp4 fp8 modelopt sglang blackwell uncensored abliterated image-text-to-text conversational base_model:orcarouter/Qwen3.8-Flash-Next-Uncensored

Related

Total size
120 GB
Files
223
Quantizations
1
Registered
2026-10-10 16:58
Last updated on HF
2026-10-10 15:57

Files by quantization

Auxiliary files 223 files 120 GB
model-ple-00008.safetensors 4.84 GB 644e0e5c download
model-ple-00001.safetensors 4.84 GB 587c4884 download
model-ple-00002.safetensors 4.84 GB 809372b4 download
model-ple-00003.safetensors 4.84 GB 76a85aca download
model-ple-00004.safetensors 4.84 GB dc9842b4 download
model-ple-00005.safetensors 4.84 GB 073f1299 download
model-ple-00006.safetensors 4.84 GB 631eede4 download
model-ple-00007.safetensors 4.84 GB 37fa44de download
model-ple-00000.safetensors 4.47 GB abd4e4ad download
model-ple-00009.safetensors 4.47 GB fb1ca31b download
model-quant-00000.safetensors 2.61 GB 73e3a036 download
model-bf16-00001.safetensors 2.29 GB ea5a6948 download
model-mtp-experts-nvfp4.safetensors 1.32 GB 8191c342 download
model-bf16-00000.safetensors 1.31 GB a603dc27 download
model-bf16-00002.safetensors 1.20 GB 3799b7d7 download
layer-00010-experts-0128-0255.safetensors 338 MB a695c4bc download
layer-00010-experts-0256-0383.safetensors 338 MB 45ba8fed download
layer-00010-experts-0384-0511.safetensors 338 MB 4faad200 download
layer-00011-experts-0128-0255.safetensors 338 MB 85ee027b download
layer-00011-experts-0256-0383.safetensors 338 MB 4a824453 download
layer-00011-experts-0384-0511.safetensors 338 MB 0b4e6f64 download
layer-00012-experts-0128-0255.safetensors 338 MB 8b3d2073 download
layer-00012-experts-0256-0383.safetensors 338 MB b338bdef download
layer-00012-experts-0384-0511.safetensors 338 MB f674ef5d download
layer-00013-experts-0128-0255.safetensors 338 MB fe1f5581 download
layer-00013-experts-0256-0383.safetensors 338 MB 83daf610 download
layer-00013-experts-0384-0511.safetensors 338 MB 8e4de574 download
layer-00014-experts-0128-0255.safetensors 338 MB 04b91607 download
layer-00014-experts-0256-0383.safetensors 338 MB 5824d598 download
layer-00014-experts-0384-0511.safetensors 338 MB 7ef853ba download
layer-00015-experts-0128-0255.safetensors 338 MB ada9efc8 download
layer-00015-experts-0256-0383.safetensors 338 MB 91a89c72 download
layer-00015-experts-0384-0511.safetensors 338 MB 48aa8159 download
layer-00016-experts-0128-0255.safetensors 338 MB f4f0b643 download
layer-00016-experts-0256-0383.safetensors 338 MB 8420febd download
layer-00016-experts-0384-0511.safetensors 338 MB 06f97049 download
layer-00017-experts-0128-0255.safetensors 338 MB 1421cc98 download
layer-00017-experts-0256-0383.safetensors 338 MB f12fc42e download
layer-00017-experts-0384-0511.safetensors 338 MB cd3236ac download
layer-00018-experts-0128-0255.safetensors 338 MB 5646eeb4 download
layer-00018-experts-0256-0383.safetensors 338 MB b8771b7e download
layer-00018-experts-0384-0511.safetensors 338 MB 9c416163 download
layer-00019-experts-0128-0255.safetensors 338 MB 32714ac5 download
layer-00019-experts-0256-0383.safetensors 338 MB eec621a2 download
layer-00019-experts-0384-0511.safetensors 338 MB 8663aa06 download
layer-00020-experts-0128-0255.safetensors 338 MB 061fb507 download
layer-00020-experts-0256-0383.safetensors 338 MB d08a45bd download
layer-00020-experts-0384-0511.safetensors 338 MB ea13e01c download
layer-00021-experts-0128-0255.safetensors 338 MB 60cff5a6 download
layer-00021-experts-0256-0383.safetensors 338 MB 367fecfe download
layer-00021-experts-0384-0511.safetensors 338 MB 94a1cbd3 download
layer-00022-experts-0128-0255.safetensors 338 MB 57614ee8 download
layer-00022-experts-0256-0383.safetensors 338 MB d18c27c3 download
layer-00022-experts-0384-0511.safetensors 338 MB b895ebcf download
layer-00023-experts-0128-0255.safetensors 338 MB 533b44c3 download
layer-00023-experts-0256-0383.safetensors 338 MB 9b526e01 download
layer-00023-experts-0384-0511.safetensors 338 MB 26fd9964 download
layer-00024-experts-0128-0255.safetensors 338 MB 5d9c3284 download
layer-00024-experts-0256-0383.safetensors 338 MB 7336fac4 download
layer-00024-experts-0384-0511.safetensors 338 MB 6359b70e download
layer-00025-experts-0128-0255.safetensors 338 MB cde77ae8 download
layer-00025-experts-0256-0383.safetensors 338 MB c1e27170 download
layer-00025-experts-0384-0511.safetensors 338 MB bdd9982c download
layer-00026-experts-0128-0255.safetensors 338 MB e4f8b86f download
layer-00026-experts-0256-0383.safetensors 338 MB 3a47f94c download
layer-00026-experts-0384-0511.safetensors 338 MB 140bdda6 download
layer-00027-experts-0128-0255.safetensors 338 MB 71529a27 download
layer-00027-experts-0256-0383.safetensors 338 MB b489a023 download
layer-00027-experts-0384-0511.safetensors 338 MB ff957329 download
layer-00028-experts-0128-0255.safetensors 338 MB 932c51b8 download
layer-00028-experts-0256-0383.safetensors 338 MB ca83818f download
layer-00028-experts-0384-0511.safetensors 338 MB 106c188d download
layer-00029-experts-0128-0255.safetensors 338 MB 1688980a download
layer-00029-experts-0256-0383.safetensors 338 MB 00871fca download
layer-00029-experts-0384-0511.safetensors 338 MB 9c0b7bf0 download
layer-00030-experts-0128-0255.safetensors 338 MB 9ec366a7 download
layer-00030-experts-0256-0383.safetensors 338 MB cc6bad44 download
layer-00030-experts-0384-0511.safetensors 338 MB 536bddea download
layer-00031-experts-0128-0255.safetensors 338 MB 26e33f17 download
layer-00031-experts-0256-0383.safetensors 338 MB c04f4b25 download
layer-00031-experts-0384-0511.safetensors 338 MB e83eecea download
layer-00032-experts-0128-0255.safetensors 338 MB 349bd598 download
layer-00032-experts-0256-0383.safetensors 338 MB 2f0b93c4 download
layer-00032-experts-0384-0511.safetensors 338 MB 507a5e90 download
layer-00033-experts-0128-0255.safetensors 338 MB 79c57fb4 download
layer-00033-experts-0256-0383.safetensors 338 MB 4cdbe450 download
layer-00033-experts-0384-0511.safetensors 338 MB 3310eb99 download
layer-00034-experts-0128-0255.safetensors 338 MB 96d9e193 download
layer-00034-experts-0256-0383.safetensors 338 MB fb837f0e download
layer-00034-experts-0384-0511.safetensors 338 MB 2d05b4fb download
layer-00035-experts-0128-0255.safetensors 338 MB be984c14 download
layer-00035-experts-0256-0383.safetensors 338 MB 26117191 download
layer-00035-experts-0384-0511.safetensors 338 MB 41607b3e download
layer-00036-experts-0128-0255.safetensors 338 MB 878a2ccc download
layer-00036-experts-0256-0383.safetensors 338 MB 63740c5d download
layer-00036-experts-0384-0511.safetensors 338 MB deb43a54 download
layer-00037-experts-0128-0255.safetensors 338 MB 4b69bc03 download
layer-00037-experts-0256-0383.safetensors 338 MB 368ff1ed download
layer-00037-experts-0384-0511.safetensors 338 MB 89894200 download
layer-00038-experts-0128-0255.safetensors 338 MB 3bf6204c download
layer-00038-experts-0256-0383.safetensors 338 MB 9de8cb51 download
layer-00038-experts-0384-0511.safetensors 338 MB 23345f11 download
layer-00039-experts-0128-0255.safetensors 338 MB 3080dde4 download
layer-00039-experts-0256-0383.safetensors 338 MB e0001a6c download
layer-00039-experts-0384-0511.safetensors 338 MB aef6ce70 download
layer-00040-experts-0128-0255.safetensors 338 MB a9ed224a download
layer-00040-experts-0256-0383.safetensors 338 MB c14bc960 download
layer-00040-experts-0384-0511.safetensors 338 MB 6d92058c download
layer-00041-experts-0128-0255.safetensors 338 MB ba168052 download
layer-00041-experts-0256-0383.safetensors 338 MB adc56bc4 download
layer-00041-experts-0384-0511.safetensors 338 MB f352a256 download
layer-00042-experts-0128-0255.safetensors 338 MB 97588568 download
layer-00042-experts-0256-0383.safetensors 338 MB edb90156 download
layer-00042-experts-0384-0511.safetensors 338 MB 86b17cb0 download
layer-00043-experts-0128-0255.safetensors 338 MB 2e94cb67 download
layer-00043-experts-0256-0383.safetensors 338 MB 71c573ff download
layer-00043-experts-0384-0511.safetensors 338 MB f0adcc45 download
layer-00044-experts-0128-0255.safetensors 338 MB 9adf5a43 download
layer-00044-experts-0256-0383.safetensors 338 MB 8caf7bd0 download
layer-00044-experts-0384-0511.safetensors 338 MB 9114fb1a download
layer-00045-experts-0128-0255.safetensors 338 MB 5871293f download
layer-00045-experts-0256-0383.safetensors 338 MB b7abc25a download
layer-00045-experts-0384-0511.safetensors 338 MB fa02dd0c download
layer-00046-experts-0128-0255.safetensors 338 MB 18b52c91 download
layer-00046-experts-0256-0383.safetensors 338 MB 534c5d00 download
layer-00046-experts-0384-0511.safetensors 338 MB cc8b231b download
layer-00047-experts-0128-0255.safetensors 338 MB 4ab14fba download
layer-00047-experts-0256-0383.safetensors 338 MB cdca158a download
layer-00047-experts-0384-0511.safetensors 338 MB 4961f9d0 download
layer-00010-experts-0000-0127.safetensors 338 MB 3abe1d32 download
layer-00011-experts-0000-0127.safetensors 338 MB fba63936 download
layer-00012-experts-0000-0127.safetensors 338 MB 18892764 download
layer-00013-experts-0000-0127.safetensors 338 MB 99a9bfbc download
layer-00014-experts-0000-0127.safetensors 338 MB 9fbdb2b6 download
layer-00015-experts-0000-0127.safetensors 338 MB ae9a7282 download
layer-00016-experts-0000-0127.safetensors 338 MB ded0d079 download
layer-00017-experts-0000-0127.safetensors 338 MB c6286104 download
layer-00018-experts-0000-0127.safetensors 338 MB 59e4aecf download
layer-00019-experts-0000-0127.safetensors 338 MB 8946447b download
layer-00020-experts-0000-0127.safetensors 338 MB 183434fd download
layer-00021-experts-0000-0127.safetensors 338 MB a8ea649c download
layer-00022-experts-0000-0127.safetensors 338 MB 5d0f0bd9 download
layer-00023-experts-0000-0127.safetensors 338 MB 08be5ba0 download
layer-00024-experts-0000-0127.safetensors 338 MB 31ce8574 download
layer-00025-experts-0000-0127.safetensors 338 MB 94eed05d download
layer-00026-experts-0000-0127.safetensors 338 MB 50db0859 download
layer-00027-experts-0000-0127.safetensors 338 MB 16c3ef9c download
layer-00028-experts-0000-0127.safetensors 338 MB a892e7b7 download
layer-00029-experts-0000-0127.safetensors 338 MB 979e7d01 download
layer-00030-experts-0000-0127.safetensors 338 MB b6405e86 download
layer-00031-experts-0000-0127.safetensors 338 MB 54312d50 download
layer-00032-experts-0000-0127.safetensors 338 MB a47fa09a download
layer-00033-experts-0000-0127.safetensors 338 MB 5534c393 download
layer-00034-experts-0000-0127.safetensors 338 MB 6bb7674b download
layer-00035-experts-0000-0127.safetensors 338 MB 7779b8e9 download
layer-00036-experts-0000-0127.safetensors 338 MB 4e14f26a download
layer-00037-experts-0000-0127.safetensors 338 MB 97791467 download
layer-00038-experts-0000-0127.safetensors 338 MB 443d28c4 download
layer-00039-experts-0000-0127.safetensors 338 MB 5d5d9428 download
layer-00040-experts-0000-0127.safetensors 338 MB 63d99598 download
layer-00041-experts-0000-0127.safetensors 338 MB 465c09b7 download
layer-00042-experts-0000-0127.safetensors 338 MB 2f791731 download
layer-00043-experts-0000-0127.safetensors 338 MB ec2ebe1e download
layer-00044-experts-0000-0127.safetensors 338 MB b2672822 download
layer-00045-experts-0000-0127.safetensors 338 MB 89f561a0 download
layer-00046-experts-0000-0127.safetensors 338 MB 0e5948ae download
layer-00047-experts-0000-0127.safetensors 338 MB b63d8ace download
layer-00000-experts-0128-0255.safetensors 338 MB c7b3dd7a download
layer-00000-experts-0256-0383.safetensors 338 MB d473f695 download
layer-00000-experts-0384-0511.safetensors 338 MB 49a11ce1 download
layer-00001-experts-0128-0255.safetensors 338 MB ada44b26 download
layer-00001-experts-0256-0383.safetensors 338 MB d4dff7fb download
layer-00001-experts-0384-0511.safetensors 338 MB 42cae3a2 download
layer-00002-experts-0128-0255.safetensors 338 MB 9f0f85a0 download
layer-00002-experts-0256-0383.safetensors 338 MB 0ee37d70 download
layer-00002-experts-0384-0511.safetensors 338 MB 7a4eccee download
layer-00003-experts-0128-0255.safetensors 338 MB 027acab8 download
layer-00003-experts-0256-0383.safetensors 338 MB 4fb4ff93 download
layer-00003-experts-0384-0511.safetensors 338 MB 8689dc73 download
layer-00004-experts-0128-0255.safetensors 338 MB 3e97111c download
layer-00004-experts-0256-0383.safetensors 338 MB de38d7f2 download
layer-00004-experts-0384-0511.safetensors 338 MB 295db691 download
layer-00005-experts-0128-0255.safetensors 338 MB b3702a71 download
layer-00005-experts-0256-0383.safetensors 338 MB 970a76e1 download
layer-00005-experts-0384-0511.safetensors 338 MB addb78d7 download
layer-00006-experts-0128-0255.safetensors 338 MB 231aee4c download
layer-00006-experts-0256-0383.safetensors 338 MB a64af6f9 download
layer-00006-experts-0384-0511.safetensors 338 MB 152a13da download
layer-00007-experts-0128-0255.safetensors 338 MB 80be18bc download
layer-00007-experts-0256-0383.safetensors 338 MB 87e5a66b download
layer-00007-experts-0384-0511.safetensors 338 MB ac06542d download
layer-00008-experts-0128-0255.safetensors 338 MB 83be8d5e download
layer-00008-experts-0256-0383.safetensors 338 MB d814f5bd download
layer-00008-experts-0384-0511.safetensors 338 MB cb1fdf27 download
layer-00009-experts-0128-0255.safetensors 338 MB 64a8d58a download
layer-00009-experts-0256-0383.safetensors 338 MB 512fac25 download
layer-00009-experts-0384-0511.safetensors 338 MB 3b8b64c7 download
layer-00000-experts-0000-0127.safetensors 338 MB 79d20243 download
layer-00001-experts-0000-0127.safetensors 338 MB 4455e722 download
layer-00002-experts-0000-0127.safetensors 338 MB 7fef76a7 download
layer-00003-experts-0000-0127.safetensors 338 MB d052a936 download
layer-00004-experts-0000-0127.safetensors 338 MB 9ad555d4 download
layer-00005-experts-0000-0127.safetensors 338 MB 6269fe77 download
layer-00006-experts-0000-0127.safetensors 338 MB 1af4b4d8 download
layer-00007-experts-0000-0127.safetensors 338 MB 26a9f72b download
layer-00008-experts-0000-0127.safetensors 338 MB 29da7f6d download
layer-00009-experts-0000-0127.safetensors 338 MB dc7c8975 download
model.safetensors.index.json 34.4 MB 2b2e87a8 download
tokenizer.json 12.2 MB 0997f410 download
config.json 9.41 MB c9ec6aaf download
hf_quant_config.json 9.38 MB 9585a5e5 download
vocab.json 6.41 MB 0aa0ce06 download
merges.txt 3.20 MB a494e019 download
MANIFEST.json 17.7 KB 9fce3720 download
tokenizer_config.json 17.5 KB 5de744b3 download
README.md 10.8 KB 1348aa02 download
chat_template.jinja 8.74 KB c0c686f9 download
LICENSE 3.16 KB 9557a896 download
.gitattributes 1.60 KB a09db2ea download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
VALIDATION.txt 238 B a694a4f4 download
generation_config.json 202 B 023756cf download

README current version from Hugging Face


license: other
license_name: qwen-community-1.0
license_link: LICENSE
base_model: orcarouter/Qwen3.8-Flash-Next-Uncensored
base_model_relation: quantized
pipeline_tag: image-text-to-text
tags:

  • nvfp4
  • fp8
  • modelopt
  • sglang
  • blackwell
  • uncensored
  • abliterated

Qwen3.8-Flash-Next-Uncensored-NVFP4 (beginning-ai)

A mixed-precision quantization of
orcarouter/Qwen3.8-Flash-Next-Uncensored
(revision e096800036ec20da7e2442dcd4044a004d4e99fa), made to run on one 96 GB
Blackwell GPU
with SGLang. That model is an abliterated (refusal-removed) BF16 build
of Qwen/Qwen3.8-Flash-Next.

This checkpoint has the same structure, formats and file layout as our censored
quant, beginning-ai/Qwen3.8-Flash-Next-NVFP4.
The two can be swapped by changing only the model path. Its activation scales were
measured again on this model.

  • 120 GiB on disk. The BF16 source is 335 GiB.
  • About 78 GB of weights on the GPU. The 47.7 GiB n-gram PLE table stays in host
    RAM or on NVMe.
  • On one RTX PRO 6000, single-stream decode runs at about 265 tok/s at 2K
    context and about 240 tok/s at 200K
    , with a 786K-token KV cache and 524K
    context
    .
  • Nothing is removed: the vision tower and the multi-token-prediction (MTP)
    draft layer are kept.

Warning. The source model has had its refusal behavior removed. It will follow
harmful, unethical or illegal requests that the original model refuses, and it has
no built-in guardrails. Use it only where you add your own safeguards, and follow
the license and the laws that apply to you. We did not change or measure its
refusal behavior; see the source model card.

How this differs from the source and the base

From Qwen's base to orcarouter's source. We compared every tensor with the
official base. orcarouter changed 149 of the 1,658 tensors, all of them weights
that write into the residual stream:

  • the output projections of all 36 linear-attention layers and 12 sparse-attention layers;
  • the routed-expert and shared-expert down projections in all 48 layers;
  • the token embeddings and one PLE projection;
  • 3 tensors in the MTP layer.

The router, the norms, the vision tower, the output head and the 47.7 GiB n-gram
table are byte-identical to the base.

From orcarouter's source to this checkpoint:

Part Source This checkpoint Size here
Routed experts (48 layers × 512) mlp.experts.*.{gate,up,down}_proj BF16 NVFP4 (4-bit weights and activations, group 16) 68.0 GiB
Shared experts mlp.shared_expert.* BF16 NVFP4 weights, BF16 activations (W4A16_NVFP4, group 16) 0.13 GiB
Linear-attention projections (36 layers) linear_attn.{in_proj_qkv,in_proj_z,out_proj} BF16 FP8, 128×128 block scales, dynamic activations (FP8_PB_WO) 1.93 GiB
Sparse-attention projections (12 layers) self_attn.{q,k,v,o}_proj BF16 FP8, 128×128 blocks (FP8_PB_WO) 0.65 GiB
MTP draft layer routed experts mtp.layers.0.mlp.experts BF16 NVFP4 (group 16) 1.49 GiB (whole MTP layer)
N-gram PLE table BF16 FP8, copied byte-for-byte from Qwen/Qwen3.8-Flash-Next-FP8 (the table is unchanged in the source) 47.7 GiB
Output head, embeddings, router, shared-expert gate, norms, linear-attention state parameters, hyper-connections, rest of the MTP layer, vision tower BF16 BF16, unchanged (byte-for-byte) 4.6 GiB (plus the MTP layer's BF16 part)

The full per-layer list is in hf_quant_config.json (quant_algo: MIXED_PRECISION).
Byte counts and the calibration report are in MANIFEST.json.

How it was made

  • Quantizer. NVIDIA ModelOpt 0.47.0 tensor quantizers (NVFP4QTensor,
    FP8QTensor), applied to orcarouter's BF16 weights.
  • Shared scale for gate and up. SGLang runs each expert's gate_proj and
    up_proj as one fused GEMM with one global scale. So both halves are quantized
    against one shared weight_scale_2, as ModelOpt does for fused layers. A build
    check confirms that all 24,624 fused pairs, plus the 512 MTP experts, have equal
    scales.
  • Activation scales, measured on this model. We served a first build of this
    model in SGLang with a forward hook on every routed-MoE layer. The hook recorded the
    MoE inputs and routing for 16 real coding-agent prompts (332,244 tokens) and saved
    100,123 rows per layer. From that capture:
    • gate_proj/up_proj: one scale per layer, from the largest MoE input over
      every token.
    • down_proj: one scale per expert, from that expert's own activations
      (silu(gate) * up), computed with the BF16 weights on the tokens routed to it.
      24,184 of the 24,576 experts had at least 32 routed tokens; the median is about
      1,350. The rest take the larger of their own maximum and the layer's 95th
      percentile.
    • MTP experts: one scale per layer, from the same capture.
    • No scales were taken from another release.
  • No quantization of the output head. It stays BF16, as in NVIDIA's and
    RadixArk's releases of the base. A 4-bit head changed the top token at about 20%
    of positions in our tests.

Results

Test machine

GPU 1 × NVIDIA RTX PRO 6000 Blackwell Workstation Edition (96 GB, SM120), driver 580
CPU AMD Ryzen 9 9950X3D (16 cores)
Host RAM 64 GB. The PLE table does not fit, so it streams from NVMe.
Storage 1 × PCIe Gen4 NVMe (2 TB) for the weights, the PLE table and the KV disk cache
Software beginning-ai/sglang tag qwen38-fn-nvfp4-2026-10-10 (see Serving); FP8 KV cache; MTP speculative decoding (3 steps, 4 draft tokens) with rejection sampling and a 64K-token draft vocabulary

Correctness checks

Check This checkpoint Our censored quant
Conversation replay: real DeepSWE agent turns at 8K–200K context that produce a valid tool call 120 of 120 120 of 120
Decode compared with prefill: severe differences (decode log-prob > −2 but prefill < −10) 0 of 5,032 tokens 0 of 6,357 tokens
Sampled tokens outside prefill's top 20 0.10% 0.14%

Coding agent: DeepSWE smoke test

DeepSWE v1.1 with the mini-swe-agent harness,
reasoning effort medium, 524K context. We ran 4 tasks that our censored quant
solves
, one attempt each, to check that the model works as an agent. This is a
smoke test, not a benchmark score.

Task Result
etree-xml-diff-patch solved
fd-deterministic-multi-key-sorting solved
opa-template-string-reconstruction solved
httpx-deterministic-cookie-store not solved: 114 of 115 new-feature tests and all 1,281 existing tests passed

Speed

All numbers are for one stream on the test machine above. The prompts are real
coding-agent prompts. Each run generates 1,024 tokens with temperature 1.0, top_p
0.95 and top_k 20.

Context Decode (tok/s) Our censored quant
2K 266 268
64K 246 262
200K 241 233–240

Decode is the median of 4 seeds at 2K and 64K. At 200K it is the total over 4
prompts × 4 seeds, with a warm page cache.

Prompt Time to first token
2K tokens, not cached 0.18 s
32K tokens, not cached 2.7 s
128K tokens, not cached 11.8 s
Agent turn: 2K new tokens after a cached 128K prefix 0.23 s
Capacity
KV cache on the GPU (FP8) 786,432 tokens
Context length 524,288 tokens (YaRN factor 2 over the native 262,144)

Serving

The results above were measured with this exact SGLang build:

pip install -e "git+https://github.com/beginning-ai/sglang@93540334d2ce687f0677bebb42dbd252b218e92f#egg=sglang&subdirectory=python"

We have not tested this checkpoint on unpatched SGLang.

python3 -m sglang.launch_server \
  --model-path beginning-ai/Qwen3.8-Flash-Next-Uncensored-NVFP4 \
  --trust-remote-code \
  --quantization modelopt_fp4 \
  --dtype bfloat16 \
  --kv-cache-dtype fp8_e4m3 \
  --page-size 64 \
  --context-length 262144 \
  --linear-attn-decode-backend flashinfer \
  --linear-attn-prefill-backend flashinfer \
  --mamba-radix-cache-strategy extra_buffer \
  --speculative-algorithm NEXTN \
  --speculative-num-steps 3 \
  --speculative-eagle-topk 1 \
  --speculative-num-draft-tokens 4 \
  --reasoning-parser qwen3 \
  --tool-call-parser qwen3_coder

Notes:

  • PLE table on NVMe. If the 47.7 GiB PLE table does not fit in host RAM, as on
    our 64 GB machine, add --ple-offload-backend nvme (patched SGLang).
  • 524K context. Add --context-length 524288 and set
    SGLANG_ALLOW_OVERWRITE_LONGER_CONTEXT_LEN=1. Also pass
    --json-model-override-args with a YaRN rope_parameters block: rope_type: "yarn", factor: 2.0, original_max_position_embeddings: 262144, and the base
    config's mrope_* / rope_theta / partial_rotary_factor values.
  • Decode kernels. SGLANG_FP8_ROW_WEIGHTS=1 (FP8 output head and FP8 draft
    projections, converted at load time) and SGLANG_NVFP4_DECODE_MOE=1 turn on the
    patched decode kernels.

Files

File Content
layer-*-experts-*.safetensors Routed experts (NVFP4)
model-quant-00000.safetensors FP8 attention projections, NVFP4 shared experts
model-mtp-experts-nvfp4.safetensors MTP draft-layer routed experts (NVFP4)
model-ple-*.safetensors N-gram PLE table (FP8, from Qwen's FP8 release)
model-bf16-*.safetensors All BF16 tensors
MANIFEST.json Provenance, byte accounting, calibration report
VALIDATION.txt Layout and fused-scale checks run at build time

License and credits

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration