license: other
license_name: swift-open-license-v1.0
base_model:
- ajgazin/Swift-1.5-Qwen3.8-27B-Uncensored-MTP
- ajgazin/Swift-1.5-Qwen3.8-27B-Uncensored-NVFP4
- z-lab/Qwen3.8-27B-DFlash2
tags: - ninfer
- nvfp4
- fp8
- qwen3.8
- speculative-decoding
language: - en
Swift 1.5 Qwen3.8 27B Uncensored, NInfer NVFP4 artifact
NInfer v3 artifact of ajgazin/Swift-1.5-Qwen3.8-27B-Uncensored-MTP
for the NInfer engine (Blackwell, sm_120a). Text, vision,
MTP and DFlash2 components, 64 calibrated NVFP4 MLP layers, FP8 everywhere else, one file.
| File | swift15_qwen3_8_27b_uncensored_nvfp4.ninfer, 22,783,241,220 bytes |
| sha256 | 81ccce0b3c9fa5a778f3081e39ae0f84aaa9c1a75f67547c49e37bbd3fdf0340 |
| Artifact name | swift-1.5-qwen3.8-27b-uncensored |
| Components | text, vision, mtp, dflash2 (+ indexed Q4 proposal head, 131,072 rows) |
| Engine | NInfer 9e163eee (converter and target) |
| Chat template | NInfer tools/chat_templates/qwen3_8.jinja |
Sources
| Repo | Revision | Role |
|---|---|---|
ajgazin/Swift-1.5-Qwen3.8-27B-Uncensored-MTP (BF16) |
dd0744c4d5936b358f16ab2c9556e64e3845141e |
base: embeddings, norms, convolutions, GDN a/b, MTP, vision; FP8 projections and the output head are encoded from it |
ajgazin/Swift-1.5-Qwen3.8-27B-Uncensored-NVFP4 (ModelOpt 0.47.0rc0) |
e4899aa5eb4d138aeefd41ad54da7c5a12a9560f |
the 64 layers of calibrated NVFP4 MLP codes, block scales and static input scales, imported verbatim |
z-lab/Qwen3.8-27B-DFlash2 |
50307d4c4cde6860d4eee73e2547cd786fe8e8a4 |
DFlash2 drafter (trained on base Qwen3.8) |
Precision
| Tensors | Format | From |
|---|---|---|
| MLP gate/up/down, all 64 layers | NVFP4 (block 16, E4M3 block scales, calibrated) | ModelOpt export |
| Attention q/k/v/o and GDN projections (208) | FP8 E4M3 with one BF16 scale per row (fp8_row_maxabs) |
BF16 checkpoint |
| Token embedding, output head | FP8 row-scaled | BF16 checkpoint |
| Proposal head | Q4 g64, 131,072 rows | BF16 output head |
| MTP, DFlash2 | Q8 g32 projections, BF16 norms | BF16 checkpoint, DFlash2 repo |
| Vision | NInfer defaults (Q4/Q5/Q6 projections, BF16 rest) | BF16 checkpoint |
Format counts: nvfp4 128, fp8_e4m3fn_row_bf16 130, q4_g64_fp16 55, q5_g64_fp16 54, q6_g64_fp16 1,
q8_g32_fp16 28, bf16 579, fp32 288, int32 1. Same allocation as the Swift 1.0 artifact built with
the same recipe; the recipe, check scripts and logs are in conversion/.
Run
ninfer-serve --artifact swift15_qwen3_8_27b_uncensored_nvfp4.ninfer --kv-dtype int8 --max-context 32768
ninfer-serve ... --spec mtp --draft-tokens 3 --lm-head-draft
ninfer-serve ... --spec dflash2 --draft-tokens 7 --lm-head-draft
ninfer ... --vision
Resident weights on a 32 GB GPU: about 18.1 GiB (no speculation), 18.9 GiB (MTP), 20.5 GiB
(DFlash2), from the Swift 1.0 artifact with the same allocation.
Reasoning effort: the template honours reasoning_effort medium and low and enable_thinking: false; with no effort set it emits Qwen's default instruction. high, minimal and max are
rejected by the template (HTTP 400).
Validation
Python checks on the conversion pod (logs in conversion/logs/): the NVFP4 objects are
byte-identical to the ModelOpt tensors, the FP8 rows reproduce the BF16 rows to the inherent E4M3
error (relative L2 about 2.6 %), the output head is closer to the BF16 head than the NVFP4 decode
is, and the chat template renders byte-identically to the stock Swift template for effort unset,
medium, low and thinking off. Engine smoke tests and the speculative-decoding benchmark matrix have not been run on this artifact yet; the Swift 1.0 artifact built by the same recipe passed them at NInfer 9e163eee.
License
Swift Open License v1.0, as the source checkpoints. Qwen3.8 base model terms apply.