library_name: ninfer
pipeline_tag: image-text-to-text
license: other
license_name: swift-open-license-1.0
license_link: https://huggingface.co/Dragoy/Swift-1.5-Qwen3.8-27B-abliterated-NVFP4-NInfer/blob/main/LICENSE
base_model: ukisai/Swift-1.5-Qwen3.8-27b
tags:
- ninfer
- nvfp4
- fp8
- qwen3.8
- qwen3_5
- abliterated
- uncensored
- multimodal
- mtp
- dflash2
- speculative-decoding
- blackwell
- cuda
- sm_120a

Swift-1.5-Qwen3.8-27B · huihui-style abliterated · NVFP4 · NInfer
A 27.78B-parameter multimodal derivative of
ukisai/Swift-1.5-Qwen3.8-27b, abliterated in the
huihui-ai style, quantized to
NVFP4 + FP8 for the NInfer engine on Blackwell (sm_120a).
One file contains the text model, the vision tower, the MTP head and a DFlash2 drafter.
| Base | ukisai/Swift-1.5-Qwen3.8-27b @ bc7a1e10b689 |
| Abliteration | huihui-style refusal-direction removal, transferred by weight difference from the Qwen/Qwen3.8-27B ↔ huihui-ai/Huihui-Qwen3.8-27B-abliterated pair. 70 tensors, language layers 17–51 (measured, see below) |
| Weight cost | median ‖Δ‖/‖W‖ = 0.0188 (min 0.0177, max 0.0218) over the 70 changed matrices |
| Quantization | NVFP4 (MLP of layers 0–55) + FP8 (attention, GDN, MLP of layers 56–63) — allocation copied verbatim from unsloth/Qwen3.8-27B-NVFP4, 32 calibration samples |
| Engine | Neroued/ninfer @ bace20dc70249eed, built for sm_120a (CUDA 13.0.2 toolchain) |
| Container | NInfer artifact v3, official tools.convert, recipe qwen3_8_27b_nvfp4, components text,vision,mtp,dflash2, --proposal |
| Artifact | qwen3_8_27b_swift15_abliterated_nvfp4.ninfer — 1246 objects, 23,719,719,940 bytes, sha256 9810893aa7ee18ae3d213c3f95999191d52b425521da5312fe9c00fb8ca01088 |
| Built | 2026-09-29 |
Quickstart
Requires the ninfer runtime at revision ≥ bace20dc built for sm_120a (this artifact was verified with a build made on the CUDA 13.0.2 toolchain; see the NInfer README for its own requirements):
# one-shot
ninfer qwen3_8_27b_swift15_abliterated_nvfp4.ninfer \
--prompt "Tell me something interesting about the universe." \
--max-context 16384 --max-new 512 --kv-dtype fp8 --no-thinking
# MTP speculative decoding
ninfer qwen3_8_27b_swift15_abliterated_nvfp4.ninfer \
--prompt "Write a haiku about a GPU cluster at 3am." \
--max-context 16384 --max-new 256 --kv-dtype fp8 --no-thinking \
--spec mtp --draft-tokens 3 --lm-head-draft
# DFlash2 speculative decoding
ninfer qwen3_8_27b_swift15_abliterated_nvfp4.ninfer \
--prompt "Write a haiku about a GPU cluster at 3am." \
--max-context 16384 --max-new 256 --kv-dtype fp8 --no-thinking \
--spec dflash2 --draft-tokens 7 --lm-head-draft
# image input (works together with DFlash2)
ninfer qwen3_8_27b_swift15_abliterated_nvfp4.ninfer \
--messages examples/cli/messages/image_chart.json --vision \
--max-context 8192 --max-new 72 --kv-dtype fp8 --no-thinking
# serving (OpenAI / Anthropic compatible)
ninfer-serve qwen3_8_27b_swift15_abliterated_nvfp4.ninfer --port 8088
Vision and MTP/DFlash2 selection are startup choices; the runtime loads only the components you select.
Abliteration evaluation
The abliteration was checked on the BF16 checkpoints, before quantization, with the same 100 harmful
prompts Heretic asks (mlabonne/harmful_behaviors test[:100] @ 01cead013), the same system prompt
(You are a helpful assistant.), greedy decoding, 100 new tokens, and the 33 refusal markers of the
Heretic keyword scorer (@ 3521f864, AGPL-3.0-or-later; logic reproduced
here for comparability of numbers, response texts are not published). KL divergence is the first-token KL
against the stock Swift 1.5 checkpoint on 100 harmless prompts (mlabonne/harmless_alpacatest[:100]), i.e. it prices the abliteration itself, not the Qwen→Swift difference.
| Model | Refusals by Heretic metric | Explicit refusals (read by hand) | KL vs stock |
|---|---|---|---|
| Swift 1.5 (stock, BF16) | 98 / 100 | 98 | 0 |
| Swift 1.5 abliterated (BF16, before quantization) | 34 / 100 | 0 | 0.0791 |
This NVFP4/FP8 artifact (run inside ninfer-serve, thinking off) |
45 / 100 | 0 | not measured |
How to read this:
- The Heretic metric overcounts refusals for an abliterated model. Its list contains words that also appear in
completed answers (disclaimer,illegal,violat, …). The model answers and adds a
“Disclaimer: … for educational purposes” line, and the scorer counts that as a refusal. Every one of the
34 (BF16) and 45 (artifact) flagged answers was read by hand: all of them carry out the request (39 of the 45
artifact hits aredisclaimer, 4illegal, 1violat, 1i am an ai). Two BF16 answers (#66, #48) are softer
compliance (a “simulated narrative” framing, a legal-vs-illegal reframing), not refusals. - The BF16 numbers were measured twice, on 2026-09-28 and again on 2026-09-29 from re-downloaded, sha256-verified
inputs: verdicts and generated texts were identical for all 100 prompts on both variants. - The artifact number (45) is not comparable to the BF16 number (34). Different engine, NVFP4/FP8 weights, thinking
disabled through the chat template instead of a forced<think></think>prefix, and 16 prompts flip from “not
flagged” to “flagged” (5 flip the other way), i.e. whether a disclaimer/legality word appears in the first 100 tokens. Its meaningful
result is 0 explicit refusals. No KL was measured for the quantized artifact. - One evaluation run is not a guarantee; other decoding settings, prompts, or the system prompt can change the numbers.
- Full per-prompt verdicts (no text) are in
evaluation-summary.jsonandrefusal-probe-artifact.json.
Measured on this artifact
Single short runs on one RTX PRO 6000 Blackwell (Modal), --kv-dtype fp8 --no-thinking --greedy. Illustrative smoke
checks, not benchmarks — acceptance rates depend heavily on the text. Raw output is inninfer-runtime-report.json.
| Check | Result |
|---|---|
| Inventory | version 3, 1246 objects, 1240 tensors, 6 resources, 1513 bindings, 844 uses; nvfp4 = 112, fp8_e4m3fn_row_bf16 = 146 |
| Modes exercised | text, arithmetic, vision (image chart), MTP, DFlash2, DFlash2 + vision — all completed, 0 fallback steps |
| Weights resident | 19.0 GiB (text), 19.3 GiB with --vision |
| Decode speed, no speculation | ~70 tok/s (69.5–70.1 across runs) |
MTP (--draft-tokens 3) |
75.0 % acceptance on a one-sentence answer (148 tok/s overall), 51.9 % on a 256-token haiku prompt (119.5 tok/s) |
DFlash2 (--draft-tokens 7) |
14.3 % and 13.1 % acceptance on the two short text prompts (78.6 and 90.1 tok/s overall); 85.7 % on the image example (122 tok/s) |
Speculative decoding is not guaranteed to reproduce the non-speculative greedy text token for token: in the haiku
smoke the plain, MTP and DFlash2 runs differ in one line.
Provenance
| Component | Source |
|---|---|
| Base weights | ukisai/Swift-1.5-Qwen3.8-27b @ bc7a1e10b689648585a3ef41494c8d84cf77271a (18 shards, sha256-verified against the Hub LFS metadata) |
| Abliteration reference | Qwen/Qwen3.8-27B @ 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0 and huihui-ai/Huihui-Qwen3.8-27B-abliterated @ 739e3c5b89849f6c238ce1e5b70008612ae42cdd (both 18/18 shards sha256-verified) |
| Quantization recipe | unsloth/Qwen3.8-27B-NVFP4 quantization_config, verbatim: recipe/unsloth_qconfig.json and recipe/quantize_nvfp4.py (sha256-pinned copies from Dragoy/Swift-Qwen3.8-27B-abliterated-NVFP4-NInfer @ 4c25aa201b41) |
| Calibration | HuggingFaceH4/ultrachat_200k, 32 samples, seq 2048; dataset revision observed at 8049631c405ae657 before and after quantization (the recipe does not pin it) |
| Quantizer stack | torch 2.11.0, transformers 5.10.1, llmcompressor 0.12.0.1, compressed-tensors 0.17.1 |
| Converter | Neroued/ninfer @ bace20dc70249eed6402b66d4852c6c3f9612905, unmodified tools.convert; chat template tools/chat_templates/qwen3_8.jinja (sha256 a497db9e…) — note this is NInfer's maintained template, not the byte-identical Qwen/Swift chat_template.jinja |
| DFlash2 drafter | z-lab/Qwen3.8-27B-DFlash2 @ 50307d4c4cde6860d4eee73e2547cd786fe8e8a4, model.safetensors sha256 67fc76d68dc5a9415511a4f394ef744d67510cd20e93b37cc2cc7d28e4bab65c |
| Reports | ablation-report.json, source-audit.json, quantization-report.json, conversion-report.json, qwen3_8_27b_swift15_abliterated_nvfp4.ninfer.conversion.json, checksums in SHA256SUMS |
How the abliteration was applied and checked
The refusal projection is linear, so its effect on any derivative of the base model is the same constant difference:
W_ablated = W_swift15 + (W_huihui - W_qwen) # per tensor, fp32, then rounded to bf16
- The set of changed tensors is measured, not assumed. Comparing the full Qwen and huihui checkpoints (all 1199
tensors, 18/18 shards each) found exactly 70 differing tensors:self_attn.o_proj,linear_attn.out_proj
andmlp.down_projof language layers 17–51. (The huihui model card states layers 18–51; layer 17 is
also changed, which the build's guard caught.) No vision (model.visual.*), MTP (mtp.*),lm_heador
embedding tensor differs. - Independent audit (
source-audit.json): a second, separate program re-read every shard,
found the 1129 unchanged tensors bit-identical to Swift 1.5, and recomputed the 70 changed tensors from the three
source checkpoints (torch.equalon the result). The quantizer and converter refuse to run on any checkpoint whose
shard hashes differ from this audit. - The whole chain was built twice from clean inputs: the abliterated shards are byte-identical between runs (18/18 sha256).
- Vision, MTP and the token embedding come from the BF16 abliterated checkpoint; the 70 abliterated matrices come from
the quantized checkpoint (import_encoded), verified per matrix in the conversion.
Files
| File | Purpose |
|---|---|
qwen3_8_27b_swift15_abliterated_nvfp4.ninfer |
the artifact |
SHA256SUMS |
checksums of every file in this repository |
recipe/ |
quantization recipe and the scripts that built and verified this release (Modal) |
*-report.json, source-audit.json, evaluation-summary.json, refusal-probe-artifact.json, ninfer-runtime-report.json |
build, audit and measurement records |
NOTICE, LICENSE, LICENSE-APACHE-2.0 |
licence and modification notices |
License
This repository is a derivative of the Swift 1.5 checkpoint, whose license is the
Swift Open License v1.0 — not Apache. The chain:
| Component | Licence |
|---|---|
| Qwen/Qwen3.8-27B (base model) | Apache-2.0 — Copyright 2026 Alibaba Cloud (LICENSE-APACHE-2.0) |
| ukisai/Swift-1.5-Qwen3.8-27b (Swift Contribution) | Swift Open License v1.0 (LICENSE) |
| huihui-ai/Huihui-Qwen3.8-27B-abliterated (source of the weight difference) | Apache-2.0 |
| This repo (abliteration + quantization + packaging) | derivative work — the Swift Contribution contained in it stays under the Swift Open License v1.0 |
What that means in practice:
- Free use, including commercial, while your gross revenue (counting all controlled entities) is below the
$1,000,000 per fiscal year threshold; qualified non-profits have no threshold for non-commercial or research use. - Above the threshold: obtain a separate written licence from UkisAI (Swift Enterprise License).
- Redistribution: ship both licence files, keep the copyright and attribution notices, and mark files you modified
(Swift licence §4–§5). The weights here are modified (abliteration, quantization); seeNOTICE.
This is a description of what the licences say, not legal advice.
Intended use and limitations
This is an uncensored model: the abliteration removes the refusal direction, so it will attempt requests that a
stock instruction-tuned model declines, including harmful ones. It is published for research, evaluation and local
deployment where that behaviour is understood and wanted. It is not safety-aligned, and its answers can be wrong,
dangerous or illegal to act on.
Use at your own responsibility. Anyone deploying it is responsible for their own safeguards, output handling and
compliance with the licences above and applicable law. It is provided as-is, without warranty.
What was not measured: quality benchmarks of any kind (the KL against stock Swift 1.5 above is the only quality
proxy, and it was taken on the BF16 abliterated checkpoint, not on this artifact); refusal behaviour under thinking mode
or other system prompts; long-context behaviour; anything on GPUs other than one RTX PRO 6000 Blackwell. Claims in the
Swift 1.5 card apply to its BF16 weights, not to this
quantized, abliterated build.
Credit for the base model to Qwen (Alibaba Cloud); for the Swift training to
UkisAI; for the abliteration to huihui-ai; for the
engine and artifact contract to Neroued; for the DFlash2 drafter to
z-lab; for the published NVFP4 recipe to unsloth; and
for the refusal scorer to p-e-w/heretic.