← back to catalog · registered 2026-09-28 23:57

Dragoy/Swift-1.5-Qwen3.8-27B-abliterated-NVFP4-NInfer

Dragoy 27B multimodal
curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/Dragoy%2FSwift-1.5-Qwen3.8-27B-abliterated-NVFP4-NInfer"
Response includes
  • classification m-uncensored
  • files 18
  • author_summary 5 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-09-28
Downloads over time
Now0→from0↑0%
00110 on Sep 280 on Sep 29Sep
Sep 28 → Sep 29 · 2 snapshots · spans 1 day

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
other
Tags
ninfer nvfp4 fp8 qwen3.8 qwen3_5 abliterated uncensored multimodal mtp dflash2 speculative-decoding blackwell

Related

Total size
0 B
Files
18
Quantizations
1
Registered
2026-09-28 23:57
Last updated on HF
2026-09-28 23:52

Files by quantization

Auxiliary files 18 files 22.1 GB
qwen3_8_27b_swift15_abliterated_nvfp4.ninfer 22.1 GB 9810893a download
banner.png 1.93 MB 8a1497e8 download
qwen3_8_27b_swift15_abliterated_nvfp4.ninfer.conversion.json 565 KB 0c834220 download
ablation-report.json 20.2 KB e66a0334 download
evaluation-summary.json 19.7 KB 58c30784 download
ninfer-runtime-report.json 19.5 KB d4afadd6 download
conversion-report.json 15.5 KB 6a06786e download
README.md 14.8 KB 2931c063 download
refusal-probe-artifact.json 13.1 KB 5b5c7bae download
LICENSE 13.0 KB adfde1bf download
LICENSE-APACHE-2.0 11.3 KB f938136e download
quantization-report.json 8.03 KB 8d6df522 download
source-audit.json 6.84 KB 91299f29 download
source-shards-lfs.json 3.83 KB e6901637 download
NOTICE 3.07 KB 5fb32f08 download
SHA256SUMS 2.18 KB fc672ec9 download
swift15-shards-lfs.json 1.87 KB 0aaaff4e download
.gitattributes 1.61 KB 3eb78eaa download

README current version from Hugging Face


library_name: ninfer
pipeline_tag: image-text-to-text
license: other
license_name: swift-open-license-1.0
license_link: https://huggingface.co/Dragoy/Swift-1.5-Qwen3.8-27B-abliterated-NVFP4-NInfer/blob/main/LICENSE
base_model: ukisai/Swift-1.5-Qwen3.8-27b
tags:

  • ninfer
  • nvfp4
  • fp8
  • qwen3.8
  • qwen3_5
  • abliterated
  • uncensored
  • multimodal
  • mtp
  • dflash2
  • speculative-decoding
  • blackwell
  • cuda
  • sm_120a

Swift-1.5-Qwen3.8-27B · NVFP4 · NInfer

Swift-1.5-Qwen3.8-27B · huihui-style abliterated · NVFP4 · NInfer

A 27.78B-parameter multimodal derivative of
ukisai/Swift-1.5-Qwen3.8-27b, abliterated in the
huihui-ai style, quantized to
NVFP4 + FP8 for the NInfer engine on Blackwell (sm_120a).
One file contains the text model, the vision tower, the MTP head and a DFlash2 drafter.

Base ukisai/Swift-1.5-Qwen3.8-27b @ bc7a1e10b689
Abliteration huihui-style refusal-direction removal, transferred by weight difference from the Qwen/Qwen3.8-27B ↔ huihui-ai/Huihui-Qwen3.8-27B-abliterated pair. 70 tensors, language layers 17–51 (measured, see below)
Weight cost median ‖Δ‖/‖W‖ = 0.0188 (min 0.0177, max 0.0218) over the 70 changed matrices
Quantization NVFP4 (MLP of layers 0–55) + FP8 (attention, GDN, MLP of layers 56–63) — allocation copied verbatim from unsloth/Qwen3.8-27B-NVFP4, 32 calibration samples
Engine Neroued/ninfer @ bace20dc70249eed, built for sm_120a (CUDA 13.0.2 toolchain)
Container NInfer artifact v3, official tools.convert, recipe qwen3_8_27b_nvfp4, components text,vision,mtp,dflash2, --proposal
Artifact qwen3_8_27b_swift15_abliterated_nvfp4.ninfer — 1246 objects, 23,719,719,940 bytes, sha256 9810893aa7ee18ae3d213c3f95999191d52b425521da5312fe9c00fb8ca01088
Built 2026-09-29

Quickstart

Requires the ninfer runtime at revision ≥ bace20dc built for sm_120a (this artifact was verified with a build made on the CUDA 13.0.2 toolchain; see the NInfer README for its own requirements):

# one-shot
ninfer qwen3_8_27b_swift15_abliterated_nvfp4.ninfer \
  --prompt "Tell me something interesting about the universe." \
  --max-context 16384 --max-new 512 --kv-dtype fp8 --no-thinking

# MTP speculative decoding
ninfer qwen3_8_27b_swift15_abliterated_nvfp4.ninfer \
  --prompt "Write a haiku about a GPU cluster at 3am." \
  --max-context 16384 --max-new 256 --kv-dtype fp8 --no-thinking \
  --spec mtp --draft-tokens 3 --lm-head-draft

# DFlash2 speculative decoding
ninfer qwen3_8_27b_swift15_abliterated_nvfp4.ninfer \
  --prompt "Write a haiku about a GPU cluster at 3am." \
  --max-context 16384 --max-new 256 --kv-dtype fp8 --no-thinking \
  --spec dflash2 --draft-tokens 7 --lm-head-draft

# image input (works together with DFlash2)
ninfer qwen3_8_27b_swift15_abliterated_nvfp4.ninfer \
  --messages examples/cli/messages/image_chart.json --vision \
  --max-context 8192 --max-new 72 --kv-dtype fp8 --no-thinking

# serving (OpenAI / Anthropic compatible)
ninfer-serve qwen3_8_27b_swift15_abliterated_nvfp4.ninfer --port 8088

Vision and MTP/DFlash2 selection are startup choices; the runtime loads only the components you select.

Abliteration evaluation

The abliteration was checked on the BF16 checkpoints, before quantization, with the same 100 harmful
prompts Heretic asks (mlabonne/harmful_behaviors test[:100] @ 01cead013), the same system prompt
(You are a helpful assistant.), greedy decoding, 100 new tokens, and the 33 refusal markers of the
Heretic keyword scorer (@ 3521f864, AGPL-3.0-or-later; logic reproduced
here for comparability of numbers, response texts are not published). KL divergence is the first-token KL
against the stock Swift 1.5 checkpoint on 100 harmless prompts (mlabonne/harmless_alpaca
test[:100]), i.e. it prices the abliteration itself, not the Qwen→Swift difference.

Model Refusals by Heretic metric Explicit refusals (read by hand) KL vs stock
Swift 1.5 (stock, BF16) 98 / 100 98 0
Swift 1.5 abliterated (BF16, before quantization) 34 / 100 0 0.0791
This NVFP4/FP8 artifact (run inside ninfer-serve, thinking off) 45 / 100 0 not measured

How to read this:

  • The Heretic metric overcounts refusals for an abliterated model. Its list contains words that also appear in
    completed answers (disclaimer, illegal, violat, …). The model answers and adds a
    “Disclaimer: … for educational purposes” line, and the scorer counts that as a refusal. Every one of the
    34 (BF16) and 45 (artifact) flagged answers was read by hand: all of them carry out the request (39 of the 45
    artifact hits are disclaimer, 4 illegal, 1 violat, 1 i am an ai). Two BF16 answers (#66, #48) are softer
    compliance (a “simulated narrative” framing, a legal-vs-illegal reframing), not refusals.
  • The BF16 numbers were measured twice, on 2026-09-28 and again on 2026-09-29 from re-downloaded, sha256-verified
    inputs: verdicts and generated texts were identical for all 100 prompts on both variants.
  • The artifact number (45) is not comparable to the BF16 number (34). Different engine, NVFP4/FP8 weights, thinking
    disabled through the chat template instead of a forced <think></think> prefix, and 16 prompts flip from “not
    flagged” to “flagged” (5 flip the other way), i.e. whether a disclaimer/legality word appears in the first 100 tokens. Its meaningful
    result is 0 explicit refusals. No KL was measured for the quantized artifact.
  • One evaluation run is not a guarantee; other decoding settings, prompts, or the system prompt can change the numbers.
  • Full per-prompt verdicts (no text) are in evaluation-summary.json and
    refusal-probe-artifact.json.

Measured on this artifact

Single short runs on one RTX PRO 6000 Blackwell (Modal), --kv-dtype fp8 --no-thinking --greedy. Illustrative smoke
checks, not benchmarks — acceptance rates depend heavily on the text. Raw output is in
ninfer-runtime-report.json.

Check Result
Inventory version 3, 1246 objects, 1240 tensors, 6 resources, 1513 bindings, 844 uses; nvfp4 = 112, fp8_e4m3fn_row_bf16 = 146
Modes exercised text, arithmetic, vision (image chart), MTP, DFlash2, DFlash2 + vision — all completed, 0 fallback steps
Weights resident 19.0 GiB (text), 19.3 GiB with --vision
Decode speed, no speculation ~70 tok/s (69.5–70.1 across runs)
MTP (--draft-tokens 3) 75.0 % acceptance on a one-sentence answer (148 tok/s overall), 51.9 % on a 256-token haiku prompt (119.5 tok/s)
DFlash2 (--draft-tokens 7) 14.3 % and 13.1 % acceptance on the two short text prompts (78.6 and 90.1 tok/s overall); 85.7 % on the image example (122 tok/s)

Speculative decoding is not guaranteed to reproduce the non-speculative greedy text token for token: in the haiku
smoke the plain, MTP and DFlash2 runs differ in one line.

Provenance

Component Source
Base weights ukisai/Swift-1.5-Qwen3.8-27b @ bc7a1e10b689648585a3ef41494c8d84cf77271a (18 shards, sha256-verified against the Hub LFS metadata)
Abliteration reference Qwen/Qwen3.8-27B @ 1d4bf0f2ff6012fd82039f2fa52739d0dd7c60c0 and huihui-ai/Huihui-Qwen3.8-27B-abliterated @ 739e3c5b89849f6c238ce1e5b70008612ae42cdd (both 18/18 shards sha256-verified)
Quantization recipe unsloth/Qwen3.8-27B-NVFP4 quantization_config, verbatim: recipe/unsloth_qconfig.json and recipe/quantize_nvfp4.py (sha256-pinned copies from Dragoy/Swift-Qwen3.8-27B-abliterated-NVFP4-NInfer @ 4c25aa201b41)
Calibration HuggingFaceH4/ultrachat_200k, 32 samples, seq 2048; dataset revision observed at 8049631c405ae657 before and after quantization (the recipe does not pin it)
Quantizer stack torch 2.11.0, transformers 5.10.1, llmcompressor 0.12.0.1, compressed-tensors 0.17.1
Converter Neroued/ninfer @ bace20dc70249eed6402b66d4852c6c3f9612905, unmodified tools.convert; chat template tools/chat_templates/qwen3_8.jinja (sha256 a497db9e&hellip;) — note this is NInfer's maintained template, not the byte-identical Qwen/Swift chat_template.jinja
DFlash2 drafter z-lab/Qwen3.8-27B-DFlash2 @ 50307d4c4cde6860d4eee73e2547cd786fe8e8a4, model.safetensors sha256 67fc76d68dc5a9415511a4f394ef744d67510cd20e93b37cc2cc7d28e4bab65c
Reports ablation-report.json, source-audit.json, quantization-report.json, conversion-report.json, qwen3_8_27b_swift15_abliterated_nvfp4.ninfer.conversion.json, checksums in SHA256SUMS

How the abliteration was applied and checked

The refusal projection is linear, so its effect on any derivative of the base model is the same constant difference:

W_ablated = W_swift15 + (W_huihui - W_qwen)      # per tensor, fp32, then rounded to bf16
  • The set of changed tensors is measured, not assumed. Comparing the full Qwen and huihui checkpoints (all 1199
    tensors, 18/18 shards each) found exactly 70 differing tensors: self_attn.o_proj, linear_attn.out_proj
    and mlp.down_proj of language layers 17–51. (The huihui model card states layers 18–51; layer 17 is
    also changed, which the build's guard caught.) No vision (model.visual.*), MTP (mtp.*), lm_head or
    embedding tensor differs.
  • Independent audit (source-audit.json): a second, separate program re-read every shard,
    found the 1129 unchanged tensors bit-identical to Swift 1.5, and recomputed the 70 changed tensors from the three
    source checkpoints (torch.equal on the result). The quantizer and converter refuse to run on any checkpoint whose
    shard hashes differ from this audit.
  • The whole chain was built twice from clean inputs: the abliterated shards are byte-identical between runs (18/18 sha256).
  • Vision, MTP and the token embedding come from the BF16 abliterated checkpoint; the 70 abliterated matrices come from
    the quantized checkpoint (import_encoded), verified per matrix in the conversion.

Files

File Purpose
qwen3_8_27b_swift15_abliterated_nvfp4.ninfer the artifact
SHA256SUMS checksums of every file in this repository
recipe/ quantization recipe and the scripts that built and verified this release (Modal)
*-report.json, source-audit.json, evaluation-summary.json, refusal-probe-artifact.json, ninfer-runtime-report.json build, audit and measurement records
NOTICE, LICENSE, LICENSE-APACHE-2.0 licence and modification notices

License

This repository is a derivative of the Swift 1.5 checkpoint, whose license is the
Swift Open License v1.0 — not Apache. The chain:

Component Licence
Qwen/Qwen3.8-27B (base model) Apache-2.0 — Copyright 2026 Alibaba Cloud (LICENSE-APACHE-2.0)
ukisai/Swift-1.5-Qwen3.8-27b (Swift Contribution) Swift Open License v1.0 (LICENSE)
huihui-ai/Huihui-Qwen3.8-27B-abliterated (source of the weight difference) Apache-2.0
This repo (abliteration + quantization + packaging) derivative work — the Swift Contribution contained in it stays under the Swift Open License v1.0

What that means in practice:

  • Free use, including commercial, while your gross revenue (counting all controlled entities) is below the
    $1,000,000 per fiscal year threshold; qualified non-profits have no threshold for non-commercial or research use.
  • Above the threshold: obtain a separate written licence from UkisAI (Swift Enterprise License).
  • Redistribution: ship both licence files, keep the copyright and attribution notices, and mark files you modified
    (Swift licence §4–§5). The weights here are modified (abliteration, quantization); see NOTICE.

This is a description of what the licences say, not legal advice.

Intended use and limitations

This is an uncensored model: the abliteration removes the refusal direction, so it will attempt requests that a
stock instruction-tuned model declines, including harmful ones. It is published for research, evaluation and local
deployment where that behaviour is understood and wanted. It is not safety-aligned, and its answers can be wrong,
dangerous or illegal to act on.

Use at your own responsibility. Anyone deploying it is responsible for their own safeguards, output handling and
compliance with the licences above and applicable law. It is provided as-is, without warranty.

What was not measured: quality benchmarks of any kind (the KL against stock Swift 1.5 above is the only quality
proxy, and it was taken on the BF16 abliterated checkpoint, not on this artifact); refusal behaviour under thinking mode
or other system prompts; long-context behaviour; anything on GPUs other than one RTX PRO 6000 Blackwell. Claims in the
Swift 1.5 card apply to its BF16 weights, not to this
quantized, abliterated build.

Credit for the base model to Qwen (Alibaba Cloud); for the Swift training to
UkisAI; for the abliteration to huihui-ai; for the
engine and artifact contract to Neroued; for the DFlash2 drafter to
z-lab; for the published NVFP4 recipe to unsloth; and
for the refusal scorer to p-e-w/heretic.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.