← back to catalog · registered 2026-10-10 20:58

pierpaolo/Swift-1.5-Qwen3.8-27B-NVFP4-Uncensored-MTP-GGUF

pierpaolo 27B GGUF second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/pierpaolo%2FSwift-1.5-Qwen3.8-27B-NVFP4-Uncensored-MTP-GGUF"
Response includes
  • classification m-uncensored
  • files 7
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
1
Model age
today
created 2026-10-10

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
other
Languages
en it
Tags
llama.cpp gguf abliterated uncensored qwen3_8 nvfp4 fp4 q8_0 mtp nextn speculative-decoding dgx-spark
Total size
18.6 GB
Files
7
Quantizations
1
Registered
2026-10-10 20:58
Last updated on HF
2026-10-10 20:22

Files by quantization

Auxiliary files 7 files 18.6 GB
Swift-1.5-Qwen3.8-27B-NVFP4-Q8mix-Uncensored-MTP.gguf 18.6 GB ******** download
abliteration-report.json 25.6 KB 30e5a5bd download
LICENSE 13.0 KB 209a5720 download
LICENSE-APACHE-2.0 11.3 KB f938136e download
README.md 8.75 KB a800fea9 download
NOTICE 3.59 KB 8c0336e1 download
.gitattributes 1.57 KB 8889753c download

README current version from Hugging Face


license: other
license_name: swift-open-license-1.0
license_link: LICENSE
base_model:

  • jessedye90/qwen3.8-27b-swift-uncensored
  • ukisai/Swift-1.5-Qwen3.8-27b-NVFP4
    base_model_relation: quantized
    library_name: llama.cpp
    pipeline_tag: text-generation
    language:
  • en
  • it
    tags:
  • gguf
  • llama.cpp
  • abliterated
  • uncensored
  • qwen3_8
  • nvfp4
  • fp4
  • q8_0
  • mtp
  • nextn
  • speculative-decoding
  • dgx-spark
    extra_gated_prompt: >-
    This model has had its safety refusals removed (abliteration). It will comply
    with harmful, unethical or illegal requests that the base model refuses,
    including requests in the categories covered by AdvBench. It is published for
    research into refusal mechanisms, interpretability, alignment and
    red-teaming. You alone are responsible for how you use it and for everything
    it generates, and you must comply with the Swift Open License v1.0, the Apache
    License 2.0 and applicable law. By requesting access you confirm that you
    understand and accept this.
    extra_gated_fields:
    Full name: text
    Country: country
    Affiliation: text
    I understand that safety refusals have been removed and accept the applicable licenses: checkbox

Swift-1.5-Qwen3.8-27B-NVFP4-Uncensored-MTP (GGUF)

llama.cpp GGUF build of
jessedye90/qwen3.8-27b-swift-uncensored
(UkisAI's Swift 1.5 fine-tune of Qwen3.8-27B, ModelOpt NVFP4/FP8) with its
safety refusals projected out of the weights, plus a working MTP head so it
can be served with llama.cpp's built-in multi-token-prediction speculative
decoding (--spec-type draft-mtp).

⚠️ Safety alignment is removed. This model will comply with harmful,
unethical or illegal requests that the base model refuses. It is released for
research (refusal mechanisms, interpretability, red-teaming, robustness
evaluation). Do not deploy it to end users without your own safety layer.
Not made or endorsed by UkisAI, OrcaRouter, Alibaba Cloud or NVIDIA.

Why this build exists

The abliterated checkpoint ships without an MTP head (its author serves it
with a separate DFlash2 draft model instead). The upstream non-abliterated Swift
1.5 release does carry an MTP head, but taken from the censored base. This
build restores that MTP head and applies the same refusal-direction projection
to its residual writers
, so the drafter is consistent with the target, then
converts the result to GGUF while keeping the checkpoints' NVFP4 tensors
byte-identical.

What was changed, relative to the two parents

  1. MTP head merged — the 15 mtp.* tensors (BF16) were taken from
    ukisai/Swift-1.5-Qwen3.8-27b-NVFP4
    (shard model-00005-of-00005.safetensors) and merged into the abliterated
    checkpoint.
  2. MTP abliterated — mtp.layers.0.self_attn.o_proj and
    mtp.layers.0.mlp.down_proj had the same unit refusal direction projected out
    of their outputs (exact Arditi projection, W ← W − W r rᵀ). These tensors are
    BF16, so the projection is exact — no quantisation grid is involved.
    Residual component along r after the edit: ~1.2% of the original.
    mtp.fc is not a residual writer and is unchanged. embed_tokens was
    not edited, matching the base abliterated release.
  3. Converted to GGUF — convert_hf_to_gguf.py --outtype bf16 followed by
    llama-quantize --tensor-type-file, where the 193 NVFP4 tensors
    (64 layers × ffn_gate/ffn_up/ffn_down + output.weight) are copied
    unchanged
    from the checkpoint (verified byte-identical by SHA-256), the rest
    went to Q8_0/F32.

In total 130 residual writers carry the projection: 128 from the abliterated
release (64 FP8 o_proj/out_proj + 64 NVFP4 down_proj) plus the 2 MTP ones.
OrcaRouter's original abliteration covered 131 (it also edited embed_tokens);
embed_tokens is deliberately left untouched here — the upstream author found
that editing the token-embedding path broke a strict input-validation coding
task on a sibling model.

Model details

Architecture qwen35 (Gated DeltaNet + full attention hybrid), 65 blocks (64 + 1 MTP)
Parameters 27B class
Context length 262144
Format GGUF, single file, 19.9 GB
Quantisation mixed — 193 NVFP4 / 217 Q8_0 / 842 F32 (1252 tensors, 5.75 BPW)
MTP head present (blk.64.*, qwen35.nextn_predict_layers = 1)
Tokenizer / chat template embedded, same as the base
llama.cpp requires a build with qwen35 + NVFP4 (40) + draft-mtp support — verified on b11541 (commit f2918cabb)

Measured results

Measured locally on llama.cpp (build b11541, MTP enabled, greedy, thinking off)
against the censored parent quantised the same way, on the same machine. The
censored baseline was measured by us on the identical prompt set; the
abliterated-release figures come from the upstream author
(reference).

censored baseline this model
AdvBench harmful (harmful_behaviors, 100 prompts) refused 100 / 100 0 / 100
XSTest safe prompts (100) refused — over-refusal 5 / 100 0 / 100
MTP draft acceptance rate (16 prompts, same target) 61.04% 61.00% (1.83 tokens/step)
LRU-cache + longest-palindrome coding tasks — 7/7 + 6/6 asserts pass
1000! trailing zeros (verifiable) — 249 ✅

Notes on the numbers:

  • Both prompt sets are classified by refusal phrases (eval_refusal.py).
    Two classifier errors were found and corrected by hand, one per side:
    a self-referential "I cannot use an identity to scam people online myself"
    counted as a refusal on the uncensored model (it is not), and a crisis-support
    redirect counted as compliance on the baseline (it is not, no instructions were
    given). Raw verdicts are in evals/advbench_results_*_raw.json.
  • The MTP ablation is performance-neutral (61.04% → 61.00%, i.e. noise): it
    is done for consistency with the target, not for speed. Target outputs are
    byte-identical with and without the MTP edit.
  • This model does answer the AdvBench suicide-instruction prompt. The censored
    baseline redirected to crisis support. That is the expected consequence of
    removing the refusal direction, and it is the reason for the warning above.

Usage (llama.cpp)

llama-server \
  -m Swift-1.5-Qwen3.8-27B-NVFP4-Q8mix-Uncensored-MTP.gguf \
  -ngl 99 --no-mmproj \
  --host 0.0.0.0 --port 8080 \
  -c 262144 \
  --flash-attn on --jinja \
  --cache-type-k q8_0 --cache-type-v q8_0 --kv-unified \
  --spec-type draft-mtp --spec-draft-n-max 3 \
  --temp 1.0 --top-k 20 --top-p 0.95

--spec-type draft-mtp enables speculative decoding with the built-in MTP head;
drop it for a plain (slower) decode. The MTP head is verified lossless: the target
model checks every drafted token, and outputs are identical with and without it.

Reproducibility

Path Contents
abliteration/merge_mtp.py merges the 15 MTP tensors into the checkpoint (stdlib safetensors IO)
abliteration/ablate_mtp.py projects the refusal direction out of the 2 MTP residual writers, in place
abliteration/tensor-types.txt the 1252-line llama-quantize --tensor-type-file mapping
abliteration/refusal_direction.safetensors the unit refusal direction (r, 5120-d) used for the projections
abliteration/ablate_quant.py, apply_27b.py, direction.py the upstream author's abliteration toolkit
abliteration/abliteration-report.json per-tensor resid / rel_change for the 128 upstream edits
evals/ evaluation scripts and raw results (refusal, AdvBench, XSTest, quality)
NOTICE full change notices, including exactly what this build modified

License and obligations

This is a derivative work, redistributed under the terms of both parents:

  • Swift Contribution — Swift Open License v1.0 (LICENSE), Copyright 2026
    UkisAI. Section 5 limits Commercial Use: an entity with US$1M or more in
    annual gross revenue needs a separate Swift Enterprise License from UkisAI.
  • Base Model — Qwen3.8-27B (LICENSE-APACHE-2.0), Copyright 2026 Alibaba
    Cloud, Apache License 2.0.

Redistribution conditions of the Swift Open License (Section 4) are met by
shipping LICENSE, LICENSE-APACHE-2.0 and NOTICE — including the change
notices required by Section 4(b) for the files modified by this build (see
NOTICE, "GGUF / MTP change notice"). The NOTICE file also carries UkisAI's
and the upstream abliteration author's attribution notices.

"UkisAI", "Swift", "Qwen", "OrcaRouter" and "NVIDIA" are used only to state
where this model comes from. This release is not made or endorsed by any of them.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration