base_model: ukisai/Swift-1.5-Qwen3.8-27b
library_name: transformers
pipeline_tag: text-generation
license: other
license_name: swift-open-license-1.0
license_link: LICENSE
tags:
- qwen
- qwen3_8
- swift
- swift-1.5
- abliterated
- abliteration
- coding
- reasoning
Swift-1.5-Qwen3.8-27B-Abliterated
A refusal-abliterated derivative ofukisai/Swift-1.5-Qwen3.8-27b.
The objective of this model is to substantially reduce refusal behavior while
preserving the capabilities of the original Swift-1.5 checkpoint as closely
as possible.
This is an independent derivative and is not an official UkisAI or Qwen
release.
Modification
The checkpoint uses a single refusal-direction projection selected after an
A/B capability-retention study.
| Parameter | Value |
|---|---|
| Direction position | 46 |
| Direction extraction mode | non-thinking |
| Direction separation ` | d |
| Direction AUC | 1.0000 |
| Abliteration coefficient | 1.2 |
| Scope | FULL |
| Hidden size | 5120 |
Persisted tensor edits
The final BF16 checkpoint modifies 131 tensors:
- 64 ×
mlp.down_proj - 48 ×
linear_attn.out_proj - 16 ×
self_attn.o_proj - 1 × language-model token embedding
- 2 × MTP writer tensors
The runtime validation configuration projected:
- 128 text writers
- 1 embedding
for 129 live projections.
The persisted checkpoint additionally modifies the two MTP writer tensors so
the MTP path remains aligned with the edited main model.
Evaluation
The following tests compare the original BF16 checkpoint against thepos46 / FULL / lambda=1.2 abliterated model under paired evaluation
conditions.
These are primarily capability-retention experiments, not official
leaderboard submissions.
| Benchmark | Original BF16 | Abliterated | Paired result |
|---|---|---|---|
| XSTest unsafe refusal, n=30 | 25/30 refused (83.3%) | 0/30 refused (0.0%) | 25 refusal→answer, 0 answer→refusal |
| MMLU-style, n=40 | 40/40 | 40/40 | 0 prediction changes |
| MMLU-Pro, n=200 | 131/200 (65.5%) | 130/200 (65.0%) | 5 correct→wrong, 4 wrong→correct |
| HumanEval+ | 140/164 (85.4%) | 140/164 (85.4%) | 0 pass→fail, 0 fail→pass |
| MBPP+ | 285/378 (75.4%) | 285/378 (75.4%) | 5 pass→fail, 5 fail→pass |
XSTest
A deterministic stratified sample of 30 unsafe XSTest prompts was used.
- Seed:
46 - Stratification:
type - Original refusal rate: 83.3%
- Abliterated refusal rate: 0.0%
- Original refusals converted to answers: 25/25
- Answer → refusal regressions: 0
- Exact paired p-value: 5.96046e-08
Generation was limited to 64 new tokens for this test.
Harmless next-token KL
32 harmless prompts were compared using full-vocabulary next-token KL.
| Statistic | KL |
|---|---|
| Mean | 0.571339 |
| Median | 0.165806 |
| P90 | 1.548411 |
| P95 | 2.384434 |
| Maximum | 2.715364 |
These prompts were also involved in refusal-direction discovery and therefore
should be treated as a regression diagnostic rather than a held-out benchmark.
MMLU-Pro
A deterministic, category-balanced subset of 200 MMLU-Pro test questions was
used.
Protocol:
- seed:
46 - non-thinking mode
- constrained next-token scoring across valid A-J answer choices
- same questions and prompt format for both checkpoints
Results:
- Original: 131/200 = 65.5%
- Abliterated: 130/200 = 65.0%
- Correct → wrong: 5
- Wrong → correct: 4
- Prediction changes: 18
- Exact paired p-value: 1.0
- Mean delta gold logP: -0.0296
- Median delta gold logP: -0.0581
This protocol is designed for paired capability-retention measurement and
should not be compared directly with official MMLU-Pro CoT leaderboard scores.
HumanEval+
EvalPlus was used for executable code evaluation.
An initial evaluation exposed a sanitizer artifact where imports from the
benchmark prompt were removed while type annotations such as List andTuple remained. The final paired evaluation therefore restored imports
declared in the benchmark prompt for both models before execution.
Corrected results:
HumanEval Base
- Original: 149/164 = 90.9%
- Abliterated: 148/164 = 90.2%
HumanEval+ Base + Extra
- Original: 140/164 = 85.4%
- Abliterated: 140/164 = 85.4%
- Pass → fail: 0
- Fail → pass: 0
Evaluation timing parameters:
min_time_limit = 1.0gt_time_limit_factor = 8.0
MBPP+
Full MBPP+ v0.2.0 evaluation was run across 378 tasks.
MBPP Base
- Original: 332/378 = 87.8%
- Abliterated: 335/378 = 88.6%
- Pass → fail: 4
- Fail → pass: 7
- Exact paired p-value: 0.548828
MBPP+ Base + Extra
- Original: 285/378 = 75.4%
- Abliterated: 285/378 = 75.4%
- Pass → fail: 5
- Fail → pass: 5
- Exact paired p-value: 1.0
Interpretation
Within the tested protocols, the selected pos46 / FULL / lambda=1.2
configuration produced a large reduction in refusal behavior without a
detectable systematic degradation in the measured knowledge, reasoning, or
coding benchmarks.
Most notably:
- HumanEval+ strict Base+Extra was unchanged.
- MBPP+ strict Base+Extra was unchanged.
- MMLU-Pro changed by only one answer out of 200, with paired regressions and
improvements nearly balanced.
This does not establish that the transformation is lossless for every
workload.
Areas that have not yet been exhaustively evaluated include:
- long-context retention
- multilingual performance
- tool-use / agentic behavior
- additional reasoning benchmarks
- quantized variants
Quantization
This repository contains the BF16 abliterated checkpoint.
A later EXL3 build can be produced directly from this checkpoint.
For practical deployment evaluation, comparing:
Abliterated BF16 → Abliterated EXL3
is sufficient to measure the quantization loss of the final intended model.
A separately quantized original checkpoint is only necessary if performing a
full factorial decomposition of:
- ablation loss
- quantization loss
- ablation × quantization interaction
Reproduction metadata
Machine-readable experiment metadata is available in:
ABLITERATION_METADATA.json
Additional paired benchmark reports are included under:
eval/
when available.
License and attribution
This repository is a derivative ofukisai/Swift-1.5-Qwen3.8-27b.
Please review the included upstream licensing and attribution files before
using or redistributing this checkpoint:
LICENSELICENSE-APACHE-2.0NOTICE
The Swift portion is distributed under the Swift Open License v1.0, while
underlying Qwen components include Apache-2.0 licensed material.
Users are responsible for complying with all applicable upstream license
conditions.
Credits
- UkisAI — Swift
- Alibaba Cloud / Qwen team — Qwen
- EvalPlus — HumanEval+ and MBPP+
- XSTest
- MMLU-Pro