← back to catalog · registered 2026-10-09 01:58

jakeatx/ATX-Swift-1.5-Qwen3.8-27B-Uncensored-2.3bpw-MTP-GGUF

jakeatx 27B GGUF second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/jakeatx%2FATX-Swift-1.5-Qwen3.8-27B-Uncensored-2.3bpw-MTP-GGUF"
Response includes
  • classification m8
  • files 6
  • author_summary 5 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-10-09

Training datasets

Corpora the author lists in the model card. Datasets tracked in our /datasets catalog carry a category badge linking to the workflow stage. Others open on Hugging Face.

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
other
Tags
llama.cpp gguf quantized exl3 qwen3.8 swift-1.5 uncensored abliterated mtp 12gb text-generation dataset:jakeatx/ATX-Swift-1.5-Qwen3.8-27B-Uncensored-2.3bpw-LiveCodeBench-Pagoda

Related

Total size
8.38 GB
Files
6
Quantizations
1
Registered
2026-10-09 01:58
Last updated on HF
2026-10-09 02:07

Files by quantization

Auxiliary files 6 files 8.38 GB
ATX-Swift-1.5-Qwen3.8-27B-Uncensored-2.3bpw-MTP.gguf 8.38 GB 3c7f5333 download
README.md 15.2 KB 70a4b308 download
LICENSE 13.0 KB 209a5720 download
LICENSE-APACHE-2.0 11.3 KB f938136e download
.gitattributes 1.57 KB bbcfdf84 download
NOTICE 1.11 KB c4ad1a71 download

README current version from Hugging Face


license: other
license_name: swift-open-license-1.0
base_model: d0xin/Swift-1.5-Qwen3.8-27B-Uncensored-BF16
base_model_relation: quantized
library_name: llama.cpp
pipeline_tag: text-generation
datasets:

  • jakeatx/ATX-Swift-1.5-Qwen3.8-27B-Uncensored-2.3bpw-LiveCodeBench-Pagoda
    tags:
  • gguf
  • quantized
  • exl3
  • qwen3.8
  • swift-1.5
  • uncensored
  • abliterated
  • mtp
  • 12gb

Swift 1.5 Qwen3.8-27B Uncensored — ATX 2.3 bpw + MTP (12 GB cards)

Uses fewer tokens than other models its size

  • 0.73x the output tokens of stock Qwen3.8-27B at 4 bits. On 24 fixed questions (3 gate prompts, 11 GPQA, 10 LiveCodeBench), each answered greedy to end of sequence, this build's total output (reasoning + answer) was 0.73x that of stock Qwen3.8-27B ATX-4-XS (geometric mean of per-question ratios, n = 24, 90% CI 0.54–0.98).
  • Pagoda agent task (reasoning effort xhigh, 204,800 context, 12 GB budget): this build used 30K–42K reasoning tokens per run and wrote a working page in 3 of 3 runs. Ternary Bonsai 2 27B used 69K–92K and the Mirai fork 1.5 bpw used 28K–87K, with 0 of 3 working pages each.

Keeps some, not all, of Swift 1.5's gains

Swift 1.5 IQ4_XS-M is 0.81x stock on the same 24 questions, so this build keeps part of Swift's shorter reasoning but not all of it: at temperature 1.0 it writes about 1.22x the reasoning of Swift 1.5 IQ4_XS-M (90% CI 1.11–1.33, 348 responses). On LiveCodeBench v6 it scores 77.7%, against 89.25% for Swift 1.5 IQ4_XS-M and 90.3% published for Qwen BF16.

Honest assessment. We view this quant as retaining roughly 85% of BF16 benchmark score. On LiveCodeBench v6 it scores 77.71% pooled over 7 runs on 12 GB cards, against 90.3% published for Qwen BF16 (86% of it) and 89.25% for our 4.56 bpw Swift IQ4_XS-M. It is a 12 GB option, not a replacement for the 4-bit build on a 24 GB card. Details and failure signatures are below.

Test results, videos and harness: jakeatx/ATX-Swift-1.5-Qwen3.8-27B-Uncensored-2.3bpw-LiveCodeBench-Pagoda

This is the Swift 1.5 uncensored BF16 checkpoint natively re-encoded at about 2.3 bits per weight, with its MTP head, sized for 12 GB cards at about 200K context with MTP drafting. The source derives from UkisAI's Swift 1.5, which is built on Qwen3.8-27B. Source revision: 15165fce17cb716934a2b15e746d9c7b061d4a0f.

The format and the per-tensor bit allocation are inspired by the Mirai models: Mirai Labs' Mirai S (a 2.4-bit Qwen3.8-27B) and its community llama.cpp port by alesha-pro. No Mirai weights are included. Every weight in this file was encoded from Swift 1.5 BF16.

The GGUF contains the text model and the MTP draft layer. It does not contain the vision tower. This is an uncensored, refusal-reduced derivative.

Artifact Details
ATX-Swift-1.5-Qwen3.8-27B-Uncensored-2.3bpw-MTP.gguf 9.00 GB (9,003,017,888 bytes), SHA-256 3c7f5333fea80bea42660934cdea13bbee3104a2c9b06d50b9f496babb95af90

Runtime: llamAmpere v0.5 required

This GGUF uses EXL3 trellis tensor types (EXL3_2, EXL3_3, EXL3_4) and is meant to run with SJ-KVaRN KV cache compression. It needs llamAmpere v0.5 (release pending). Stock llama.cpp will not load it, and the 12 GB configuration uses SJ-KVaRN and drafter options that first ship in v0.5. The commands below are written for v0.5 and will run as-is once it is published. llamAmpere targets NVIDIA Ampere (RTX 30-series) cards.

12 GB cards (RTX 3060 12 GB, 3080 12 GB, 3080 Ti): 204,800 context, MTP

LLAMA_MTP_DRAFT_COMPUTE_LEAN=1 llama-server \
  -hf jakeatx/ATX-Swift-1.5-Qwen3.8-27B-Uncensored-2.3bpw-MTP-GGUF \
  -hff ATX-Swift-1.5-Qwen3.8-27B-Uncensored-2.3bpw-MTP.gguf \
  -c 204800 --parallel 1 -ngl 99 -fa on -fit off -b 4096 -ub 512 \
  --no-context-shift --cache-ram 0 --jinja \
  -ctk kvarn3 -ctv kvarn2 --kvarn-body-type kvarn4t \
  --kvarn-sink 128 --kvarn-sink-type f16 --kvarn-staging-type tq6_0 \
  --kvarn-tail 4096 --kvarn-tail-max 8192 \
  --cache-type-s q8_0 \
  --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0 \
  --spec-draft-vocab-map auto:65536 \
  --spec-draft-type-k q8_0 --spec-draft-type-v q8_0 --spec-draft-window 8192 \
  --temp 1.0 --top-k 20 --top-p 0.95 --min-p 0

What the settings do:

  • -ctk kvarn3 -ctv kvarn2 --kvarn-body-type kvarn4t ...: SJ-KVaRN KV cache, 3-bit K and 2-bit V with a trellis body, a 128-token f16 sink, tq6_0 staging and an adaptive 4,096–8,192-token tail.
  • --cache-type-s q8_0: the Gated DeltaNet recurrent state is stored at 8 bits.
  • --spec-type draft-mtp --spec-draft-n-max 4: the built-in MTP head drafts up to 4 tokens per step. --spec-draft-vocab-map auto:65536 uses the 65,536-token draft vocabulary built into llamAmpere for Qwen3.8-27B (the same list as atx_65536). --spec-draft-window 8192 and the q8_0 draft KV keep the drafter's cache small.
  • LLAMA_MTP_DRAFT_COMPUTE_LEAN=1: caps the MTP draft context's micro-batch at 64 tokens, which shrinks the drafter's compute buffer. It is part of the measured 12 GB fit, so keep it.
  • -ub 512: smaller prefill micro-batches, which keep the compute buffers inside the budget.

Measured fit: at ctx 204,800, a prompt of 203,568 tokens plus 256 generated tokens peaked at 10,690 MiB whole-card memory, under the 11,000 MiB budget we use for 12 GB cards (it leaves room for the desktop and the driver). Measured on an RTX 3090 with nvidia-smi sampling every 0.5 s. On the LiveCodeBench rentals the same configuration was calibrated on each 12 GB card itself and ran at 229,376 context on the RTX 3060 and 221,184 on the RTX 3080 / 3080 Ti.

The measured runs pointed --spec-draft-vocab-map at the file atx_65536.txt and also set GGML_CUDA_PREFILL_KV_MIB=256 and GGML_KVARN_PREFILL_MIB=256. Those two variables only restate the built-in default of 256 MiB for the bounded-prefill workspaces, so they are left out above.

24 GB cards (RTX 3090 / 3090 Ti): 131,072 context, MTP

The configuration used for the Hermes agent test below:

llama-server \
  -hf jakeatx/ATX-Swift-1.5-Qwen3.8-27B-Uncensored-2.3bpw-MTP-GGUF \
  -hff ATX-Swift-1.5-Qwen3.8-27B-Uncensored-2.3bpw-MTP.gguf \
  -c 131072 --parallel 1 -ngl 99 -fa on -b 4096 -ub 1024 \
  -ctk tq5_0 -ctv turbo4 \
  --spec-type draft-mtp --spec-draft-vocab-map auto:65536 \
  --spec-draft-type-k tq5_0 --spec-draft-type-v turbo4 \
  --jinja --temp 1.0 --top-k 20 --top-p 0.95

On a 24 GB card, our Swift 1.5 IQ4_XS-M build is the stronger choice. This 2.3 bpw build is for cards where that one does not fit.

Sampling and reasoning

All tests used temperature 1.0, top-k 20, top-p 0.95, min-p 0. Reasoning effort is set per request with "chat_template_kwargs": {"reasoning_effort": "xhigh"} (or medium). In the agent test at xhigh, the counted runs also set a thinking budget (--reasoning-budget 180000 on the server, lowered per request so the answer always fits the context). Two earlier xhigh runs at 131,072 context with no budget spent the whole context planning and wrote no file (see the dataset).

How it was made

Tensor class Format
Attention and Gated DeltaNet inputs (q/k/v, qkv, gate) EXL3_3 (EXL3_2 in layers 1, 2, 4 and 62)
Writers (ffn_down, ssm_out, attn_output) EXL3_2
ffn_gate / ffn_up EXL3_2, EXL3_3 in layers 53–61
Output head EXL3_3
MTP block (Swift 1.5's own MTP head) EXL3_4
Token embedding Q8_0 (kept in host RAM, no VRAM)
96 small attention/GDN tensors, norms Q8_0 / F32

The allocation follows the Mirai S pattern (3 bits on every attention input, 2 bits on the writers, more bits in the late FFN layers and the head) mapped onto EXL3 trellis types, at about the same GPU byte budget as the Mirai fork we tested (7,287 MiB GPU-resident including MTP, by tensor-table sum). The EXL3 tensors were encoded with the exllamav3 quantizer, capturing Hessians layer by layer from Swift 1.5's own activations through the already-quantized prefix, then converted to GGUF and checked bit-exact against the exllamav3 encoding. The file is 9.0 GB on disk because the token embedding is Q8_0 (1,288 MiB). It stays in host RAM and costs no VRAM.

KL vs Swift 1.5 BF16: 0.157 nats mean (median 0.060, p90 0.321) over 8 windows of 2,048 tokens (16,376 scored tokens). On the same windows and teacher, the 1.5 bpw Mirai fork, a Qwen3.8-27B model, measured 0.218 (Qwen BF16 itself measures 0.008 against Swift BF16).

LiveCodeBench v6 on 12 GB cards

Pinned 100-task LiveCodeBench v6 subset (the same tasks and protocol as our IQ4_XS-M 89.25% run), reasoning effort xhigh, temperature 1.0 / top-p 0.95 / top-k 20 / min-p 0, no max_tokens. Each run used one rented 12 GB card with the 12 GB configuration above at the context calibrated on that card. Complete runs only (100/100 tasks).

Card Runs (seeds) pass@1 per run Mean pass@1 Truncated MTP acceptance Gen tok/s
RTX 3060 12 GB 3 (0, 1, 2) 79, 79, 79 79.00% 0 / 300 0.517 28.9
RTX 3080 12 GB 1 (0) 77 77.00% 1 / 100 0.503 61.1
RTX 3080 Ti 3 (0, 1, 2) 74, 80, 76 76.67% 1 / 300 0.513 37.4
Pooled 7 77.71% (544/700) 2 / 700 0.513 34.7
  • Pooled 95% CI: 75.74–79.69 (t over runs), 71.14–83.86 (task bootstrap).
  • vs Qwen BF16 published 90.3%: −12.59 pt; 77.71 / 90.3 = 86%.
  • vs Swift IQ4_XS-M 89.25% (4 seeds, same tasks): −11.54 pt (paired task bootstrap 95% CI −15.86 to −7.61).
  • Gen tok/s is token-weighted over each card type and includes slow rental hosts (one 3080 Ti host ran at 19.4 tok/s; the other two at about 69). The power limits were the rental hosts' own.

Grader note (secondary). The pinned verifier grades five multi-answer or float-tolerance tasks by exact match, which also affects the 89.25% reference. An audit re-graded the first 430 completed rows with special judges for those five tasks: 354/430 (82.3%) became 366/430 (85.1%). The headline above uses the original grader so it compares directly with 89.25%. A docker re-grade of the same 430 rows with the reference harness matched 430/430.

Known failure signatures (from the audit of 76 failed rows):

  • Token-slip syntax errors in about 2% of answers (9 of 428 extracted programs), such as max(max(a), max(b) or seen.add(nums[j). The reasoning often contains the correct line and only the final code is corrupted. None of 861 IQ4_XS-M answers showed this pattern.
  • Longer reasoning: median completion length about 1.3x the IQ4_XS-M reference per task (geomean 1.45x).
  • Occasional runaway generation: on one hard task, two runs used the full 221K context without emitting code.

Agent reliability: Hermes

We used this model as the backend of the Hermes agent (CLI) for a real research-and-build task on one RTX 3090 Ti (350 W), with the 24 GB command above (131,072 context, tq5_0/turbo4 KV, MTP with the 65,536-token draft vocabulary). It completed 62 API calls with 39,604 output tokens in 598.0 s of summed request latency, 66.2 tok/s end to end (prefill of the uncached suffix and time to first token included), at contexts from 16.7K to 67.3K tokens. The longest generations ran at 73.6–83.8 tok/s end to end. The agent produced a self-contained oil-market report page with charts. The prices and events in it came from the model's web searches and have not been checked.

Recording of the page · the page itself · per-call log

Pagoda test: agentic coding on a 12 GB budget

The voxel-pagoda prompt (from Google AI's post), run through the pi coding agent with read/write/edit/bash tools, xhigh reasoning, ctx 204,800, the 12 GB budget, 3 seeds per model, a 45-minute wall clock and a 30-turn cap. Every comparator used the same 3-bit K / 2-bit V SJ-KVaRN cache.

Model Working pages (renders, no errors)
This model (2.3 bpw + MTP) 3 / 3
Qwen3.8-27B UD-IQ2_XXS (unsloth, 2.16 bpw) 0 / 3 (no file written in any run)
Mirai fork 1.5 bpw + MTP 0 / 3 (one page with a syntax error, one truncated script that never runs, one degenerate render)
Ternary Bonsai 2 27B PQ2_0 0 / 3 (two pages with syntax errors, one run wrote no file)

All three of our runs wrote a working three.js scene and spent the remaining turns checking and polishing it until the turn cap.

Seed 0 Seed 1 Seed 2
seed 0 seed 1 seed 2
video video video

Every page, video, comparator run and the full harness notes are in the dataset.

License and credits

Distributed under the Swift Open License v1.0, inherited from Swift 1.5. See LICENSE, LICENSE-APACHE-2.0, and NOTICE for the applicable terms and attribution. Credit goes to UkisAI for Swift 1.5, d0xin for the uncensored derivative, and the Qwen team for Qwen3.8-27B. The bit allocation is inspired by Mirai Labs' Mirai S and alesha-pro's llama.cpp port; no Mirai weights are included. The EXL3 format and quantizer are from exllamav3 by turboderp.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration