license: other
license_name: swift-open-license-1.0
base_model: d0xin/Swift-1.5-Qwen3.8-27B-Uncensored-BF16
base_model_relation: quantized
library_name: llama.cpp
pipeline_tag: text-generation
datasets:
- jakeatx/ATX-Swift-1.5-Qwen3.8-27B-Uncensored-2.3bpw-LiveCodeBench-Pagoda
tags: - gguf
- quantized
- exl3
- qwen3.8
- swift-1.5
- uncensored
- abliterated
- mtp
- 12gb
Swift 1.5 Qwen3.8-27B Uncensored — ATX 2.3 bpw + MTP (12 GB cards)
Uses fewer tokens than other models its size
- 0.73x the output tokens of stock Qwen3.8-27B at 4 bits. On 24 fixed questions (3 gate prompts, 11 GPQA, 10 LiveCodeBench), each answered greedy to end of sequence, this build's total output (reasoning + answer) was 0.73x that of stock Qwen3.8-27B ATX-4-XS (geometric mean of per-question ratios, n = 24, 90% CI 0.54–0.98).
- Pagoda agent task (reasoning effort xhigh, 204,800 context, 12 GB budget): this build used 30K–42K reasoning tokens per run and wrote a working page in 3 of 3 runs. Ternary Bonsai 2 27B used 69K–92K and the Mirai fork 1.5 bpw used 28K–87K, with 0 of 3 working pages each.
Keeps some, not all, of Swift 1.5's gains
Swift 1.5 IQ4_XS-M is 0.81x stock on the same 24 questions, so this build keeps part of Swift's shorter reasoning but not all of it: at temperature 1.0 it writes about 1.22x the reasoning of Swift 1.5 IQ4_XS-M (90% CI 1.11–1.33, 348 responses). On LiveCodeBench v6 it scores 77.7%, against 89.25% for Swift 1.5 IQ4_XS-M and 90.3% published for Qwen BF16.
Honest assessment. We view this quant as retaining roughly 85% of BF16 benchmark score. On LiveCodeBench v6 it scores 77.71% pooled over 7 runs on 12 GB cards, against 90.3% published for Qwen BF16 (86% of it) and 89.25% for our 4.56 bpw Swift IQ4_XS-M. It is a 12 GB option, not a replacement for the 4-bit build on a 24 GB card. Details and failure signatures are below.
Test results, videos and harness: jakeatx/ATX-Swift-1.5-Qwen3.8-27B-Uncensored-2.3bpw-LiveCodeBench-Pagoda
This is the Swift 1.5 uncensored BF16 checkpoint natively re-encoded at about 2.3 bits per weight, with its MTP head, sized for 12 GB cards at about 200K context with MTP drafting. The source derives from UkisAI's Swift 1.5, which is built on Qwen3.8-27B. Source revision: 15165fce17cb716934a2b15e746d9c7b061d4a0f.
The format and the per-tensor bit allocation are inspired by the Mirai models: Mirai Labs' Mirai S (a 2.4-bit Qwen3.8-27B) and its community llama.cpp port by alesha-pro. No Mirai weights are included. Every weight in this file was encoded from Swift 1.5 BF16.
The GGUF contains the text model and the MTP draft layer. It does not contain the vision tower. This is an uncensored, refusal-reduced derivative.
| Artifact | Details |
|---|---|
ATX-Swift-1.5-Qwen3.8-27B-Uncensored-2.3bpw-MTP.gguf |
9.00 GB (9,003,017,888 bytes), SHA-256 3c7f5333fea80bea42660934cdea13bbee3104a2c9b06d50b9f496babb95af90 |
Runtime: llamAmpere v0.5 required
This GGUF uses EXL3 trellis tensor types (EXL3_2, EXL3_3, EXL3_4) and is meant to run with SJ-KVaRN KV cache compression. It needs llamAmpere v0.5 (release pending). Stock llama.cpp will not load it, and the 12 GB configuration uses SJ-KVaRN and drafter options that first ship in v0.5. The commands below are written for v0.5 and will run as-is once it is published. llamAmpere targets NVIDIA Ampere (RTX 30-series) cards.
12 GB cards (RTX 3060 12 GB, 3080 12 GB, 3080 Ti): 204,800 context, MTP
LLAMA_MTP_DRAFT_COMPUTE_LEAN=1 llama-server \
-hf jakeatx/ATX-Swift-1.5-Qwen3.8-27B-Uncensored-2.3bpw-MTP-GGUF \
-hff ATX-Swift-1.5-Qwen3.8-27B-Uncensored-2.3bpw-MTP.gguf \
-c 204800 --parallel 1 -ngl 99 -fa on -fit off -b 4096 -ub 512 \
--no-context-shift --cache-ram 0 --jinja \
-ctk kvarn3 -ctv kvarn2 --kvarn-body-type kvarn4t \
--kvarn-sink 128 --kvarn-sink-type f16 --kvarn-staging-type tq6_0 \
--kvarn-tail 4096 --kvarn-tail-max 8192 \
--cache-type-s q8_0 \
--spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0 \
--spec-draft-vocab-map auto:65536 \
--spec-draft-type-k q8_0 --spec-draft-type-v q8_0 --spec-draft-window 8192 \
--temp 1.0 --top-k 20 --top-p 0.95 --min-p 0
What the settings do:
-ctk kvarn3 -ctv kvarn2 --kvarn-body-type kvarn4t ...: SJ-KVaRN KV cache, 3-bit K and 2-bit V with a trellis body, a 128-token f16 sink, tq6_0 staging and an adaptive 4,096–8,192-token tail.--cache-type-s q8_0: the Gated DeltaNet recurrent state is stored at 8 bits.--spec-type draft-mtp --spec-draft-n-max 4: the built-in MTP head drafts up to 4 tokens per step.--spec-draft-vocab-map auto:65536uses the 65,536-token draft vocabulary built into llamAmpere for Qwen3.8-27B (the same list asatx_65536).--spec-draft-window 8192and the q8_0 draft KV keep the drafter's cache small.LLAMA_MTP_DRAFT_COMPUTE_LEAN=1: caps the MTP draft context's micro-batch at 64 tokens, which shrinks the drafter's compute buffer. It is part of the measured 12 GB fit, so keep it.-ub 512: smaller prefill micro-batches, which keep the compute buffers inside the budget.
Measured fit: at ctx 204,800, a prompt of 203,568 tokens plus 256 generated tokens peaked at 10,690 MiB whole-card memory, under the 11,000 MiB budget we use for 12 GB cards (it leaves room for the desktop and the driver). Measured on an RTX 3090 with nvidia-smi sampling every 0.5 s. On the LiveCodeBench rentals the same configuration was calibrated on each 12 GB card itself and ran at 229,376 context on the RTX 3060 and 221,184 on the RTX 3080 / 3080 Ti.
The measured runs pointed --spec-draft-vocab-map at the file atx_65536.txt and also set GGML_CUDA_PREFILL_KV_MIB=256 and GGML_KVARN_PREFILL_MIB=256. Those two variables only restate the built-in default of 256 MiB for the bounded-prefill workspaces, so they are left out above.
24 GB cards (RTX 3090 / 3090 Ti): 131,072 context, MTP
The configuration used for the Hermes agent test below:
llama-server \
-hf jakeatx/ATX-Swift-1.5-Qwen3.8-27B-Uncensored-2.3bpw-MTP-GGUF \
-hff ATX-Swift-1.5-Qwen3.8-27B-Uncensored-2.3bpw-MTP.gguf \
-c 131072 --parallel 1 -ngl 99 -fa on -b 4096 -ub 1024 \
-ctk tq5_0 -ctv turbo4 \
--spec-type draft-mtp --spec-draft-vocab-map auto:65536 \
--spec-draft-type-k tq5_0 --spec-draft-type-v turbo4 \
--jinja --temp 1.0 --top-k 20 --top-p 0.95
On a 24 GB card, our Swift 1.5 IQ4_XS-M build is the stronger choice. This 2.3 bpw build is for cards where that one does not fit.
Sampling and reasoning
All tests used temperature 1.0, top-k 20, top-p 0.95, min-p 0. Reasoning effort is set per request with "chat_template_kwargs": {"reasoning_effort": "xhigh"} (or medium). In the agent test at xhigh, the counted runs also set a thinking budget (--reasoning-budget 180000 on the server, lowered per request so the answer always fits the context). Two earlier xhigh runs at 131,072 context with no budget spent the whole context planning and wrote no file (see the dataset).
How it was made
| Tensor class | Format |
|---|---|
| Attention and Gated DeltaNet inputs (q/k/v, qkv, gate) | EXL3_3 (EXL3_2 in layers 1, 2, 4 and 62) |
Writers (ffn_down, ssm_out, attn_output) |
EXL3_2 |
ffn_gate / ffn_up |
EXL3_2, EXL3_3 in layers 53–61 |
| Output head | EXL3_3 |
| MTP block (Swift 1.5's own MTP head) | EXL3_4 |
| Token embedding | Q8_0 (kept in host RAM, no VRAM) |
| 96 small attention/GDN tensors, norms | Q8_0 / F32 |
The allocation follows the Mirai S pattern (3 bits on every attention input, 2 bits on the writers, more bits in the late FFN layers and the head) mapped onto EXL3 trellis types, at about the same GPU byte budget as the Mirai fork we tested (7,287 MiB GPU-resident including MTP, by tensor-table sum). The EXL3 tensors were encoded with the exllamav3 quantizer, capturing Hessians layer by layer from Swift 1.5's own activations through the already-quantized prefix, then converted to GGUF and checked bit-exact against the exllamav3 encoding. The file is 9.0 GB on disk because the token embedding is Q8_0 (1,288 MiB). It stays in host RAM and costs no VRAM.
KL vs Swift 1.5 BF16: 0.157 nats mean (median 0.060, p90 0.321) over 8 windows of 2,048 tokens (16,376 scored tokens). On the same windows and teacher, the 1.5 bpw Mirai fork, a Qwen3.8-27B model, measured 0.218 (Qwen BF16 itself measures 0.008 against Swift BF16).
LiveCodeBench v6 on 12 GB cards
Pinned 100-task LiveCodeBench v6 subset (the same tasks and protocol as our IQ4_XS-M 89.25% run), reasoning effort xhigh, temperature 1.0 / top-p 0.95 / top-k 20 / min-p 0, no max_tokens. Each run used one rented 12 GB card with the 12 GB configuration above at the context calibrated on that card. Complete runs only (100/100 tasks).
| Card | Runs (seeds) | pass@1 per run | Mean pass@1 | Truncated | MTP acceptance | Gen tok/s |
|---|---|---|---|---|---|---|
| RTX 3060 12 GB | 3 (0, 1, 2) | 79, 79, 79 | 79.00% | 0 / 300 | 0.517 | 28.9 |
| RTX 3080 12 GB | 1 (0) | 77 | 77.00% | 1 / 100 | 0.503 | 61.1 |
| RTX 3080 Ti | 3 (0, 1, 2) | 74, 80, 76 | 76.67% | 1 / 300 | 0.513 | 37.4 |
| Pooled | 7 | 77.71% (544/700) | 2 / 700 | 0.513 | 34.7 |
- Pooled 95% CI: 75.74–79.69 (t over runs), 71.14–83.86 (task bootstrap).
- vs Qwen BF16 published 90.3%: −12.59 pt; 77.71 / 90.3 = 86%.
- vs Swift IQ4_XS-M 89.25% (4 seeds, same tasks): −11.54 pt (paired task bootstrap 95% CI −15.86 to −7.61).
- Gen tok/s is token-weighted over each card type and includes slow rental hosts (one 3080 Ti host ran at 19.4 tok/s; the other two at about 69). The power limits were the rental hosts' own.
Grader note (secondary). The pinned verifier grades five multi-answer or float-tolerance tasks by exact match, which also affects the 89.25% reference. An audit re-graded the first 430 completed rows with special judges for those five tasks: 354/430 (82.3%) became 366/430 (85.1%). The headline above uses the original grader so it compares directly with 89.25%. A docker re-grade of the same 430 rows with the reference harness matched 430/430.
Known failure signatures (from the audit of 76 failed rows):
- Token-slip syntax errors in about 2% of answers (9 of 428 extracted programs), such as
max(max(a), max(b)orseen.add(nums[j). The reasoning often contains the correct line and only the final code is corrupted. None of 861 IQ4_XS-M answers showed this pattern. - Longer reasoning: median completion length about 1.3x the IQ4_XS-M reference per task (geomean 1.45x).
- Occasional runaway generation: on one hard task, two runs used the full 221K context without emitting code.
Agent reliability: Hermes
We used this model as the backend of the Hermes agent (CLI) for a real research-and-build task on one RTX 3090 Ti (350 W), with the 24 GB command above (131,072 context, tq5_0/turbo4 KV, MTP with the 65,536-token draft vocabulary). It completed 62 API calls with 39,604 output tokens in 598.0 s of summed request latency, 66.2 tok/s end to end (prefill of the uncached suffix and time to first token included), at contexts from 16.7K to 67.3K tokens. The longest generations ran at 73.6–83.8 tok/s end to end. The agent produced a self-contained oil-market report page with charts. The prices and events in it came from the model's web searches and have not been checked.
Recording of the page · the page itself · per-call log
Pagoda test: agentic coding on a 12 GB budget
The voxel-pagoda prompt (from Google AI's post), run through the pi coding agent with read/write/edit/bash tools, xhigh reasoning, ctx 204,800, the 12 GB budget, 3 seeds per model, a 45-minute wall clock and a 30-turn cap. Every comparator used the same 3-bit K / 2-bit V SJ-KVaRN cache.
| Model | Working pages (renders, no errors) |
|---|---|
| This model (2.3 bpw + MTP) | 3 / 3 |
| Qwen3.8-27B UD-IQ2_XXS (unsloth, 2.16 bpw) | 0 / 3 (no file written in any run) |
| Mirai fork 1.5 bpw + MTP | 0 / 3 (one page with a syntax error, one truncated script that never runs, one degenerate render) |
| Ternary Bonsai 2 27B PQ2_0 | 0 / 3 (two pages with syntax errors, one run wrote no file) |
All three of our runs wrote a working three.js scene and spent the remaining turns checking and polishing it until the turn cap.
| Seed 0 | Seed 1 | Seed 2 |
|---|---|---|
![]() |
![]() |
![]() |
| video | video | video |
Every page, video, comparator run and the full harness notes are in the dataset.
License and credits
Distributed under the Swift Open License v1.0, inherited from Swift 1.5. See LICENSE, LICENSE-APACHE-2.0, and NOTICE for the applicable terms and attribution. Credit goes to UkisAI for Swift 1.5, d0xin for the uncensored derivative, and the Qwen team for Qwen3.8-27B. The bit allocation is inspired by Mirai Labs' Mirai S and alesha-pro's llama.cpp port; no Mirai weights are included. The EXL3 format and quantizer are from exllamav3 by turboderp.


