base_model: grimlee/Swift-1.5-Qwen3.8-27B-Abliterated
base_model_relation: quantized
library_name: exllamav3
license: other
license_name: swift-open-license-1.0
license_link: LICENSE
pipeline_tag: text-generation
tags:
- exl3
- exllamav3
- tabbyapi
- quantized
- swift
- abliterated
- coding
- reasoning
- vision
- tensor-parallel
Swift 1.5 Qwen3.8-27B Abliterated · EXL3 4.25
Built for dual RTX 5060 Ti 16 GB cards. A calibrated EXL3 conversion of grimlee's abliterated Swift 1.5 Qwen3.8-27B, with a serving profile tuned for three concurrent long-context requests on 32 GB total VRAM. Native MTP and vision weights are included.
| Specs | |
|---|---|
| Target hardware | 2× NVIDIA RTX 5060 Ti · 16 GB each |
| Format / runtime | EXL3 · ExLlamaV3 1.5.2 / TabbyAPI |
| Quantization | 4.25bpw decoder target · 6-bit head · 4-bit MTP · 5-bit vision |
| Tested concurrency | 3 simultaneous requests |
| Configured context per request | 115,200 tokens, including generated output |
| Shared cache | 350,208 tokens · Q6 K/V |
| Weight shards | 17.2 GB / 16.0 GiB, across 3 shards |
| Peak GPU memory in the test sweep | 14,422 / 13,474 MiB used |
| Input | Text; image support passed a basic smoke test |
| Speculative decoding | Native MTP included · tested at 3 draft tokens |
The hardware tuning is a quantization and serving-memory balance, not a new fine-tune or a custom 5060 Ti kernel. Vision weights are offloaded to system RAM in the tested profile, leaving more VRAM for the main model and shared cache. The test machine has 64 GB system RAM.
Performance on dual RTX 5060 Ti 16 GB
PCIe-limited test box: these results come from a Ryzen 7 5800X / Gigabyte B550 AORUS ELITE AX V2 system with very uneven GPU connectivity: one card on PCIe Gen3 x2 through the chipset, the other on Gen4 x8 CPU lanes. The x2 card also shares the chipset uplink with other devices. This is a poor topology for multi-GPU inference, especially tensor-parallel communication and large prompt processing. A properly provisioned dual-GPU system should perform better on communication-bound workloads; the improvement has not been measured, and no speedup multiplier is claimed. Treat these as results from this constrained box—not a performance ceiling for the quant or the GPUs.
| Workload | Server decode, per stream | Client batch wall | Aggregate output / client wall |
|---|---|---|---|
| Short single request | 66.1 tok/s | 15.64 s | 65.5 tok/s |
| Three short requests | 41.9–43.3 tok/s | 24.90 s | 123.4 tok/s |
| Three cold ~114k contexts | 21.3–21.8 tok/s | 780.76 s | 3.9 tok/s, including cold prefill |
| Exact repeat of the three ~114k contexts | 22.1–22.5 tok/s | 50.87 s | 60.4 tok/s |
Measured October 8, 2026, with native tensor parallel, Q6 main/draft KV and MTP depth 3. Each request generated 1,024 tokens, with thinking disabled, temperature 0.6, top-p 0.95 and top-k 20. One sample per workload, after the server's startup warmup and smoke checks; these are synthetic throughput measurements, not coding-quality scores.
Each long request contained 113,890 input + 1,024 output = 114,914 total tokens. All six cold/repeated completions recalled their three unique early/middle/late markers. Cold prompt processing took 724–731 seconds per request under concurrent load, with first streamed output around 733 seconds. Repeats reused 113,664 input tokens per request and produced first output around 4.8 seconds. Cached performance is not cold-prefill performance.
The figures describe this complete serving setup, including speculative decoding and scheduling—not quantization alone. Per-stream server decode and aggregate client-wall throughput are intentionally reported separately.
Run it
Use TabbyAPI with a compatible ExLlamaV3 installation. This is EXL3, not GGUF, MLX or a directly loadable Transformers BF16 checkpoint.
Download the whole model repository into a directory named swift15-abliterated-exl3-4_25bpw under your TabbyAPI models directory. Adapt these fields in config.yml:
model:
model_dir: ./models
model_name: swift15-abliterated-exl3-4_25bpw
backend: exllamav3
max_seq_len: 115200
cache_size: 350208
cache_mode: 6,6
tensor_parallel: true
tensor_parallel_backend: native
autosplit_reserve: [1600, 1600]
chunk_size: 2048
recurrent_checkpoint_interval_pp: 4096
max_batch_size: 3
vision: true
vision_offload: true
warmup: true
tool_format: qwen3_coder
reasoning: true
template_vars_default:
enable_thinking: true
reasoning_effort: medium
draft_model:
draft_mode: mtp
draft_cache_mode: 6,6
draft_num_tokens: 3
dynamic_draft: true
memory:
sysmem_recurrent_cache: 16384
sysmem_kv_cache: 8192
The system-memory cache budgets above are in MiB. The shared GPU cache is not three permanently partitioned slots. dynamic_draft is enabled in the saved configuration, but the tested native-TP MTP path uses fixed depth rather than confidence-based draft shortening.
Start TabbyAPI using its project instructions, then connect an OpenAI-compatible client to your configured /v1 endpoint. Retrieve the model ID from /v1/models. Keep the server local or configure authentication before exposing it to a network.
Reserve room for the answer and reasoning: 115,200 is the total request budget, not the usable prompt length. A 16,384-token output allowance leaves at most 98,816 input tokens, including chat formatting and tools. The base model advertises 262,144 tokens; that is not the context validated here on these two cards.
For thinking-enabled use, start with medium effort, temperature 1.0, top-p 0.95, top-k 20 and min-p 0. The template supports low, medium and xhigh—not high. Reduce the shared cache and/or request context if your machine lacks memory headroom; other GPU workloads and image inputs can change memory demand.
Quantization
Converted directly from grimlee's abliterated BF16 weights, with no additional fine-tuning or new abliteration by Dankpaws.
| Component | Precision |
|---|---|
| Main decoder | Mixed precision · 4.25bpw target |
| Language-model output head | 6-bit |
| MTP weights | 4-bit target |
| Vision weights | 5-bit target |
| Token embeddings | BF16 |
| Small non-quantized tensors | Floating point |
Converter: ExLlamaV3 1.5.2, mul1 codebook, output scales enabled, calibration dimensions 250 × 2,048. No custom calibration dataset was supplied. 4.25 is the decoder quantization target, not the precision of every tensor or a measured average over the complete repository. See quantization_config.json for tensor-level storage details.
Validation, provenance and limitations
BF16 source: grimlee/Swift-1.5-Qwen3.8-27B-Abliterated, downloaded at revision 0e35abf5d32233d3b574ed1b8951264221b8872d. It is an abliterated derivative of UkisAI's Swift 1.5 Qwen3.8-27B, itself derived from Qwen3.8-27B.
The October 8 sweep completed in 2,198.8 seconds, with serving configuration unchanged. Sampled minimum free GPU memory was 1,520 / 2,468 MiB. A 128 MiB free-memory abort guard was used for the test; it is not a permanent serving watchdog or a guarantee for arbitrary workloads.
Additional checks on this quant:
- A parsed tool call, exact tool-result marker round trip and red-image recognition passed.
- Three small code/state-analysis tasks scored 1/3 without thinking, 3/3 with medium, and 3/3 with xhigh. Medium batch wall was 7.89 s versus 9.41 s for xhigh. These are synthetic smoke checks, not broad agent or coding benchmarks.
- Nine ~100k-context completions across rotated no-thinking / medium / xhigh requests recalled their markers and stopped normally, with up to 8,192 generated tokens allowed per request.
- Those mixed-effort runs also showed substantial scheduling stalls: cached streams sharing the server with cold long-prefill requests took 241–450 seconds end-to-end, with reported generation rates as low as 1.6–3.0 tok/s. They are not clean effort-speed comparisons. The headline throughput table uses the uniform-workload tests instead.
No paired BF16-versus-EXL3 task benchmark or next-token agreement evaluation has been completed for this release. HumanEval+, MBPP+ and MMLU-Pro scores in the upstream card describe its BF16 abliteration study; they are not scores for this quant. ABLITERATION_METADATA.json preserves upstream provenance, not evidence of post-quantization quality.
Long synthetic marker recall does not establish general long-context reasoning or autonomous-agent reliability. Image testing here is a basic smoke test; video, real-world vision accuracy and concurrent image-heavy workloads have not been validated.
License
Community quantization by Dankpaws; not an official grimlee, UkisAI or Qwen release.
Both bundled licenses apply:
- Swift Open License v1.0 covers UkisAI's contribution; commercial use above its US$1 million gross-revenue threshold requires a separate enterprise license.
- Apache License 2.0 covers the underlying Qwen material as identified in the source release.
Read the full license texts before use or redistribution. Retain upstream attribution in NOTICE and the abliteration metadata. This conversion replaces the BF16 weights with mixed-precision EXL3 tensors; this model card replaces the source README.