← back to catalog · registered 2026-10-08 22:58

Dankpaws/Swift1.5-Qwen3.8-27B-Abliterated-EXL3-4.25bpw

Dankpaws 27B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/Dankpaws%2FSwift1.5-Qwen3.8-27B-Abliterated-EXL3-4.25bpw"
Response includes
  • classification unknown
  • files 23
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
?
Primary method

Unclassified

No clear signals of an abliteration technique in this model.
Confidence
UNKNOWN
Why this label 1 signal
No classification signals present. This may not be an abliterated model at all - it could be a repackaging, a merge with unrelated goals, or unrelated content that mentions the term.
  • no classification signals present (no abliterated, uncensored, or known producer/method markers)
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-10-08

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
other
Tags
exllamav3 safetensors qwen3_5 exl3 tabbyapi quantized swift abliterated coding reasoning vision tensor-parallel

Related

Total size
16.0 GB
Files
23
Quantizations
1
Registered
2026-10-08 22:58
Last updated on HF
2026-10-08 21:07

Files by quantization

Auxiliary files 23 files 16.1 GB
model-00002-of-00003.safetensors 8.00 GB 0ff8f47e download
model-00001-of-00003.safetensors 7.88 GB 205b7b5f download
model-00003-of-00003.safetensors 149 MB 51efe5a1 download
tokenizer.json 12.2 MB 0997f410 download
vocab.json 6.41 MB 0aa0ce06 download
merges.txt 3.20 MB a494e019 download
quantization_config.json 614 KB dd4944cd download
model.safetensors.index.json 289 KB 373b87ad download
benchmark-summary.json 33.8 KB 30d77d2b download
tokenizer_config.json 17.5 KB 5de744b3 download
LICENSE 13.0 KB 209a5720 download
LICENSE-APACHE-2.0 11.3 KB f938136e download
README.md 9.72 KB f7c7e6af download
chat_template.jinja 8.74 KB c0c686f9 download
config.json 4.54 KB 4bc0706d download
release-manifest.json 3.10 KB 4b2c1e6a download
NOTICE 1.93 KB 6b63b944 download
.gitattributes 1.53 KB 52373fe2 download
ABLITERATION_METADATA.json 1.53 KB 003a37d9 download
tabby-config.example.yml 676 B 3c9079b4 download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
generation_config.json 202 B 023756cf download

README current version from Hugging Face


base_model: grimlee/Swift-1.5-Qwen3.8-27B-Abliterated
base_model_relation: quantized
library_name: exllamav3
license: other
license_name: swift-open-license-1.0
license_link: LICENSE
pipeline_tag: text-generation
tags:

  • exl3
  • exllamav3
  • tabbyapi
  • quantized
  • swift
  • abliterated
  • coding
  • reasoning
  • vision
  • tensor-parallel

Swift 1.5 Qwen3.8-27B Abliterated · EXL3 4.25

Built for dual RTX 5060 Ti 16 GB cards. A calibrated EXL3 conversion of grimlee's abliterated Swift 1.5 Qwen3.8-27B, with a serving profile tuned for three concurrent long-context requests on 32 GB total VRAM. Native MTP and vision weights are included.

Specs
Target hardware 2× NVIDIA RTX 5060 Ti · 16 GB each
Format / runtime EXL3 · ExLlamaV3 1.5.2 / TabbyAPI
Quantization 4.25bpw decoder target · 6-bit head · 4-bit MTP · 5-bit vision
Tested concurrency 3 simultaneous requests
Configured context per request 115,200 tokens, including generated output
Shared cache 350,208 tokens · Q6 K/V
Weight shards 17.2 GB / 16.0 GiB, across 3 shards
Peak GPU memory in the test sweep 14,422 / 13,474 MiB used
Input Text; image support passed a basic smoke test
Speculative decoding Native MTP included · tested at 3 draft tokens

The hardware tuning is a quantization and serving-memory balance, not a new fine-tune or a custom 5060 Ti kernel. Vision weights are offloaded to system RAM in the tested profile, leaving more VRAM for the main model and shared cache. The test machine has 64 GB system RAM.

Performance on dual RTX 5060 Ti 16 GB

PCIe-limited test box: these results come from a Ryzen 7 5800X / Gigabyte B550 AORUS ELITE AX V2 system with very uneven GPU connectivity: one card on PCIe Gen3 x2 through the chipset, the other on Gen4 x8 CPU lanes. The x2 card also shares the chipset uplink with other devices. This is a poor topology for multi-GPU inference, especially tensor-parallel communication and large prompt processing. A properly provisioned dual-GPU system should perform better on communication-bound workloads; the improvement has not been measured, and no speedup multiplier is claimed. Treat these as results from this constrained box—not a performance ceiling for the quant or the GPUs.

Workload Server decode, per stream Client batch wall Aggregate output / client wall
Short single request 66.1 tok/s 15.64 s 65.5 tok/s
Three short requests 41.9–43.3 tok/s 24.90 s 123.4 tok/s
Three cold ~114k contexts 21.3–21.8 tok/s 780.76 s 3.9 tok/s, including cold prefill
Exact repeat of the three ~114k contexts 22.1–22.5 tok/s 50.87 s 60.4 tok/s

Measured October 8, 2026, with native tensor parallel, Q6 main/draft KV and MTP depth 3. Each request generated 1,024 tokens, with thinking disabled, temperature 0.6, top-p 0.95 and top-k 20. One sample per workload, after the server's startup warmup and smoke checks; these are synthetic throughput measurements, not coding-quality scores.

Each long request contained 113,890 input + 1,024 output = 114,914 total tokens. All six cold/repeated completions recalled their three unique early/middle/late markers. Cold prompt processing took 724–731 seconds per request under concurrent load, with first streamed output around 733 seconds. Repeats reused 113,664 input tokens per request and produced first output around 4.8 seconds. Cached performance is not cold-prefill performance.

The figures describe this complete serving setup, including speculative decoding and scheduling—not quantization alone. Per-stream server decode and aggregate client-wall throughput are intentionally reported separately.

Run it

Use TabbyAPI with a compatible ExLlamaV3 installation. This is EXL3, not GGUF, MLX or a directly loadable Transformers BF16 checkpoint.

Download the whole model repository into a directory named swift15-abliterated-exl3-4_25bpw under your TabbyAPI models directory. Adapt these fields in config.yml:

model:
  model_dir: ./models
  model_name: swift15-abliterated-exl3-4_25bpw
  backend: exllamav3
  max_seq_len: 115200
  cache_size: 350208
  cache_mode: 6,6
  tensor_parallel: true
  tensor_parallel_backend: native
  autosplit_reserve: [1600, 1600]
  chunk_size: 2048
  recurrent_checkpoint_interval_pp: 4096
  max_batch_size: 3
  vision: true
  vision_offload: true
  warmup: true
  tool_format: qwen3_coder
  reasoning: true
  template_vars_default:
    enable_thinking: true
    reasoning_effort: medium
draft_model:
  draft_mode: mtp
  draft_cache_mode: 6,6
  draft_num_tokens: 3
  dynamic_draft: true
memory:
  sysmem_recurrent_cache: 16384
  sysmem_kv_cache: 8192

The system-memory cache budgets above are in MiB. The shared GPU cache is not three permanently partitioned slots. dynamic_draft is enabled in the saved configuration, but the tested native-TP MTP path uses fixed depth rather than confidence-based draft shortening.

Start TabbyAPI using its project instructions, then connect an OpenAI-compatible client to your configured /v1 endpoint. Retrieve the model ID from /v1/models. Keep the server local or configure authentication before exposing it to a network.

Reserve room for the answer and reasoning: 115,200 is the total request budget, not the usable prompt length. A 16,384-token output allowance leaves at most 98,816 input tokens, including chat formatting and tools. The base model advertises 262,144 tokens; that is not the context validated here on these two cards.

For thinking-enabled use, start with medium effort, temperature 1.0, top-p 0.95, top-k 20 and min-p 0. The template supports low, medium and xhigh—not high. Reduce the shared cache and/or request context if your machine lacks memory headroom; other GPU workloads and image inputs can change memory demand.

Quantization

Converted directly from grimlee's abliterated BF16 weights, with no additional fine-tuning or new abliteration by Dankpaws.

Component Precision
Main decoder Mixed precision · 4.25bpw target
Language-model output head 6-bit
MTP weights 4-bit target
Vision weights 5-bit target
Token embeddings BF16
Small non-quantized tensors Floating point

Converter: ExLlamaV3 1.5.2, mul1 codebook, output scales enabled, calibration dimensions 250 × 2,048. No custom calibration dataset was supplied. 4.25 is the decoder quantization target, not the precision of every tensor or a measured average over the complete repository. See quantization_config.json for tensor-level storage details.

Validation, provenance and limitations

BF16 source: grimlee/Swift-1.5-Qwen3.8-27B-Abliterated, downloaded at revision 0e35abf5d32233d3b574ed1b8951264221b8872d. It is an abliterated derivative of UkisAI's Swift 1.5 Qwen3.8-27B, itself derived from Qwen3.8-27B.

The October 8 sweep completed in 2,198.8 seconds, with serving configuration unchanged. Sampled minimum free GPU memory was 1,520 / 2,468 MiB. A 128 MiB free-memory abort guard was used for the test; it is not a permanent serving watchdog or a guarantee for arbitrary workloads.

Additional checks on this quant:

  • A parsed tool call, exact tool-result marker round trip and red-image recognition passed.
  • Three small code/state-analysis tasks scored 1/3 without thinking, 3/3 with medium, and 3/3 with xhigh. Medium batch wall was 7.89 s versus 9.41 s for xhigh. These are synthetic smoke checks, not broad agent or coding benchmarks.
  • Nine ~100k-context completions across rotated no-thinking / medium / xhigh requests recalled their markers and stopped normally, with up to 8,192 generated tokens allowed per request.
  • Those mixed-effort runs also showed substantial scheduling stalls: cached streams sharing the server with cold long-prefill requests took 241–450 seconds end-to-end, with reported generation rates as low as 1.6–3.0 tok/s. They are not clean effort-speed comparisons. The headline throughput table uses the uniform-workload tests instead.

No paired BF16-versus-EXL3 task benchmark or next-token agreement evaluation has been completed for this release. HumanEval+, MBPP+ and MMLU-Pro scores in the upstream card describe its BF16 abliteration study; they are not scores for this quant. ABLITERATION_METADATA.json preserves upstream provenance, not evidence of post-quantization quality.

Long synthetic marker recall does not establish general long-context reasoning or autonomous-agent reliability. Image testing here is a basic smoke test; video, real-world vision accuracy and concurrent image-heavy workloads have not been validated.

License

Community quantization by Dankpaws; not an official grimlee, UkisAI or Qwen release.

Both bundled licenses apply:

  • Swift Open License v1.0 covers UkisAI's contribution; commercial use above its US$1 million gross-revenue threshold requires a separate enterprise license.
  • Apache License 2.0 covers the underlying Qwen material as identified in the source release.

Read the full license texts before use or redistribution. Retain upstream attribution in NOTICE and the abliteration metadata. This conversion replaces the BF16 weights with mixed-precision EXL3 tensors; this model card replaces the source README.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration