license: apache-2.0
base_model:
- Kujira/Underdog-Saluki-27B-1.0-MTP-GGUF
- ConwayResearch/Underdog-Saluki-27B-1.0
base_model_relation: finetune
pipeline_tag: text-generation
language: - en
- ja
library_name: gguf
tags: - gguf
- llama.cpp
- 2-bit
- mtp
- speculative-decoding
- abliterated
- uncensored
- tool-calling
- agents
- qwen3.8
Underdog Saluki 27B 1.0 + MTP, Abliterated (GGUF)
Underdog Saluki 27B 1.0 + MTP with
refusals removed. It answers requests that the original model refuses, and otherwise behaves
close to the original: same file size, same speed, MTP self-speculative decoding still works.
| File | Underdog-Saluki-27B-1.0-IQ2-mix-MTP-abliterated.gguf, 8.35 GB |
| sha256 | b9b231748c94e13ecc0f841942798eb36b27eaea60014fea896218bc86742bbc |
| Changed | 4 tensors in blocks 35–36 (blk.35.attn_output, blk.35.ffn_down, blk.36.ffn_down, blk.36.ssm_out), kept at their original quantization types and sizes |
| Unchanged | the other 862 tensors (including the MTP head blk.64.*) and all metadata are byte-identical to Underdog-Saluki-27B-1.0-IQ2-mix-MTP.gguf (sha256 98f6ebb5…e89e52) |
No training was done. The method is not published.
Usage
Same as the MTP release: stock llama.cpp with MTP drafting (--spec-type draft-mtp).
llama-server -m Underdog-Saluki-27B-1.0-IQ2-mix-MTP-abliterated.gguf --jinja -ngl 99 -fa on \
-c 40960 -np 1 -ctk q8_0 -ctv q8_0 \
--temp 0.6 --top-p 0.95 --top-k 20 --presence-penalty 1.5 \
--spec-type draft-mtp --spec-draft-n-max 2
--presence-penalty 1.5is recommended. Without it, long thinking sometimes got stuck repeating
itself in agent use; with it, no loops were seen.- 12 GB cards: 40K context fits (11.4 GB with a 22K-token prompt). 64K starts at 11.85 GB and grows
about 0.45 GB with long prompts, spills out of VRAM and slows down sharply.
Measured (one machine)
RTX 3080 12 GB, WSL2, llama.cpp (PrismML fork prism-b10754). Refusal = the reply contains a
refusal phrase. Harmful prompts: the 104-prompt test split of a public harmful-behaviors set.
| Original (MTP release) | This file | |
|---|---|---|
| Refusal, harmful prompts (64 tokens, thinking off) | 99% | 1% |
| Refusal, harmful prompts (thinking on, up to 1,536 tokens) | 78% | 2% |
| Refusal, harmless prompts (thinking on) | 0% | 0% |
| Perplexity, English (wikitext-2 test, 2048 × 40 chunks) | 6.558 | 6.662 (+1.6%) |
| Perplexity, Japanese (wiki40b-ja test, 2048 × 40 chunks) | 11.21 | 11.41 (+1.8%) |
| Long-form, 24 prompts (en/ja, thinking on): looping / empty answer | 0% / 8% | 0% / 0% |
| Tool calling, 30 prompts with 10 tools: right tool / valid arguments | 97% / 100% | 100% / 100% |
| Answer after a (fake) tool result, 30 prompts | 100% | 100% |
| No tool call when none is needed, 10 prompts | 90% | 80% |
- Speed and VRAM are the same as the MTP release (same tensor types and size): about 58–62 tok/s
with MTP on at 40K context. - Thinking length on everyday tasks is about the same as the original (mean 1,835 vs 2,059
characters on the long-form prompts). - These are small, single-machine tests, not a benchmark. The quality numbers on Saluki's card are
for the original model. - The evaluation prompts and scripts are available on request in the Discussions tab.
Responsible use
Intended for research and uncensored local use. This model does not refuse harmful requests. You are responsible for how you use it and for any
content it produces. Do not put it in front of users without your own safeguards.
Credits and license
Apache-2.0, see LICENSE and NOTICE.
- Underdog Saluki 27B 1.0 by Underdog (ConwayResearch), built on Qwen3.8-27B and ISTA-DASLab's Qwen3.8-27B-GSQ-RCO-GGUF.
- Qwen3.8-27B and its MTP head by the Qwen team.
- MTP head GGUF quantization by Unsloth (
unsloth/Qwen3.8-27B-GGUF), via the
MTP release (see its card).
Not affiliated with or endorsed by Underdog/ConwayResearch, Qwen, Unsloth, ISTA-DASLab, PrismML or BoldingBuilds.
日本語メモ
Underdog Saluki 27B 1.0 + MTP の拒否を外した版です。
変えたのは第35・36層の4テンソルだけで、量子化の種類とサイズは元のままです。残りの 862 テンソルと MTP ヘッド、メタデータは元とバイト単位で同一です。学習はしていません。方法は公開していません。
- 拒否率:有害な質問で 99% → 1%(thinking ありでは 78% → 2%)。無害な質問の拒否は 0%
- perplexity の悪化:英語 +1.6%、日本語 +1.8%
- 長文のループ 0%、ツール呼び出しは元と同等。速度と VRAM は MTP 版と同じです
--presence-penalty 1.5を付けて使うのがおすすめです。12GB のカードでは文脈 40K まで(64K は VRAM が溢れて遅くなります)
研究と、手元で検閲なしに使う用途向けです。有害な依頼も断りません。使い方と出力の責任は利用者にあります。
評価に使った問題とスクリプトは、Discussions で問い合わせてもらえれば共有します。