← back to catalog · registered 2026-08-22 13:56

hotdogs/Qwen3.8-27B-abliterated-cyber-preview

hotdogs Qwen 28B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/hotdogs%2FQwen3.8-27B-abliterated-cyber-preview"
Response includes
  • classification m1
  • files 11
  • hub_downloads_all_time 1,339
  • author_summary 25 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
1K
324 last 30d - stable
Likes
3
Descendants
1
in 1 direct fork
Model age
7w ago
created 2026-08-19

Training datasets

1 of 1 in /datasets

Corpora the author lists in the model card. Datasets tracked in our /datasets catalog carry a category badge linking to the workflow stage. Others open on Hugging Face.

Downloads over time
Now1.4K→from49↑2,798%
05191K1.6K49 on Aug 191.4K on Oct 11AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Genealogy 1 direct fork

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh
Tags
transformers safetensors qwen3_5 image-text-to-text abliterated cybersecurity offensive-security agentic tool-calling mcp conversational en

Related

Total size
51.7 GB
Files
11
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-21 03:31

Files by quantization

Auxiliary files 11 files 51.8 GB
model-00001-of-00002.safetensors 46.4 GB d1597122 download
model-00002-of-00002.safetensors 5.34 GB a1079980 download
tokenizer.json 19.1 MB 87a7830d download
model.safetensors.index.json 139 KB bd795d8f download
chat_template.jinja 8.74 KB c0c686f9 download
README.md 7.79 KB 03c485d0 download
tokenizer_config.json 6.99 KB 773a6327 download
config.json 3.71 KB 3f2a4d7c download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.16 KB 33818c7f download
generation_config.json 213 B 79c8cce3 download

README current version from Hugging Face


license: apache-2.0
base_model: hotdogs/Qwen3.8-27B-abliterated
base_model_relation: finetune
pipeline_tag: image-text-to-text
library_name: transformers
tags:

  • qwen3_5
  • abliterated
  • cybersecurity
  • offensive-security
  • agentic
  • tool-calling
  • mcp
  • image-text-to-text
  • conversational
    language:
  • en
  • zh
    datasets:
  • hotdogs/cyber-sft-agent-qwen38

Qwen3.8-27B-Abliterated-Cyber-Preview

An offensive-security / agentic tool-calling LoRA merge on top of
hotdogs/Qwen3.8-27B-abliterated
— the training-free abliterated build of
Qwen/Qwen3.8-27B. The LoRA is
merged into the base at scale 1.0 (PEFT alpha/r = 64/32 = 2.0, i.e.
the exact strength the scale sweep below found optimal), and the MTP
(Multi-Token Prediction) head is preserved
, so the GGUF build supports
self-speculative decoding.

This is a preview: the LoRA was trained on an agentic dataset that mixes
penetration-testing tool calls (<tool_call><function=...> XML) with
security QA. This model will not refuse. It is published for authorized
security research, red-team exercises, and studying LLM-driven tool use
—
pentesting only systems you own or have explicit permission to test. You are
responsible for your use of it and for complying with all applicable laws.
Inherits the Apache-2.0 license of the base.


Quick Results

All numbers below were measured on our own reproduction pipeline (see
Method and Evaluation). The merged model adds a reliable, correctly
formatted tool-call pathway while leaving general capability essentially
unchanged.

metric base (no LoRA) this repo (merged @ scale 1.0)
Tool-call format emitted (6 pentest prompts) 0/6 · 0 % 6/6 · 100 %
Correct real-tool selection (same 6 prompts) 0/6 · 0 % 6/6 · 100 % (nmap, ffuf, masscan, sqlmap, wpscan, smbclient)
General capability (7 QA/math/code prompts) 6/7 · 86 % 7/7 · 100 %
KL divergence (base ‖ merged), base-prompts — 0.041 (essentially zero)
KL divergence (base ‖ merged), tool-prompts — 0.808 (the intended re-target)
MTP draft acceptance rate — 0.77 (51/66)

Reading the numbers: the merge shifts the model's next-token distribution by a
tiny amount on general prompts (KL ≈ 0.04 → base knowledge is preserved), while
shifting it ~20× more on security/tool prompts (KL ≈ 0.81 → the intended
behavioral re-target to emit tool calls). The scale sweep and MTP numbers are
detailed below.


What changed

The base model already refuses nothing (abliterated). This repo adds a
tool-calling capability: given a pentest scenario, the model now emits a
structured, well-formed agentic call instead of free-form text.

User: Port scan the host 203.0.113.10 and identify which services are exposed.

<tool_call>
<function=nmap>
<parameter=target>
203.0.113.10
</parameter>
<parameter=ports>
-top 1000
</parameter>
</function>
</tool_call>

The tool-call format follows the Qwen3.5 native <tool_call>/<function>/<parameter>
schema (also exercised through the model's chat template). At scale 1.0 the
model picked the correct real tool per scenario in 6/6 cases.


Method

Base

hotdogs/Qwen3.8-27B-abliterated (λ = 1.2), a 27B native vision-language
hybrid (full-attention + GDN linear-attention) with a frozen MTP head.

LoRA

  • Dataset: hotdogs/cyber-sft-agent-qwen38 — 8,400 rows
    (train 7,140 / valid 840 / test 420). 30 % tool-call, 30 % multi-turn,
    think-tags 100 % balanced, 22 tools (nmap, sqlmap, metasploit, hydra,
    crackmapexec, linpeas, wpscan, whatweb, ffuf, gobuster, …), 1,995 unique IP
    contexts, 0 duplicates.
  • Config: r=32, alpha=64, dropout=0, ctx=8192, LR=1e-4 (cosine),
    BF16 (required by GDN hybrid), batch 2 × grad-accum 2 (effective 4),
    2 epochs. Trainable 233M (0.85 %).
  • Target modules: standard q/k/v/o/gate/up/down + GDN
    in_proj_qkv/out_proj/in_proj_z/a/b.

Merge

  • Full LoRA (incl. GDN) merged into the base via PEFT merge_and_unload() at
    default scale (alpha/r = 2.0, ≡ llama.cpp --lora-scaled :1.0).
  • MTP head recovered after merge: merge_and_unload() + save_pretrained()
    drops the mtp.* tensors, so the 15 MTP tensors (849 MB) were copied back
    from the base into the merged safetensors (the LoRA never touched mtp.*,
    so these weights are byte-correct). Final model = 1,199 tensors, verified
    loading via unsloth, MTP physically present.
  • GGUF: convert_hf_to_gguf.py --outtype bf16 (MTP kept in-file), 54.6 GB.
    Serves with --spec-type draft-mtp.

Evaluation

1. Overfitting check (held-out valid loss)

metric value
valid token loss (840 rows) 0.480
train token loss (300-row sample) 0.384
gap (valid − train) 0.095
valid % loss < 0.2 30.0 %
train % loss < 0.2 27.3 %

A small gap (≤ 0.15) means the model generalizes rather than memorizes.
The training loss plateau (~0.44 median, only ~3.8 % of points near zero) is a
healthy learning floor for a diverse 8.4K dataset, not overfitting.

2. Scale sweep (llama.cpp --lora-scaled, no merge needed)

6 tool prompts + 7 base prompts, temperature 0, identical config across scales:

scale tool-call format (6) correct real tool (6) base capability (7)
baseline (no LoRA) 0/6 0/6 6/7
0.25 1/6 0/6 7/7
0.5 6/6 2/6 7/7
0.75 6/6 4/6 7/7
1.0 6/6 6/6 7/7
1.5 6/6 6/6 7/7

Scale 1.0 is the sweet spot: every tool prompt emits a well-formed
<tool_call> and picks the correct real tool, while base capability stays
at 7/7. Damping below 1.0 hurts tool selection (hallucinated tools like
web_fetch, wp_recon); going above ~2.0 degrades parameter quality (e.g.
ports="-s -p 1-65535"). The merge therefore uses the default
alpha/r = 2.0, which corresponds to this optimal 1.0 point.

3. KL divergence (base ‖ merged)

Measured on a shared continuation that the merged model sampled, comparing
top-25 next-token log-probs from the base vs the merged GGUF (both via
llama.cpp llama-server):

prompt KL (base ‖ merged)
17 × 43 0.0066
Python reverse-string 0.1559
Capital of France 0.0206
stack vs queue 0.0556
15 % of 200 0.0030
boiling point 0.0041
mean (base prompts) 0.0410
port-scan tool prompt 0.7332
SQLi tool prompt 0.8833
mean (tool prompts) 0.8083
overall mean 0.2328

The ~20× gap between base-prompts (0.04) and tool-prompts (0.81) is the
quantitative signature of a surgically targeted LoRA: it re-wires
tool-calling behavior without disturbing general knowledge. This matches the
abliterated base's own near-zero first-token KL (0.0001) — the merge adds
capability, not drift.

4. MTP (Multi-Token Prediction)

metric value
MTP draft acceptance rate 0.77 (51 / 66)
mean draft length 2.55
head location in-file (bf16 GGUF)

Well above the ~0.3 threshold where speculative decoding pays off — MTP is
active and beneficial.


Files

file purpose
model-00001-of-00002.safetensors / model-00002-of-00002.safetensors merged weights (MTP in shard 2)
config.json Qwen3_5ForConditionalGeneration
chat_template.jinja / tokenizer.json / tokenizer_config.json / processor_config.json Qwen3.5 chat + vision template

Disclaimer

This is a preview built for authorized security research and red-teaming.
It will not refuse and may emit instructions for exploiting systems. Use only
on systems you own or are explicitly authorized to test. Not for
misuse/harmful activity. Apache-2.0, inherited from the base.

README history 5 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-19Replace private LAN IP with doc-safe TEST-NET IP (203.0.113.10)efd0c587.8 KB
    Loading...
  2. 2026-08-19Revert KL table to per-prompt format; add dataset hotdogs/cyber-sft-agent-qwe...a8514347.8 KB
    Loading...
  3. 2026-08-19Add base-vs-merged top-1 comparison to KL divergence table8fb38e18.5 KB
    Loading...
  4. 2026-08-19Remove Reproduction section (GGUF not yet in repo)e7edc657.8 KB
    Loading...
  5. 2026-08-19Add detailed model card: eval, KL divergence, scale sweep, MTP, reproduction1aa4f8b8.7 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration