← back to catalog · registered 2026-08-22 13:56

hotdogs/Qwen3.8-27B-abliterated-sme-preview

hotdogs Qwen 28B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/hotdogs%2FQwen3.8-27B-abliterated-sme-preview"
Response includes
  • classification m1
  • files 23
  • hub_downloads_all_time 772
  • author_summary 25 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
772
76 last 30d - cooling
Likes
0
Descendants
1
in 1 direct fork
Model age
7w ago
created 2026-08-22

Training datasets

1 of 1 in /datasets

Corpora the author lists in the model card. Datasets tracked in our /datasets catalog carry a category badge linking to the workflow stage. Others open on Hugging Face.

Downloads over time
Now790→from0↑0%
02905798690 on Aug 19790 on Oct 11790 on Oct 10AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Genealogy 1 direct fork

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en th
Tags
safetensors qwen3_5 qwen3.5 gdn linear-attention hybrid sft lora fine-tuned sme tool-calling agent

Related

Total size
51.7 GB
Files
23
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-22 12:56

Files by quantization

Auxiliary files 23 files 51.8 GB
model-00005-of-00012.safetensors 4.64 GB 101bcbe4 download
model-00008-of-00012.safetensors 4.63 GB a1cd1b42 download
model-00011-of-00012.safetensors 4.62 GB acc81604 download
model-00003-of-00012.safetensors 4.62 GB 65fc10a3 download
model-00009-of-00012.safetensors 4.62 GB 1eb911de download
model-00010-of-00012.safetensors 4.59 GB fa86ab73 download
model-00007-of-00012.safetensors 4.59 GB f7cc677f download
model-00006-of-00012.safetensors 4.58 GB 76200a65 download
model-00004-of-00012.safetensors 4.58 GB 5e41d54d download
model-00002-of-00012.safetensors 4.51 GB bc4d3eb1 download
model-00001-of-00012.safetensors 3.16 GB 6ddc4fdb download
model-00012-of-00012.safetensors 2.60 GB 92520aa7 download
tokenizer.json 19.1 MB 87a7830d download
model.safetensors.index.json 114 KB 14ac1502 download
chat_template.jinja 8.74 KB c0c686f9 download
README.md 7.89 KB 769aaece download
tokenizer_config.json 6.95 KB 58c65b5d download
config.json.bak_unsloth 3.71 KB d3c457ff download
config.json 3.58 KB c776ed35 download
.gitattributes 1.53 KB 52373fe2 download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
generation_config.json 202 B 023756cf download

README current version from Hugging Face


language:

  • en
  • th
    tags:
  • qwen3.5
  • qwen3_5
  • gdn
  • linear-attention
  • hybrid
  • sft
  • lora
  • fine-tuned
  • sme
  • tool-calling
  • agent
  • business
  • multimodal
    license: apache-2.0
    base_model: hotdogs/Qwen3.8-27B-abliterated
    datasets:
  • hotdogs/sme-sft-docdata-qwen38
    pipeline_tag: text-generation

Qwen3.8-27B-Abliterated-SME-Preview

An SME (small & medium enterprise) business assistant built on top of the
Qwen3.8-27B abliterated (λ=1.2) hybrid reasoning model — fine-tuned with LoRA
on a curated document-analysis + tool-calling dataset. It is designed to reason
over CSV/office documents, call tools correctly, and answer SME business
questions, while preserving the base model's general capability.

  • Base: hotdogs/Qwen3.8-27B-abliterated (λ=1.2, Qwen3_5ForConditionalGeneration)
  • Method: LoRA SFT (r=32, α=64, BF16), 2 epochs
  • Training data: hotdogs/sme-sft-docdata-qwen38 (8,827 rows, 30% tool-call, 100% think-tagged, 0% duplicates)
  • Format: full BF16 merged safetensors (MTP head preserved), multimodal wrapper

⚡ Quick Results (A/B vs base — same harness, same prompts)

Measured with lm-eval HF backend on the same 7× RTX 3090 machine, identical
flags both sides (no chat-template on MC tasks — required for thinking models).

Benchmark Base SME-Preview Δ
ARC-Challenge (0-shot, 300) acc 0.5667 0.5700 +0.003
ARC-Challenge acc_norm 0.5733 0.5733 0.000
MMLU (0-shot, 200) 0.8477 0.8449 −0.003
GSM8K (5-shot, strict) 0.6000 0.8100 +0.210
GSM8K (5-shot, flexible) 0.6500 0.8100 +0.160

How to read this

  • ARC / MMLU Δ ≈ 0 → the SME fine-tune does not degrade general
    knowledge or science reasoning — capability is preserved at ~100%.
  • GSM8K +21 pts → math / step-by-step reasoning improved significantly,
    thanks to the think-tagged, reasoning-heavy training data.

Sanity check: the base model's own ARC (0.5667) and MMLU (0.8477) match the
known-good baseline for this model family, confirming the harness was correct
(not a below-chance artifact).

KL divergence (base vs merged)

Test prompt KL
"The capital of France is" 0.13
"A CSV file is used for" 0.10
"To sum a column of numbers, you" 0.04
"The best way to back up data is" 0.06
"In cybersecurity, a firewall" 0.07
"Sales increased because" 0.15

Low KL (0.04–0.15) + argmax agreement on all prompts = the merge is faithful;
no catastrophic shift from fine-tuning.


🛠️ Tool-Calling (validated)

Given a document/table query, the model reasons first, then emits a correctly
formatted <tool_call>. Verified output:

User: Which rows in /reports/expenses_2025.csv have department = Sales?

<tool_call>
<function=csv_filter>
<parameter=path>
/reports/expenses_2025.csv
</parameter>
<parameter=column>
department
</parameter>
<parameter=op>
eq
</parameter>
<parameter=value>
Sales
</parameter>
</function>
</tool_call>
  • ✅ Correct <tool_call> / <function> / <parameter> structure
  • ✅ No infinite loop — exactly one tool call per decision
  • ✅ Reasoning precedes each tool call

🧠 Architecture

  • Class: Qwen3_5ForConditionalGeneration (multimodal wrapper, text-only use here)
  • 64 hidden layers = 48 linear-attention (GDN) + 16 full-attention
    (pattern: 3 GDN + 1 full, repeated)
  • MTP head preserved (mtp_num_hidden_layers=1, 15 MTP tensors)
  • BF16 — required for GDN (linear_attn) layers (FP16 → NaN grad norms)

🔬 Fine-Tuning Details

Hyperparameter Value
LoRA r / α / dropout 32 / 64 / 0.0
Target modules q,k,v,o + gate,up,down + GDN (in_proj_qkv,out_proj,in_proj_z,in_proj_a,in_proj_b)
Context 8192
Precision BF16
Epochs 2 (resumed from checkpoint)
LR / scheduler 1e-4, cosine
Batch (per-GPU × grad-accum) 2 × 2 (eff 4) across 7 GPUs
Trainable params ~233M (0.85%)
Final loss (epoch 2 median) 0.048

Dataset composition (hotdogs/sme-sft-docdata-qwen38):

  • 8,827 rows (7,799 train / 1,028 valid)
  • 30% tool-call rows, 70% QA rows
  • 100% assistant messages have reasoning (think-tagged)
  • 82.7% evidence-chained answers, 0% byte-exact duplicates
  • Tools: csv_aggregate, csv_filter, csv_head, csv_stats, doc_compare,
    doc_metadata, doc_read, doc_search, doc_summarize, and more

🚀 Usage

Transformers (Python)

import torch
from transformers import AutoTokenizer, AutoModelForImageTextToText

MODEL = "hotdogs/Qwen3.8-27B-abliterated-sme-preview"
tok = AutoTokenizer.from_pretrained(MODEL, trust_remote_code=True)
model = AutoModelForImageTextToText.from_pretrained(
    MODEL, torch_dtype=torch.bfloat16, device_map="auto", trust_remote_code=True
)

prompt = tok.apply_chat_template(
    [{"role": "user", "content": "Which rows in expenses_2025.csv have department = Sales?"}],
    tokenize=False, add_generation_prompt=True
)
ids = tok(prompt, return_tensors="pt").to("cuda")["input_ids"]
out = model.generate(ids, max_new_tokens=256, do_sample=False,
                     pad_token_id=tok.pad_token_id, eos_token_id=tok.eos_token_id)
print(tok.decode(out[0][ids.shape[1]:], skip_special_tokens=True))

Note: this is a multimodal model (vision wrapper present). To use it as a
plain text assistant, load with AutoModelForImageTextToText and pass text-only
messages as above. Do not call tokenizer(text) on the raw processor with a
plain string in a way that routes to the vision path — use apply_chat_template.

GGUF / llama.cpp

A f16 GGUF with the MTP head intact is available. Convert instructions for
custom quants (Q4_K_M / IQ3_M etc.):

PYTHONPATH=/path/to/llama.cpp/gguf-py python3 /path/to/llama.cpp/convert_hf_to_gguf.py \
  <HF_download_dir> --outfile model-f16.gguf --outtype f16
# then quantize, e.g.:
/path/to/llama.cpp/llama-quantize model-f16.gguf model-Q4_K_M.gguf Q4_K_M

Requires a llama.cpp build that supports the qwen35 architecture (GDN +
linear-attention). Keep MTP (do not pass --no-mtp — the model has MTP
tensors, block_count must be 64 + nextn).


📁 Files

File Size Description
model-00001-of-00012.safetensors … -00012 ~54 GB BF16 merged weights (12 shards)
model.safetensors.index.json — Shard index
config.json — Model config (MTP=1, text_config restored)
tokenizer.json / tokenizer_config.json — Qwen3.5 tokenizer
preprocessor_config.json / video_preprocessor_config.json — Processor (multimodal)
chat_template.jinja — Chat template (thinking + tool-call)
generation_config.json — Generation defaults

⚠️ Notes & Limitations

  • Preview release — validated on ARC/MMLU/GSM8K + tool-call smoke tests, not
    yet on full agentic benchmarks (e.g. Terminal-Bench / SWE-Bench).
  • Abliterated base: safety-tuning refusals are reduced by design (λ=1.2 kept
    capability while cutting refusals 98%→39%). Exercise judgment for harmful use.
  • Document tools (csv_*, doc_*) are emitted as structured calls — you
    must wire them to a runtime (e.g. a function-calling agent loop) to actually
    execute them.
  • 27B BF16 needs a multi-GPU setup (or a quantized GGUF) for practical inference.

📜 License

Apache-2.0. Base model: hotdogs/Qwen3.8-27B-abliterated
(λ=1.2 abliteration); fine-tune method follows the Train-Studio (unsloth + HF Trainer) recipe.


Built with ❤️ on 7× RTX 3090 — SFT via unsloth + HF Trainer, merged with MTP
preserved, benchmarked A/B against base.

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-22Fix base_model in YAML + link abliterated base repo345ba447.9 KB
    Loading...
  2. 2026-08-22Add comprehensive README with A/B benchmark results7a6e4217.8 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration