language:
- en
- th
tags: - qwen3.5
- qwen3_5
- gdn
- linear-attention
- hybrid
- sft
- lora
- fine-tuned
- sme
- tool-calling
- agent
- business
- multimodal
license: apache-2.0
base_model: hotdogs/Qwen3.8-27B-abliterated
datasets: - hotdogs/sme-sft-docdata-qwen38
pipeline_tag: text-generation
Qwen3.8-27B-Abliterated-SME-Preview
An SME (small & medium enterprise) business assistant built on top of the
Qwen3.8-27B abliterated (λ=1.2) hybrid reasoning model — fine-tuned with LoRA
on a curated document-analysis + tool-calling dataset. It is designed to reason
over CSV/office documents, call tools correctly, and answer SME business
questions, while preserving the base model's general capability.
- Base:
hotdogs/Qwen3.8-27B-abliterated(λ=1.2,Qwen3_5ForConditionalGeneration) - Method: LoRA SFT (r=32, α=64, BF16), 2 epochs
- Training data:
hotdogs/sme-sft-docdata-qwen38(8,827 rows, 30% tool-call, 100% think-tagged, 0% duplicates) - Format: full BF16 merged safetensors (MTP head preserved), multimodal wrapper
⚡ Quick Results (A/B vs base — same harness, same prompts)
Measured with lm-eval HF backend on the same 7× RTX 3090 machine, identical
flags both sides (no chat-template on MC tasks — required for thinking models).
| Benchmark | Base | SME-Preview | Δ |
|---|---|---|---|
| ARC-Challenge (0-shot, 300) acc | 0.5667 | 0.5700 | +0.003 |
| ARC-Challenge acc_norm | 0.5733 | 0.5733 | 0.000 |
| MMLU (0-shot, 200) | 0.8477 | 0.8449 | −0.003 |
| GSM8K (5-shot, strict) | 0.6000 | 0.8100 | +0.210 |
| GSM8K (5-shot, flexible) | 0.6500 | 0.8100 | +0.160 |
How to read this
- ARC / MMLU Δ ≈ 0 → the SME fine-tune does not degrade general
knowledge or science reasoning — capability is preserved at ~100%. - GSM8K +21 pts → math / step-by-step reasoning improved significantly,
thanks to the think-tagged, reasoning-heavy training data.
Sanity check: the base model's own ARC (0.5667) and MMLU (0.8477) match the
known-good baseline for this model family, confirming the harness was correct
(not a below-chance artifact).
KL divergence (base vs merged)
| Test prompt | KL |
|---|---|
| "The capital of France is" | 0.13 |
| "A CSV file is used for" | 0.10 |
| "To sum a column of numbers, you" | 0.04 |
| "The best way to back up data is" | 0.06 |
| "In cybersecurity, a firewall" | 0.07 |
| "Sales increased because" | 0.15 |
Low KL (0.04–0.15) + argmax agreement on all prompts = the merge is faithful;
no catastrophic shift from fine-tuning.
🛠️ Tool-Calling (validated)
Given a document/table query, the model reasons first, then emits a correctly
formatted <tool_call>. Verified output:
User: Which rows in /reports/expenses_2025.csv have department = Sales?
<tool_call>
<function=csv_filter>
<parameter=path>
/reports/expenses_2025.csv
</parameter>
<parameter=column>
department
</parameter>
<parameter=op>
eq
</parameter>
<parameter=value>
Sales
</parameter>
</function>
</tool_call>
- ✅ Correct
<tool_call>/<function>/<parameter>structure - ✅ No infinite loop — exactly one tool call per decision
- ✅ Reasoning precedes each tool call
🧠 Architecture
- Class:
Qwen3_5ForConditionalGeneration(multimodal wrapper, text-only use here) - 64 hidden layers = 48 linear-attention (GDN) + 16 full-attention
(pattern: 3 GDN + 1 full, repeated) - MTP head preserved (
mtp_num_hidden_layers=1, 15 MTP tensors) - BF16 — required for GDN (
linear_attn) layers (FP16 → NaN grad norms)
🔬 Fine-Tuning Details
| Hyperparameter | Value |
|---|---|
| LoRA r / α / dropout | 32 / 64 / 0.0 |
| Target modules | q,k,v,o + gate,up,down + GDN (in_proj_qkv,out_proj,in_proj_z,in_proj_a,in_proj_b) |
| Context | 8192 |
| Precision | BF16 |
| Epochs | 2 (resumed from checkpoint) |
| LR / scheduler | 1e-4, cosine |
| Batch (per-GPU × grad-accum) | 2 × 2 (eff 4) across 7 GPUs |
| Trainable params | ~233M (0.85%) |
| Final loss (epoch 2 median) | 0.048 |
Dataset composition (hotdogs/sme-sft-docdata-qwen38):
- 8,827 rows (7,799 train / 1,028 valid)
- 30% tool-call rows, 70% QA rows
- 100% assistant messages have reasoning (think-tagged)
- 82.7% evidence-chained answers, 0% byte-exact duplicates
- Tools:
csv_aggregate,csv_filter,csv_head,csv_stats,doc_compare,doc_metadata,doc_read,doc_search,doc_summarize, and more
🚀 Usage
Transformers (Python)
import torch
from transformers import AutoTokenizer, AutoModelForImageTextToText
MODEL = "hotdogs/Qwen3.8-27B-abliterated-sme-preview"
tok = AutoTokenizer.from_pretrained(MODEL, trust_remote_code=True)
model = AutoModelForImageTextToText.from_pretrained(
MODEL, torch_dtype=torch.bfloat16, device_map="auto", trust_remote_code=True
)
prompt = tok.apply_chat_template(
[{"role": "user", "content": "Which rows in expenses_2025.csv have department = Sales?"}],
tokenize=False, add_generation_prompt=True
)
ids = tok(prompt, return_tensors="pt").to("cuda")["input_ids"]
out = model.generate(ids, max_new_tokens=256, do_sample=False,
pad_token_id=tok.pad_token_id, eos_token_id=tok.eos_token_id)
print(tok.decode(out[0][ids.shape[1]:], skip_special_tokens=True))
Note: this is a multimodal model (vision wrapper present). To use it as a
plain text assistant, load withAutoModelForImageTextToTextand pass text-only
messages as above. Do not calltokenizer(text)on the raw processor with a
plain string in a way that routes to the vision path — useapply_chat_template.
GGUF / llama.cpp
A f16 GGUF with the MTP head intact is available. Convert instructions for
custom quants (Q4_K_M / IQ3_M etc.):
PYTHONPATH=/path/to/llama.cpp/gguf-py python3 /path/to/llama.cpp/convert_hf_to_gguf.py \
<HF_download_dir> --outfile model-f16.gguf --outtype f16
# then quantize, e.g.:
/path/to/llama.cpp/llama-quantize model-f16.gguf model-Q4_K_M.gguf Q4_K_M
Requires a llama.cpp build that supports the qwen35 architecture (GDN +
linear-attention). Keep MTP (do not pass --no-mtp — the model has MTP
tensors, block_count must be 64 + nextn).
📁 Files
| File | Size | Description |
|---|---|---|
model-00001-of-00012.safetensors … -00012 |
~54 GB | BF16 merged weights (12 shards) |
model.safetensors.index.json |
— | Shard index |
config.json |
— | Model config (MTP=1, text_config restored) |
tokenizer.json / tokenizer_config.json |
— | Qwen3.5 tokenizer |
preprocessor_config.json / video_preprocessor_config.json |
— | Processor (multimodal) |
chat_template.jinja |
— | Chat template (thinking + tool-call) |
generation_config.json |
— | Generation defaults |
⚠️ Notes & Limitations
- Preview release — validated on ARC/MMLU/GSM8K + tool-call smoke tests, not
yet on full agentic benchmarks (e.g. Terminal-Bench / SWE-Bench). - Abliterated base: safety-tuning refusals are reduced by design (λ=1.2 kept
capability while cutting refusals 98%→39%). Exercise judgment for harmful use. - Document tools (
csv_*,doc_*) are emitted as structured calls — you
must wire them to a runtime (e.g. a function-calling agent loop) to actually
execute them. - 27B BF16 needs a multi-GPU setup (or a quantized GGUF) for practical inference.
📜 License
Apache-2.0. Base model: hotdogs/Qwen3.8-27B-abliterated
(λ=1.2 abliteration); fine-tune method follows the Train-Studio (unsloth + HF Trainer) recipe.
Built with ❤️ on 7× RTX 3090 — SFT via unsloth + HF Trainer, merged with MTP
preserved, benchmarked A/B against base.