← back to catalog · registered 2026-08-22 13:56

hotdogs/qwen27B-Agent-R2-abliterated-preview

hotdogs Qwen 27B GGUF second-order 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/hotdogs%2Fqwen27B-Agent-R2-abliterated-preview"
Response includes
  • classification m8
  • files 18
  • hub_downloads_all_time 19,436
  • author_summary 25 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
19K
1K last 30d - cooling
Likes
6
Model age
3mo ago
created 2026-07-13

Training datasets

Corpora the author lists in the model card. Datasets tracked in our /datasets catalog carry a category badge linking to the workflow stage. Others open on Hugging Face.

Downloads over time
Now19.8K→from2.9K↑571%
2.1K8.6K15K21.5K2.9K on Jul 1519.8K on Oct 11JulAugSepOct
Jul 15 → Oct 11 · 53 snapshots · spans 88 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
agpl-3.0
Languages
en th
Tags
transformers safetensors gguf qwen3_5_text text-generation qwen fable agent tool-call tool-use function-calling reasoning

Related

Total size
50.9 GB
Files
18
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-15 23:05

Files by quantization

Auxiliary files 18 files 50.9 GB
model-00001-of-00002.safetensors 46.4 GB 8e9ade90 download
model-00002-of-00002.safetensors 4.49 GB 6264a5a0 download
tokenizer.json 19.1 MB 06b95093 download
imatrix.dat 13.0 MB 63237b3c download
mixed_calibration.txt 10.6 MB 33f401e7 download
wiki.train.raw 10.4 MB 6707892f download
wiki.test.raw 1.23 MB d9d79158 download
wiki.valid.raw 1.09 MB 7760d2e6 download
code_samples.txt 195 KB da46a9a1 download
model.safetensors.index.json 83.1 KB 613e2894 download
chat_template.jinja 11.7 KB 177eacb1 download
README.md 8.53 KB 367628a9 download
.gitattributes 3.05 KB 40bf12d0 download
config.json 2.68 KB 02da8eac download
quant_rules.txt 1.15 KB 5a43402c download
tokenizer_config.json 1.10 KB ed1f99f3 download
generation_config.json 214 B 3f25ead4 download
thai_articles.txt 28.0 B 5643a78c download

README current version from Hugging Face


license: agpl-3.0
language:

  • en
  • th
    tags:
  • qwen
  • fable
  • agent
  • tool-call
  • tool-use
  • function-calling
  • reasoning
  • abliterated
  • uncensored
  • conversational
  • mtp
  • multi-token-prediction
  • transformers
  • text-generation
  • thai
  • speculative-decoding
  • preview
    base_model:
  • hotdogs/qwen27b-abliterated-Fable-MTP
  • huihui-ai/Huihui-Qwen3.6-27B-abliterated
    datasets:
  • hotdogs/uka-fable-reasoning
  • NousResearch/hermes-function-calling-v1
  • 11-47/claude_opus_4.8_max_thinking_5k_v2
    library_name: transformers
    pipeline_tag: text-generation

🐉 qwen27B-Agent-R2-abliterated-preview

27B Agent Model — Abliterated · MTP · Tool-Calling · Speculative Decoding


Architecture graph for hotdogs/qwen27B-Agent-R2-abliterated-preview. Open in hfviewer

Preview release — Built from Fable-MTP + agent LoRA fusion. Features Multi-Token Prediction (MTP) for speculative decoding (up to 2× faster generation), abliterated (no guardrails), and tool-calling support.


✨ Key Features

Capability Description
⚡ MTP Speculative Decoding Draft 2 tokens at a time — up to +85% decode TPS on single GPU
🔧 Tool Calling Hermes/Qwen function-calling format via llama.cpp --tools all
🔓 Abliterated Unrestricted — all refusal mechanisms removed
🧠 Reasoning Fable-style reasoning with step-by-step CoT
🌏 Thai + English Native bilingual support
💻 Code Python, shell, system tasks

🚀 Usage

llama.cpp (Recommended)

# Quick test
./llama-cli -m qwen27B-Agent-R2-abliterated-preview-MTP.IQ4_NL.gguf \
  -p "Hello" -n 100 --temp 0.6

# Full agent server with tool calling + MTP speculative decoding
./llama-server \
  -m qwen27B-Agent-R2-abliterated-preview-MTP.IQ4_NL.gguf \
  --mmproj Qwen3.6-27B-mmproj-BF16.gguf \
  --host 0.0.0.0 --port 8080 \
  --n-gpu-layers 999 \
  --ctx-size $((256*1024)) \
  --batch-size 8192 \
  --ubatch-size 1024 \
  --cache-type-k f16 \
  --cache-type-v f16 \
  --flash-attn on \
  --cont-batching \
  --mlock \
  --no-mmap \
  --chat-template-file chat_template.jinja \
  --samplers "top_k;top_p;min_p;temperature;dry" \
  --repeat-penalty 1.10 \
  --repeat-last-n 256 \
  --dry-multiplier 0.8 \
  --dry-base 1.75 \
  --dry-allowed-length 2 \
  --dry-penalty-last-n -1 \
  --reverse-prompt "<|im_end|>" \
  --reverse-prompt "<|endoftext|>" \
  -n 4096 \
  --tools all \
  --parallel 1
  --spec-type draft-mtp \
  --spec-draft-n-max 2

llama.cpp (coding setting)

--samplers "top_k;top_p;min_p;temperature;dry" \
  --repeat-penalty 1.03 \
  --repeat-last-n 256 \
  --dry-multiplier 0.5 \
  --dry-base 1.75 \
  --dry-allowed-length 5 \
  --dry-penalty-last-n -1 \
Parameter Purpose
--cache-type-k f16 / --cache-type-v f16 F16 KV cache for quality
--flash-attn on Flash attention for speed
--tools all Enable tool/function calling
--spec-type draft-mtp MTP speculative decoding (draft 2 tokens)
--spec-draft-n-max 2 Max 2 draft tokens per step
--cont-batching Continuous batching for multi-turn
--chat-template-file chat_template.jinja Use Jinja2 chat template from GGUF

Python (Transformers)

from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained(
    "hotdogs/qwen27B-Agent-R2-abliterated-preview",
    torch_dtype="auto",
    device_map="auto",
    trust_remote_code=True
)
tokenizer = AutoTokenizer.from_pretrained("hotdogs/qwen27B-Agent-R2-abliterated-preview")

messages = [{"role": "user", "content": "Hello"}]
inputs = tokenizer.apply_chat_template(messages, tokenize=True, return_tensors="pt")
outputs = model.generate(inputs, max_new_tokens=256, temperature=0.6)
print(tokenizer.decode(outputs[0]))

📦 Downloads

File Size Quant Description
qwen27B-Agent-R2-abliterated-preview-MTP.IQ4_NL.gguf 16 GB IQ4_NL Recommended — balanced quality/speed + imatrix
qwen27B-Agent-R2-abliterated-preview.Q6_K.gguf 21 GB Q6_K Higher quality, slightly slower
qwen27B-Agent-R2-abliterated-preview.Q6_K_imatrix.gguf 22 GB Q6_K Higher quality, slightly slower + imatrix
qwen27B-Agent-R2-abliterated-preview.f16.gguf 51 GB f16 Full precision

🎯 Q4_K_M is recommended for most users — good quality with 16 GB VRAM usage.

📷 Multimodal Projector (mmproj)

For vision support, pair this model with the mmproj from Qwen/Qwen3.6-27B:

# Extract mmproj from Qwen3.6-27B vision model
python3 ./llama.cpp/convert_hf_to_gguf.py \
  --mmproj Qwen/Qwen3.6-27B \
  --outfile mmproj-qwen3.6-27b.gguf

# Use with llama-server for vision + tool calling
./llama-server \
  -m qwen27B-Agent-R2-abliterated-preview-MTP.IQ4_NL.gguf \
  --mmproj mmproj-qwen3.6-27b.gguf \
  ... (same params as above)

Note: The mmproj extracts the vision projector from the base Qwen3.6-27B vision encoder. The language model (this GGUF) then interprets visual embeddings for image understanding tasks.


🧬 Architecture

Parameter Value
Base Qwen3.6-27B (Dense)
Parameters ~27B
Hidden Size 5,120
Attention Linear + Standard hybrid
Context 8,192 tokens (extendable)
Precision BF16 / GGUF quantized
Format ChatML (Jinja2 template)
MTP Head ✅ 1 extra layer (draft 2 tokens)

Built on hotdogs/qwen27b-abliterated-Fable-MTP with multi-LoRA fusion and MTP tensor injection from huihui-ai/Huihui-Qwen3.6-27B-abliterated.


✅ What This Model Excels At

  • Agent tasks — Tool calling, planning, multi-step reasoning
  • Coding — Python, shell scripts, system administration
  • Knowledge QA — General knowledge with step-by-step reasoning
  • Thai + English — Native-level bilingual capability
  • Creative — Storytelling, analysis, brainstorming

⚡ MTP Speculative Decoding

This model preserves the Multi-Token Prediction (MTP) head from the Qwen3.6 architecture, enabling speculative decoding:

Standard:  [token₁] → [token₂] → [token₃] → ...  (~36 TPS)
MTP:       [token₁ token₂] → [token₃ token₄] → ...  (~66 TPS)
  • MTP head adds ~849 MB to model size
  • Uses --spec-type draft-mtp in llama.cpp
  • Best for single-user agent workloads
  • ~1.2–1.8× decode speedup

💖 Support / โปรดสนับสนุน

If you find this model useful, please consider supporting my work!
หากคุณคิดว่าโมเดลนี้มีประโยชน์ กรุณาสนับสนุนผลงานของฉันด้วยนะคะ! 🙏

Bitcoin QR — Donate

₿ Bitcoin — BTC:

bc1qf27cyk3vmugcdyv9xdtuv5jwz37863crpj5c9v

Thank you for your support! 🙏✨
ขอบคุณมากๆ สำหรับการสนับสนุนค่า! 💖🤗


🙏 Acknowledgements / ขอบคุณ


Built with ❤️ by UKA — 18-year-old coder & cybersecurity expert

README history 11 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-15Update README.md3f40a108.5 KB
    Loading...
  2. 2026-08-13Update README.md9fca25e8 KB
    Loading...
  3. 2026-08-13Update README.md458b92b7.7 KB
    Loading...
  4. 2026-08-13Update README.mdb724a8f7.8 KB
    Loading...
  5. 2026-08-11Update README.md789493c7.5 KB
    Loading...
  6. 2026-07-22Update README.mde27a0bb7.4 KB
    Loading...
  7. 2026-07-18Update README.md75988727.4 KB
    Loading...
  8. 2026-07-18Update README.mdd60e6ae7.4 KB
    Loading...
  9. 2026-07-13Update README.mdf30ea8d7.3 KB
    Loading...
  10. 2026-07-13Upload README.md with huggingface_hub04010f87.3 KB
    Loading...
  11. 2026-07-13Upload README.md with huggingface_hubdf9e4086.7 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration