← back to catalog · registered 2026-08-23 12:02

hotdogs/Qwen3.8-27B-abliterated-code-analysis-preview-mtp-GGUF

hotdogs Qwen 27B GGUF second-order 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/hotdogs%2FQwen3.8-27B-abliterated-code-analysis-preview-mtp-GGUF"
Response includes
  • classification m8
  • files 5
  • hub_downloads_all_time 1,005
  • author_summary 25 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
1K
276 last 30d - stable
Likes
1
Model age
7w ago
created 2026-08-23
Downloads over time
Now1.1K→from330↑235%
2915898871.2K330 on Aug 261.1K on Oct 11AugSepOct
Aug 26 → Oct 11 · 47 snapshots · spans 46 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 2 formats · 276 downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
mit
Quantizations
F16 Q4_K Q6_K
Tags
transformers gguf qwen3 code-review code-analysis lora sft abliterated multi-token-prediction mtp llama.cpp text-generation

Related

Total size
87.5 GB
Files
5
Quantizations
4
Registered
2026-08-23 12:02
Last updated on HF
2026-08-23 12:05

Files by quantization

F16 1 file 50.9 GB
Qwen3.8-27B-code-analysis-preview-v2-mtp-f16.gguf 50.9 GB 8d6cd92a download
Q6_K 1 file 20.9 GB
Qwen3.8-27B-code-analysis-preview-v2-mtp-Q6_K.gguf 20.9 GB b2c2bc10 download
Q4_K 1 file 15.7 GB
Qwen3.8-27B-code-analysis-preview-v2-mtp-Q4_K_M.gguf 15.7 GB 98c3902b download
Auxiliary files 2 files 4.48 KB
README.md 2.74 KB 718bcfff download
.gitattributes 1.74 KB 6957a042 download

README current version from Hugging Face


base_model:

  • hotdogs/Qwen3.8-27B-abliterated-code-analysis-preview
    library_name: transformers
    model_type: qwen3_5
    pipeline_tag: text-generation
    tags:
  • gguf
  • qwen3
  • code-review
  • code-analysis
  • lora
  • sft
  • abliterated
  • multi-token-prediction
  • mtp
  • llama.cpp
    license: mit

Qwen3.8-27B Code Analysis Preview (v2) — MTP GGUF

GGUF quantizations of the code-analysis fine-tune hotdogs/Qwen3.8-27B-abliterated-code-analysis-preview, with Multi-Token-Prediction (MTP) tensors preserved for speculative decoding.

Given a code snippet, it returns a structured, multi-paragraph review — real bugs, line-level reasoning, severity, and a concrete fix in a code block. It is a reasoning model: it thinks first, then answers.

v2 fixes the template-collapse of v1. v1 was trained on a synthetic placeholder dataset and answered in one line ("No bugs found. Code is clean."). v2 is retrained on 21,009 real code+bug+answer rows across 5 languages (Python, JS, Go, Rust, C) with detailed 550–880 char answers — the model now actually finds the bugs and generalizes to unseen bug types.

Files

File Size Quant Bits/weight
Qwen3.8-27B-code-analysis-preview-v2-mtp-f16.gguf 51 GB F16 16.0
Qwen3.8-27B-code-analysis-preview-v2-mtp-Q6_K.gguf 21 GB Q6_K 6.56
Qwen3.8-27B-code-analysis-preview-v2-mtp-Q4_K_M.gguf 16 GB Q4_K_M 4.92

All 3 files: 866 tensors, MTP preserved — 15 blk.64.* tensors (11 transformer-layer + 4 blk.64.nextn.*).

MTP / speculative decoding

The MTP head lives in block 64 (blk.64.nextn.*). With llama.cpp you can use it as a draft model for speculative decoding:

llama-server -m Qwen3.8-27B-code-analysis-preview-v2-mtp-Q6_K.gguf \
  --n-gpu-layers 999 --ctx-size 262144 --parallel 1 \
  --cache-type-k f16 --cache-type-v f16 --flash-attn on \
  --temp 1 --top-k 20 --top-p 0.95 --min-p 0.0 --jinja --tools all

(The blk.64.nextn.* tensors load automatically; enable speculative decoding via the predictor/draft options of your llama.cpp build.)

Smoke test (v2)

Case Result
Off-by-one (in-archetype) 🟢 Found it + fix + docstring note
Async race (unseen) 🟢 "no cache-hit fast path" + concurrency
Clean code (hallucination test) 🟢 "correct, no bugs" + minor float/bool note

Measured on Q6_K (5×3090 / 2 GPUs, flash-attn): ~34 t/s generation, ~141 t/s prompt eval.

Recommended

  • Q6_K — best quality/size balance (21 GB) — recommended
  • Q4_K_M — fastest / smallest (16 GB), great for consumer GPUs
  • F16 — max fidelity (51 GB)

License

MIT (inherits the abliterated base).

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-23Upload README.md with huggingface_hubcabdcf42.7 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration