← back to catalog · registered 2026-08-22 13:56

ababaka/Huihui-Qwen3.8-27B-Abliterated-W4A16-AutoRound

ababaka Qwen 28B second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/ababaka%2FHuihui-Qwen3.8-27B-Abliterated-W4A16-AutoRound"
Response includes
  • classification m1
  • files 20
  • hub_downloads_all_time 2,363
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
2K
498 last 30d - stable
Likes
3
Model age
7w ago
created 2026-08-20
Downloads over time
Now2.5K→from216↑1,072%
1009881.9K2.8K216 on Aug 192.5K on Oct 11AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en ru da
Tags
safetensors qwen3_5 qwen qwen3.8 mllm compressed-tensors w4a16 auto-round 4-bit int8 abliterated speculative-decoding

Related

Total size
15.6 GB
Files
20
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-25 09:18

Files by quantization

Auxiliary files 20 files 15.6 GB
model-00004-of-00007.safetensors 3.00 GB 62a42262 download
model-00001-of-00007.safetensors 2.99 GB d23bf46d download
model-00002-of-00007.safetensors 2.98 GB da4e28d1 download
model-00003-of-00007.safetensors 2.98 GB 4e458f33 download
model-00006-of-00007.safetensors 1.20 GB b2e0854e download
model-00007-of-00007.safetensors 1.20 GB ce55d80e download
model-00005-of-00007.safetensors 667 MB 0994bb8d download
model_extra_tensors.safetensors 615 MB e565c306 download
mtp_draft_vocab_ids.pt 322 KB 8af90286 download
tokenizer.json 19.1 MB 06b95093 download
model.safetensors.index.json 195 KB 1f1321de download
config.json 21.7 KB 3ac99868 download
quantization_config.json 15.8 KB 263f7f84 download
chat_template.jinja 8.74 KB c0c686f9 download
README.md 6.43 KB 1a3d56e2 download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.16 KB 33818c7f download
tokenizer_config.json 1.14 KB 1d134cd2 download
preprocessor_config.json 390 B 2ea84a43 download
generation_config.json 214 B 8b9f95da download

README current version from Hugging Face


license: apache-2.0
language:

  • en
  • ru
  • da
    tags:
  • qwen
  • qwen3.8
  • mllm
  • compressed-tensors
  • w4a16
  • auto-round
  • 4-bit
  • int8
  • abliterated
  • speculative-decoding
  • mtp
    base_model: huihui-ai/Huihui-Qwen3.8-27B-abliterated
    inference: false

Huihui-Qwen3.8-27B-Abliterated-W4A16-AutoRound

Quantized from scratch (FP16 → W4A16) version of
huihui-ai/Huihui-Qwen3.8-27B-abliterated
in compressed-tensors pack-quantized format, following the recipe of
dbirk/Qwen3.8-27B-W4A16-AutoRound.

Fully loadable by vLLM 0.27+ out of the box (quant_method: compressed-tensors),
with the MTP (multi-token-prediction) speculative-decoding pipeline applied:
lm_head / embed_tokens / MTP module quantized to int8 and a 40k-token
draft-vocabulary draft head for fast speculation.

Средняя скорость генерации на RTX 3090 (24 GB): ≈80 tok/s декод с MTP k=3
(44 сэмпла live-мониторинга: средняя 74, во время активной генерации ~81,
медиана 84, пики до 121 tok/s при acceptance MTP ~65–90%; на длинных
контекстах acceptance падает до 5–30% → ~40–55 tok/s; префилл 1.8–7.6k tok/s).
Без MTP — примерно втрое медленнее.

Как сделана модель (рецепт, пайплайн, формат)

Полный цикл конвертации FP16 → vLLM-ready на RTX 3090 занимает ≈2 ч 45 мин
(квантование 64 слоёв ≈2 ч 37 мин при ~137–155 с/слой + пост-пайплайн ≈6 мин;
пики: VRAM 15.5 GB / RAM 23 GB).

Квантование (auto-round 0.14.2)

Параметр Значение
Toolchain auto-round 0.14.2, transformers 5.15.0, torch 2.13+cu130, compressed-tensors
Scheme W4A16, int4, group 128, symmetric (pack-quantized)
Dataset NeelNanda/pile-10k, 128 samples, seqlen 2048, batch 4
Iters 200
seed / trust_remote_code 42 / True
quant_nontext_module False (vision tower не трогается)
Export format llm_compressor

Остались в BF16 через layer_config: linear_attn.in_proj_a/b (все 48 слоёв),
вся vision-башня visual.*, модуль mtp и lm_head.

Пост-квантование (пайплайн, в том же каталоге)

Шаг Что
quant_lm_head lm_head → int8 g128 (−1.3 GB VRAM)
quant_embed embed_tokens → int8 g128 (untied, ещё −1.3 GB)
quant_mtp mtp.fc + 7 Linear MTP → int8 g128
build_draft_vocab draft head на 40960 токенов для MTP-спекуляции (частотный топ)

config_groups (порядок важен для матчинга compressed-tensors):

group_0  targets=["Linear"]             int4 g128 sym   — основной корпус
group_1  targets=["re:.*lm_head$"]      int8 g128       — lm_head + mtp.draft_lm_head
group_2  targets=["re:.*embed_tokens$"] int8 g128       — embed_tokens
group_3  targets=["re:^mtp\\..*"]       int8 g128       — модуль MTP

Веса: weight_packed int32 + weight_scale fp16 (bf16 для embedding) + weight_shape.
MTP и draft head лежат в model_extra_tensors.safetensors, рядом
mtp_draft_vocab_ids.pt (карта id для урезанного словаря драфтера).

Карточка базы (кратко)

  • Архитектура: Qwen3.5 (qwen3_5), 27B, 64 слоя (48 linear-attention DeltaNet +
    16 full attention), словарь 248k, модуль MTP, vision-башня.
  • Abliterated от huihui-ai: refusal-поведение удалено, способности сохранены;
    ответственность за использование — на пользователе.
  • Untied embeddings: lm_head и embed_tokens — отдельные матрицы.
Как запускать (vLLM, 150k контекста на 3090)

Требования

  • vLLM 0.27+ (формат compressed-tensors грузится из коробки, квантование указывать не нужно)
  • torch 2.13+cu130 (CUDA 13-совместимый драйвер)
  • RTX 3090 24 GB: полный 150k-контекст влезает (KV-пул ~178k токенов, fp8)

Запуск

python -m vllm.entrypoints.openai.api_server \
  --model ababaka/Huihui-Qwen3.8-27B-Abliterated-W4A16-AutoRound \
  --served-model-name qwen3.8-27b \
  --host 0.0.0.0 --port 18020 \
  --gpu-memory-utilization 0.93 \
  --max-model-len 150000 \
  --max-num-seqs 8 \
  --kv-cache-dtype fp8 \
  --mamba-ssm-cache-dtype float16 \
  --max-num-batched-tokens 2048 \
  --language-model-only \
  --speculative-config '{"method":"mtp","num_speculative_tokens":3,"draft_sample_method":"probabilistic"}'

Ключевые флаги:

Флаг Зачем
--mamba-ssm-cache-dtype float16 обязателен: конфиг модели просит float32, дефолт не влезает в 24 GB
--kv-cache-dtype fp8 KV-кэш fp8 (вместе с attention-блоком 832 токена)
--max-num-seqs 8 >8 не влезает: Mamba-блоки по 1 на последовательность
--speculative-config MTP-спекуляция k=3 с урезанным 40k словарём драфтера

Проверка

curl http://localhost:18020/v1/chat/completions \
  -H "Authorization: Bearer <ключ>" -H "Content-Type: application/json" \
  -d '{"model":"qwen3.8-27b","messages":[{"role":"user","content":"Hi!"}],"max_tokens":64}'

Ожидаемый старт: загрузка весов ~0.2 с, torch.compile ~26 с (первый раз),
KV-пул 151–178k токенов, декод ~55–110 tok/s с MTP (accept 45–90% на
коротких промптах, 15–30% на длинных).

Кредиты

README history 5 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-25Upload folder using huggingface_hubc20530b6.6 KB
    Loading...
  2. 2026-08-20Upload README.md with huggingface_hub92600106.4 KB
    Loading...
  3. 2026-08-20Upload README.md with huggingface_huba56e14b6.4 KB
    Loading...
  4. 2026-08-20Upload README.md with huggingface_hub135b3386.3 KB
    Loading...
  5. 2026-08-20Upload README.md with huggingface_hub287a2916 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration