← back to catalog · registered 2026-08-22 13:56

lyf/Qwen3.8-27B-Huihui-Abliterated-NVFP4-MTP-VL

lyf Qwen 24B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/lyf%2FQwen3.8-27B-Huihui-Abliterated-NVFP4-MTP-VL"
Response includes
  • classification m1
  • files 24
  • hub_downloads_all_time 7,770
  • author_summary 10 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
8K
3K last 30d - stable
Likes
3
Model age
7w ago
created 2026-08-19
Downloads over time
Now8.3K→from61↑13,572%
03.1K6.1K9.2K61 on Aug 198.3K on Oct 11AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh
Tags
transformers safetensors qwen3_5 image-text-to-text qwen3.8 nvfp4 compressed-tensors vllm blackwell rtx-5090 sm120 multimodal

Related

Total size
19.1 GB
Files
24
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-19 03:53

Files by quantization

Auxiliary files 24 files 19.2 GB
model-00001-of-00002.safetensors 9.27 GB 3e48d5c7 download
model-00002-of-00002.safetensors 9.08 GB a61b78bb download
model-mtp-extra.safetensors 810 MB b44a958e download
tokenizer.json 19.1 MB 06b95093 download
vocab.json 6.41 MB 0aa0ce06 download
merges.txt 3.20 MB a494e019 download
model.safetensors.index.json 272 KB d4096555 download
config.json 29.1 KB a732e096 download
LICENSE 11.3 KB f938136e download
chat_template.jinja 8.74 KB c0c686f9 download
README.md 3.32 KB c20a075a download
hf_quant_config.json 3.28 KB a306db26 download
SHA256SUMS.source 2.60 KB 1fbd15cc download
SHA256SUMS 1.93 KB d11b6517 download
.gitattributes 1.53 KB 52373fe2 download
VALIDATION_REPORT.json 1.16 KB 077677c1 download
tokenizer_config.json 1.10 KB d1a20cc3 download
BUILD_MANIFEST.json 645 B 4606dd92 download
recipe.yaml 495 B e2c4a1c4 download
STATIC_VALIDATION_REPORT.json 465 B 4c5f98e6 download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
crc32.txt 238 B 6de5ee6a download
generation_config.json 218 B edcb0e7f download

README current version from Hugging Face


license: apache-2.0
license_link: https://huggingface.co/Qwen/Qwen3.8-27B/blob/main/LICENSE
base_model: huihui-ai/Huihui-Qwen3.8-27B-abliterated
base_model_relation: quantized
pipeline_tag: image-text-to-text
library_name: transformers
language:

  • en
  • zh
    tags:
  • qwen3_5
  • qwen3.8
  • nvfp4
  • compressed-tensors
  • vllm
  • blackwell
  • rtx-5090
  • sm120
  • multimodal
  • vision
  • video
  • mtp
  • speculative-decoding
  • tool-calling
  • reasoning
  • abliterated

Huihui Qwen3.8-27B Abliterated NVFP4 MTP VL

A compressed-tensors NVFP4 W4A4 release of huihui-ai/Huihui-Qwen3.8-27B-abliterated at revision d42ca8978c5a66e92c3446d46e8adfe03ef692ff. The same-source BF16 vision/video tower (333 tensors) and all 15 BF16 MTP tensors are retained.

Quick start — RTX 5090 / Blackwell

hf download lyf/Qwen3.8-27B-Huihui-Abliterated-NVFP4-MTP-VL --local-dir ./huihui-qwen38-nvfp4

docker run --rm --gpus all --ipc=host --network=host   -e VLLM_NVFP4_GEMM_BACKEND=flashinfer-cutlass   -e VLLM_USE_FLASHINFER_SAMPLER=1   -v "$PWD/huihui-qwen38-nvfp4:/model:ro"   vllm/vllm-openai:qwen38-x86_64-cu130   /model   --served-model-name qwen38-huihui-nvfp4   --host 0.0.0.0 --port 8000   --max-model-len 4096   --kv-cache-dtype fp8   --gpu-memory-utilization 0.92   --max-num-seqs 1   --max-num-batched-tokens 1024   --enable-prefix-caching   --trust-remote-code   --speculative-config '{"method":"qwen3_5_mtp","num_speculative_tokens":3}'

The command above is the verified multimodal/MTP smoke profile. Larger context values require workload-specific VRAM validation.

Quantization and lineage

Component Format/source
Language-model Linear layers NVFP4 W4A4, group size 16
Vision/video tower 333 tensors, BF16, same Huihui checkpoint
MTP head 15 tensors, BF16, same Huihui checkpoint
lm_head, token embedding, GDN conv1d BF16
Calibration CNN/DailyMail 3.0.0, 20 × 8192 tokens
Packaging compressed-tensors nvfp4-pack-quantized

No official Qwen, Unsloth, Blackfrost, or other behavioral variant weights were grafted into this release.

Validation

Validated on one RTX 5090, 450 W cap, using vllm/vllm-openai:qwen38-x86_64-cu130:

  • GET /health and /v1/models: passed
  • OpenAI-compatible 1024-token text generation: passed
  • Native MTP n=3: passed
  • MTP draft tokens: 1362; accepted: 572; acceptance: 42.0%
  • Mean acceptance length: 2.33
  • Per-position acceptance: 0.602 / 0.429 / 0.295
  • Client elapsed for 1024 output tokens: 10.73 s
  • Runtime VRAM under generation: ~28,944 MiB

The earlier full release workflow also verifies 333 vision tensors and 15 same-source MTP tensors statically. The NInfer derivative in the companion repository carries a separate real-image runtime validation.

Files

  • model-00001-of-00002.safetensors, model-00002-of-00002.safetensors: compressed checkpoint
  • model-mtp-extra.safetensors: 15 same-source BF16 MTP tensors
  • model.safetensors.index.json: complete 2687-tensor index
  • BUILD_MANIFEST.json, VALIDATION_REPORT.json, STATIC_VALIDATION_REPORT.json, recipe.yaml, SHA256SUMS: provenance and reproducibility

Intended use

This is an abliterated behavioral derivative intended for model research and local inference. Users are responsible for downstream use and applicable policies.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-19Add files using upload-large-folder tool3e2cce23.3 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration