← back to catalog · registered 2026-08-22 13:56

letechlead/Huihui-Qwen3.8-27B-Abliterated-INT4-W4A16-AutoRound

letechlead Qwen 3.1B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/letechlead%2FHuihui-Qwen3.8-27B-Abliterated-INT4-W4A16-AutoRound"
Response includes
  • classification m1
  • files 19
  • hub_downloads_all_time 4,890
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
5K
2K last 30d - stable
Likes
3
Model age
7w ago
created 2026-08-18
Downloads over time
Now5.3K→from363↑1,354%
1172K3.9K5.8K363 on Aug 195.3K on Oct 11AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
transformers safetensors qwen3_5 image-text-to-text qwen3.8 qwen3.5 autoround int4 w4a16 abliterated multimodal conversational

Related

Total size
17.7 GB
Files
19
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-20 06:49

Files by quantization

Auxiliary files 19 files 17.7 GB
model-00003-of-00007.safetensors 3.00 GB 27de5df1 download
model-00004-of-00007.safetensors 3.00 GB bfd4d5d7 download
model-00001-of-00007.safetensors 3.00 GB f1bfa22a download
model-00002-of-00007.safetensors 2.97 GB af16ad69 download
model-00006-of-00007.safetensors 2.37 GB 55a14ee7 download
model-00007-of-00007.safetensors 2.37 GB 6866cf8a download
model-00005-of-00007.safetensors 734 MB f0cd6396 download
model_extra_tensors.safetensors 284 MB 94102b67 download
tokenizer.json 19.1 MB 06b95093 download
model.safetensors.index.json 188 KB 41cb3359 download
config.json 15.2 KB 418ebd0b download
quantization_config.json 10.7 KB 011526aa download
chat_template.jinja 8.74 KB c0c686f9 download
README.md 5.19 KB de368c70 download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.16 KB 33818c7f download
tokenizer_config.json 1.14 KB 1d134cd2 download
preprocessor_config.json 443 B 8ed39680 download
generation_config.json 214 B 8b9f95da download

README current version from Hugging Face


library_name: transformers
license: apache-2.0
pipeline_tag: image-text-to-text
base_model:

  • huihui-ai/Huihui-Qwen3.8-27B-abliterated
  • Qwen/Qwen3.8-27B
    tags:
  • qwen3.8
  • qwen3.5
  • autoround
  • int4
  • w4a16
  • abliterated
  • multimodal
  • safetensors

Huihui-Qwen3.8-27B-Abliterated-INT4-W4A16-AutoRound

This is the SafeTensors INT4 AutoRound version of huihui-ai/Huihui-Qwen3.8-27B-abliterated, which is based on Qwen/Qwen3.8-27B.

The artifact is multimodal: it contains the Qwen3.5 language model, MTP path, and vision tower. AutoRound INT4 quantization is applied to the language-model transformer weights and the exported MTP path as described below. The vision tower remains in BF16.

Quantization

The export was produced with AutoRound 0.14.2. The settings below are taken directly from quantization_config.json and config.json:

  • Weight precision: INT4
  • Format: W4A16 (4-bit weights with floating-point activations)
  • Group size: 128
  • Symmetric quantization: enabled
  • Packing format: auto_round:auto_gptq
  • Calibration sequence length: 512 tokens
  • Batch size: 1
  • Quantization method: auto-round
  • Quantization targets: model.language_model.layers and mtp.layers
  • Selected linear-attention input projections are retained at 16-bit through extra_config
  • mtp.fc is retained at 16-bit through extra_config

The exported configuration reports BF16 as the model dtype. The repository is approximately 19.02 GB decimal (17.71 GiB), including configuration, tokenizer, processor files, and SafeTensors weights.

Model details

  • Model name: Huihui-Qwen3.8-27B-Abliterated-INT4-W4A16-AutoRound
  • Architecture: Qwen3_5ForConditionalGeneration
  • Text layers: 64
  • Hidden size: 5120
  • Vocabulary size: 248,320
  • Maximum position embeddings: 262,144
  • Attention layout: 48 linear-attention layers and 16 full-attention layers
  • Parameters represented by the indexed model weights: 6,284,446,960
  • Vision tower: BF16, with Qwen3_5ForConditionalGeneration multimodal processor configuration

The source model retains its first 15 layers without ablation and does not modify its MTP or visual components. That is source-model information; it is not an additional transformation performed by this AutoRound export.

Usage with vLLM

The following command launches this exact model repository with vLLM and exposes both the stable local alias and the full model name. The --served-model-name values are the names accepted by the OpenAI-compatible API; they do not rename the files on disk.

vllm serve letechlead/Huihui-Qwen3.8-27B-Abliterated-INT4-W4A16-AutoRound \
  --served-model-name local Huihui-Qwen3.8-27B-Abliterated-INT4-W4A16-AutoRound \
  --tensor-parallel-size 2 \
  --max-model-len 262144 \
  --port 18080

This command is for the exact artifact letechlead/Huihui-Qwen3.8-27B-Abliterated-INT4-W4A16-AutoRound; do not replace the model argument with a generic path such as /models. The tensor-parallel size, maximum context that is practical, and KV-cache dtype must be chosen for the available hardware. The model supports text input and multimodal processor inputs when the installed vLLM release supports this architecture and its AutoRound/AWQ-compatible INT4 format.

Usage with Transformers

Use a recent Transformers release that supports Qwen3_5ForConditionalGeneration and the model's processor:

from transformers import AutoModelForMultimodalLM, AutoProcessor

model_id = "letechlead/Huihui-Qwen3.8-27B-Abliterated-INT4-W4A16-AutoRound"

processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForMultimodalLM.from_pretrained(
    model_id,
    device_map="auto",
    dtype="auto",
)

For text-only requests, pass text through the processor. For image or video requests, pass the media together with the text prompt according to the Qwen3.5 Transformers documentation.

Files

  • model-00001-of-00007.safetensors through model-00007-of-00007.safetensors: sharded model weights
  • model_extra_tensors.safetensors: additional exported tensors, including the exported MTP path
  • model.safetensors.index.json: weight-to-shard index and parameter metadata
  • quantization_config.json: AutoRound quantization metadata
  • config.json: model architecture, text configuration, vision configuration, and context configuration
  • tokenizer.json, tokenizer_config.json, chat_template.jinja: tokenizer and chat template
  • processor_config.json, preprocessor_config.json: multimodal processor configuration

Limitations and responsible use

This model inherits the source model's abliterated behavior and may produce sensitive, controversial, or otherwise unsafe output. It is intended for research, evaluation, and controlled use. Users are responsible for complying with applicable laws, platform rules, and organizational policies, and for reviewing outputs before relying on them.

License and attribution

This derived model is published under Apache-2.0, matching the license declared by the source model card. Review the source model card and repository for the original model's terms and attribution requirements.

Quantization artifact published by LeTechLead.

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-20Audit README and use exact model name in vLLM example2e8f0245.2 KB
    Loading...
  2. 2026-08-18Fix Hugging Face namespace casing in READMEebbd2cd4.7 KB
    Loading...
  3. 2026-08-18Add README.md7386b5a4.7 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration