← back to catalog · registered 2026-08-22 13:56

caiovicentino1/Huihui-Qwopus3.5-27B-v3-abliterated-HLWQ-Q5

caiovicentino1 Qwen 24B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/caiovicentino1%2FHuihui-Qwopus3.5-27B-v3-abliterated-HLWQ-Q5"
Response includes
  • classification m1
  • files 10
  • hub_downloads_all_time 3,308
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
3K
90 last 30d - cooling
Likes
15
Model age
6mo ago
created 2026-04-06
Downloads over time
Now3.3K→from2.5K↑34%
2.5K2.8K3.1K3.4K2.5K on Apr 153.3K on Oct 11AprMayJunJulAugSepOct
Apr 15 → Oct 11 · 65 snapshots · spans 179 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh ko ja
Tags
transformers safetensors qwen3_5 image-text-to-text hlwq quantized compressed-tensors int4 marlin vllm qwen3.5 hybrid-attention

Related

Total size
14.1 GB
Files
10
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-04-18 23:25

Files by quantization

Auxiliary files 10 files 14.1 GB
model-00001.safetensors 6.40 GB 74596882 download
model-00002.safetensors 4.67 GB 8462a3fe download
model-00003.safetensors 3.02 GB 57b74321 download
tokenizer.json 19.1 MB 87a7830d download
model.safetensors.index.json 118 KB ce40f613 download
tokenizer_config.json 5.27 KB e6415a99 download
README.md 4.61 KB f855658a download
config.json 4.46 KB a3d6c353 download
.gitattributes 1.53 KB 52373fe2 download
hlwq_config.json 658 B ff151643 download

README current version from Hugging Face


license: apache-2.0
base_model: Jackrong/Qwopus3.5-27B-v3
language:

  • en
  • zh
  • ko
  • ja
    tags:
  • hlwq
  • quantized
  • compressed-tensors
  • int4
  • marlin
  • vllm
  • qwen3.5
  • hybrid-attention
  • gated-deltanet
  • abliterated
    library_name: transformers
    pipeline_tag: text-generation

⚡ Huihui-Qwopus3.5-27B-v3-abliterated — HLWQ CT INT4

CompressedTensors INT4 of Jackrong/Qwopus3.5-27B-v3 (abliterated) via HLWQ (Hadamard-Lloyd Weight Quantization)

Native vLLM. Marlin kernel. Zero plugin. 168 tok/s on A100.

📊 Compression

Compression

Metric Value
📦 Format CompressedTensors INT4 symmetric (gs=128)
💾 Model size 15.2 GB (3 shards)
📉 Compression 72% (54 → 15.2 GB)
⚡ Kernel Marlin (fused dequant+matmul)
🏗️ Architecture Qwen3.5 hybrid — 64 layers (48 GDN + 16 Full Attention)
🔢 Parameters 27B dense

🏎️ Quick Start — One Command

pip install vllm
vllm serve caiovicentino1/Huihui-Qwopus3.5-27B-v3-abliterated-HLWQ-Q5 \
  --language-model-only --enforce-eager

No plugin. No pip install polarquant. No custom code.

📊 Speed

Speed Benchmark

GPU tok/s VRAM
A100 80GB 168 15.2 GB
RTX PRO 6000 96GB 18 15.2 GB
RTX 4090 24GB ~10 15.2 GB

🎯 Quality

Quality PPL

Method PPL (WikiText-2) Delta
BF16 baseline 6.37 —
HLWQ → INT4 (ours) 6.56 +0.19
Direct INT4 (naive) 6.68 +0.31

HLWQ produces better INT4 weights than direct quantization — 0.12 PPL improvement from Hadamard rotation + Lloyd-Max preprocessing.

💻 GPU Compatibility

GPU Compatibility

GPU VRAM Status
RTX 4060 Ti 16 GB ⚠️ Tight (15.2 GB model + KV cache)
RTX 4090 24 GB ✅ Comfortable
A100 / H100 80 GB ✅ Full speed
RTX PRO 6000 96 GB ✅ Full speed

🧬 Architecture

Architecture

Spec Value
Parameters 27B (dense)
Layers 64 (48 GDN + 16 Full Attention)
Hidden dim 3584
Head dim 128
Context 131,072 tokens
Vision Multimodal ViT (skipped with --language-model-only)

🔬 How HLWQ Works

Standard INT4 quantizes weights directly — outliers cause high error.
HLWQ adds a preprocessing step before INT4:

BF16 weights
    │
    ▼
[1] Hadamard rotation → distributes energy uniformly
    │  (eliminates outliers, weights become Gaussian)
    │
    ▼
[2] Lloyd-Max Q5 → MSE-optimal 5-bit quantization
    │  (best possible codebook for Gaussian distribution)
    │
    ▼
[3] Dequant → BF16 → INT4 symmetric (gs=128)
    │  (cleaner weights = better INT4)
    │
    ▼
CompressedTensors (Marlin kernel) → vLLM serve

Same speed as GPTQ/AWQ, better quality.

🔧 Usage with Transformers

pip install polarquant
import polarengine_vllm  # auto-registers with transformers
from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained(
    "caiovicentino1/Huihui-Qwopus3.5-27B-v3-abliterated-HLWQ-Q5",
    device_map="auto", trust_remote_code=True
)
tokenizer = AutoTokenizer.from_pretrained(
    "caiovicentino1/Huihui-Qwopus3.5-27B-v3-abliterated-HLWQ-Q5",
    trust_remote_code=True
)

inputs = tokenizer("Hello!", return_tensors="pt").to("cuda")
out = model.generate(**inputs, max_new_tokens=100)
print(tokenizer.decode(out[0], skip_special_tokens=True))

⚠️ Important Flags

Flag Why
--language-model-only Qwen3.5 is multimodal — skip vision encoder (text weights only)
--enforce-eager Required on Blackwell GPUs (cc 12.0). Optional on Ampere/Hopper

📖 Citation

@misc{hlwq2026,
  title={HLWQ: Hadamard-Lloyd Weight Quantization for Large Language Models},
  author={Caio Vicentino},
  year={2026},
  url={https://arxiv.org/abs/2603.29078}
}

🔗 Links

Resource Link
📄 Paper arXiv:2603.29078
🔧 Code GitHub
📦 PyPI pip install polarquant
🏠 Base model Jackrong/Qwopus3.5-27B-v3
🔀 vLLM plugin polarengine-vllm

README history 4 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-04-18Update model card with charts, benchmarks, and architecture detail923258f4.6 KB
    Loading...
  2. 2026-04-13HLWQ rebrand: title, tags, notice, self-links62bf0c63.8 KB
    Loading...
  3. 2026-04-10docs: add HLWQ rebrand notice (cite Han et al. PolarQuant prior art)1a4283d3.9 KB
    Loading...
  4. 2026-04-09chore: squash history to reclaim orphaned LFS objects (HEAD unchanged)e1b49c33 KB
    Loading...

Discussions 1 thread

  1. 2026-04-18How can I enable vision abilityopen2 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration