← back to catalog · registered 2026-08-22 13:56

gsting/Qwen3.6-35B-A3B-abliterated-FP8

gsting Qwen 34B MoE multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/gsting%2FQwen3.6-35B-A3B-abliterated-FP8"
Response includes
  • classification m1
  • files 39
  • hub_downloads_all_time 387
  • author_summary 15 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
387
110 last 30d - stable
Likes
0
Model age
4mo ago
created 2026-05-27
Downloads over time
Now422→from10↑4,120%
015430946310 on Jun 10422 on Oct 11JunJulAugSepOct
Jun 10 → Oct 11 · 57 snapshots · spans 123 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 2 formats · 126 downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
apache-2.0
Tags
transformers safetensors qwen3_5_moe image-text-to-text qwen3.6 moe vlm fp8 quantized blockwise vllm dgx-spark

Related

Total size
34.9 GB
Files
39
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-05-27 02:59

Files by quantization

Auxiliary files 39 files 34.9 GB
model-00001-of-00026.safetensors 2.41 GB 001d014d download
model-00006-of-00026.safetensors 1.84 GB f802ddd8 download
model-00016-of-00026.safetensors 1.84 GB e5c8ea41 download
model-00010-of-00026.safetensors 1.84 GB 2e63777d download
model-00018-of-00026.safetensors 1.84 GB 45d18d69 download
model-00008-of-00026.safetensors 1.84 GB 13963106 download
model-00025-of-00026.safetensors 1.79 GB 8de81a9d download
model-00014-of-00026.safetensors 1.59 GB 0f185431 download
model-00020-of-00026.safetensors 1.59 GB da9c9263 download
model-00012-of-00026.safetensors 1.59 GB f33e3f6c download
model-00022-of-00026.safetensors 1.57 GB be8d6441 download
model-00024-of-00026.safetensors 1.57 GB 27fbfae2 download
model-00004-of-00026.safetensors 1.57 GB caedcc10 download
model-00023-of-00026.safetensors 1.57 GB 3eb93c88 download
model-00003-of-00026.safetensors 1.57 GB 07d97e75 download
model-00005-of-00026.safetensors 1.57 GB ac089faf download
model-00026-of-00026.safetensors 1.55 GB 01213276 download
model-00002-of-00026.safetensors 963 MB e4de9027 download
model-00015-of-00026.safetensors 781 MB f1bacfcd download
model-00013-of-00026.safetensors 781 MB 595c28cb download
model-00021-of-00026.safetensors 781 MB 86e1cb31 download
model-00007-of-00026.safetensors 525 MB fba88c30 download
model-00011-of-00026.safetensors 525 MB 1004e820 download
model-00019-of-00026.safetensors 525 MB 67d5e4f8 download
model-00017-of-00026.safetensors 525 MB 8df7f73c download
model-00009-of-00026.safetensors 525 MB 929aded8 download
tokenizer.json 12.2 MB 5f9e4d49 download
model.safetensors.index.json 6.73 MB b4057feb download
vocab.json 6.41 MB 0aa0ce06 download
merges.txt 3.20 MB a494e019 download
config.json 36.1 KB b69eb1fd download
tokenizer_config.json 18.3 KB c6adc1b9 download
chat_template.jinja 11.0 KB 7e823502 download
README.md 4.79 KB 6a411d84 download
.gitattributes 1.53 KB 52373fe2 download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
generation_config.json 202 B 023756cf download
configuration.json 58.0 B d24dba94 download

README current version from Hugging Face


license: apache-2.0
base_model:

  • huihui-ai/Huihui-Qwen3.6-35B-A3B-abliterated
  • Qwen/Qwen3.6-35B-A3B
    tags:
  • qwen3.6
  • moe
  • vlm
  • fp8
  • quantized
  • blockwise
  • vllm
  • dgx-spark
    pipeline_tag: image-text-to-text
    library_name: transformers

Update 2025-05-06: Replaced chat_template in tokenizer_config.json with the fixed version from froggeric/Qwen-Fixed-Chat-Templates.

Huihui-Qwen3.6-35B-A3B-abliterated-FP8

Vision-capable FP8-quantized abliterated Qwen3.6-35B-A3B (MoE, hybrid mamba/attention) for Nvidia DGX Spark and other FP8-capable hardware (~80 GB VRAM for full 262k context).

I've tested many abliterated models from HF, and only Huihui makes really good ones.
Check "Claude" version if you like: batsclamp/Huihui-Qwen3.6-35B-A3B-Claude-4.6-Opus-abliterated-FP8

This one will give you ±50tps on full context when used with Eugr's vLLM (DGX Spark)

Model Lineage

Why FP8

Qwen3.6-35B-A3B in BF16 is ~72 GB on disk. FP8 cuts that to ~37 GB while preserving vision layers and precision-sensitive modules in BF16. The expected throughput uplift on DGX Spark is on par with what we saw for Qwen3.5 (31 → 51 t/s, ~65%).

Quantization Details

Scheme: native FP8 blockwise, identical on-disk format to the official Qwen/Qwen3.6-35B-A3B-FP8.

Field Value
quant_method fp8
activation_scheme dynamic (per-token, at inference)
fmt e4m3
weight_block_size [128, 128]
Scale dtype / key bf16, *.weight_scale_inv
Scale shape (ceil(out/128), ceil(in/128))

Quantized (weights → FP8 e4m3, per-block [128, 128] scales):

  • All 2D Linear *.weight in language layers that aren't in the exclusion list, including:
    • self_attn.{q,k,v,o}_proj (full-attention layers)
    • linear_attn.{in_proj_qkv, in_proj_z, out_proj} (linear-attention / mamba layers)
    • mlp.shared_expert.{gate,up,down}_proj
    • All 256 experts per MoE layer, un-fused to match Qwen's official per-expert layout:
      • mlp.experts.{0..255}.{gate_proj, up_proj, down_proj}.weight

Kept in BF16 (matches Qwen's modules_to_not_convert):

Module Reason
lm_head Output head — precision-sensitive
model.language_model.embed_tokens Embedding layer
*.input_layernorm, *.post_attention_layernorm LayerNorms
*.self_attn.{q_norm, k_norm} QK norms
*.linear_attn.{A_log, conv1d, dt_bias, in_proj_a, in_proj_b, in_proj_ba, norm} Mamba state-space params (small, sensitive)
*.mlp.gate, *.mlp.shared_expert_gate MoE router gates — routing precision matters
model.visual.* Entire visual encoder (patch_embed, 27 ViT blocks, deepstack mergers, merger)
mtp.* Multi-token prediction module

Notable Implementation Notes

  • Source experts were fused 3D (mlp.experts.gate_up_proj[256, 1024, 2048], mlp.experts.down_proj[256, 2048, 512]) — we un-fuse them to the per-expert layout the official Qwen FP8 uses (mlp.experts.{E}.{gate, up, down}_proj.weight). This is what vLLM's Fp8 MoE loader expects.
  • Streaming quantization: processed one source shard at a time on the GPU; peak host memory ~6 GB. Avoids the llmcompressor pitfall where peak VM grew to 168 GB during the Compressing phase and got OOM-killed on the 128 GB DGX Spark unified-memory budget.
  • Sanity check: round-trip dequantization median relative error ~2.2% per tensor (as expected for E4M3 blockwise).

Numbers

BF16 source This FP8
Size on disk ~72 GB ~37 GB
Tensors in index 1045 (fused experts) 64189 (un-fused)
FP8 weight tensors — 31738
BF16 weight tensors — 32451 (incl. 31738 weight_scale_inv)

Loading

from transformers import AutoModelForImageTextToText, AutoProcessor
model = AutoModelForImageTextToText.from_pretrained(
    "batsclamp/Huihui-Qwen3.6-35B-A3B-abliterated-FP8",
    dtype="auto",
    device_map="auto",
)
processor = AutoProcessor.from_pretrained("batsclamp/Huihui-Qwen3.6-35B-A3B-abliterated-FP8")

For vLLM, point at the repo — the quantization_config is already correctly set (quant_method: fp8, weight-block [128, 128], dynamic activations).

Disclaimer

Abliterated model. Not recommended if you expect a polite corporate assistant.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-05-27Duplicate from batsclamp/Huihui-Qwen3.6-35B-A3B-abliterated-FP8f2f03084.8 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration