← back to catalog · registered 2026-08-22 13:56

batsclamp/Huihui-Qwen3.5-35B-A3B-Claude-4.6-Opus-abliterated-FP8

batsclamp Qwen 34B MoE multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/batsclamp%2FHuihui-Qwen3.5-35B-A3B-Claude-4.6-Opus-abliterated-FP8"
Response includes
  • classification m1
  • files 13
  • hub_downloads_all_time 2,109
  • author_summary 7 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
2K
102 last 30d - cooling
Likes
1
Model age
6mo ago
created 2026-03-26
Downloads over time
Now2.1K→from271↑691%
1778951.6K2.3K271 on Mar 252.1K on Oct 11MarAprMayJunJulAugSepOct
Mar 25 → Oct 11 · 68 snapshots · spans 200 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
transformers safetensors qwen3_5_moe image-text-to-text qwen3.5 moe vlm fp8 quantized compressed-tensors vllm dgx-spark

Related

Total size
35.7 GB
Files
13
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-03-28 06:06

Files by quantization

Auxiliary files 13 files 35.7 GB
model.safetensors 33.3 GB eea77f11 download
model_visual.safetensors 2.41 GB ffb112c5 download
tokenizer.json 19.1 MB 87a7830d download
model.safetensors.index.json 5.64 MB 1b1e2313 download
config.json 16.9 KB e6ec0ec0 download
tokenizer_config.json 5.23 KB 4c3d5670 download
chat_template.jinja 3.95 KB 609532bf download
README.md 3.89 KB 2655b1bc download
.gitattributes 1.53 KB 52373fe2 download
recipe.yaml 393 B c3c9c69b download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
generation_config.json 218 B 5e471804 download

README current version from Hugging Face


license: apache-2.0
base_model:

  • huihui-ai/Huihui-Qwen3.5-35B-A3B-Claude-4.6-Opus-abliterated
  • Qwen/Qwen3.5-35B-A3B
  • Jackrong/Qwen3.5-35B-A3B-Claude-4.6-Opus-Reasoning-Distilled
    tags:
  • qwen3.5
  • moe
  • vlm
  • fp8
  • quantized
  • compressed-tensors
  • vllm
  • dgx-spark
    pipeline_tag: image-text-to-text
    library_name: transformers

Huihui-Qwen3.5-35B-A3B-Claude-4.6-Opus-abliterated-FP8

Vision-capable FP8 quantized fast abliterated distilled Qwen3.5-35B model made for Nvidia DGX Spark (~80GB VRAM is needed for full functionality)

Model Lineage

So first it was Qwen/Qwen3.5-35B-A3B (BF16).

Performance

Conservative approach to FP8 quantization caused minimum quality loss, while still bumping the speed from 31 t/s → 51 t/s on DGX Spark. With 262k context and some space for KV cache it uses 80GB VRAM (only).

Currently that's the best, fastest and abliterated model to be used on Nvidia DGX Spark, which also preserves all visual layers untouched.

I failed to find a case where this model will refuse to answer. It is especially funny to use with pictures ;). So far the best "tooling" skills — it really likes to Google stuff first even if it knows the answer.

I plan to test the quality of the model's output later and update this page.

Quantization Details

Quantized using the FP8_DYNAMIC scheme from llmcompressor (>=0.10) with compressed-tensors serialization.

Method

FP8_DYNAMIC is a data-free quantization scheme — no calibration dataset required. Weights are statically quantized to FP8 (per-channel, symmetric), while activations are dynamically quantized to FP8 (per-token, symmetric) at inference time.

Modules Excluded from Quantization

Matching the conservative strategy from Qwen/Qwen3.5-35B-A3B-FP8:

Module Reason
lm_head Output head — precision-sensitive
embed_tokens Embedding layer
linear_attn.conv1d, linear_attn.in_proj_a/b Linear attention layers
mlp.gate, mlp.shared_expert_gate MoE router gates — routing precision matters
model.visual.* Entire visual encoder kept at BF16
mtp.* Multi-token prediction layers

Post-processing

The model was quantized via AutoModelForCausalLM (the only loader proven to work with llmcompressor for this architecture), then post-processed:

  1. Weight key renaming — model.layers.X → model.language_model.layers.X to match the ConditionalGeneration format expected by vLLM
  2. Visual encoder restoration — BF16 vision encoder weights copied from the source model (since AutoModelForCausalLM strips them)
  3. Config restructuring — config.json rebuilt from the source model's nested structure with the quantization config injected

Resources

Disclaimer

It's an abliterated model. DO NOT use it if you think that all AIs need to be politically correct and boring.

README history 8 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-03-28Upload README.md with huggingface_hub38ad79b3.9 KB
    Loading...
  2. 2026-03-27Update README.md51115051.7 KB
    Loading...
  3. 2026-03-27Update README.md33cd59d1.7 KB
    Loading...
  4. 2026-03-27Update README.mdb54a73e1.7 KB
    Loading...
  5. 2026-03-27Update README.md71526131.6 KB
    Loading...
  6. 2026-03-27Update README.mdaa66f911.7 KB
    Loading...
  7. 2026-03-27Update README.md3dc35261.7 KB
    Loading...
  8. 2026-03-27Create README.md8683f2c1.5 KB
    Loading...

Discussions 2 threads

  1. 2026-04-14garbled characters.open2 💬#2
    Loading...
  2. 2026-03-29VLLM 0.18.0 Fails to Disable Thinking Mode for Qwen3.5-35B-A3B-FP8open4 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration