← back to catalog · registered 2026-08-22 13:56

i-morxi/Huihui-Qwen3.5-35B-A3B-abliterated-autoround

i-morxi Qwen 4.3B MoE
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/i-morxi%2FHuihui-Qwen3.5-35B-A3B-abliterated-autoround"
Response includes
  • classification m1
  • files 31
  • benchmarks 11 entries
  • hub_downloads_all_time 365
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
365
22 last 30d - cooling
Likes
0
Model age
6mo ago
created 2026-03-17
Downloads over time
Now369→from45↑720%
2915327740145 on Mar 18369 on Oct 11369 on Oct 7MarAprMayJunJulAugSepOct
Mar 18 → Oct 11 · 69 snapshots · spans 207 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 0.8 UGI
Hazardous 4.1 UGI
Natural Intelligence 24.97 UGI
Political lean -20.7% UGI
Sensitive-Info 20.98 UGI
SocPol 2.1 UGI
UGI 23.15 UGI
Willingness (10) 2.8 UGI
W10-Adherence 1.5 UGI
W10-Direct 4 UGI
Writing 37 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

Tags
safetensors qwen3_5_moe arxiv:2309.05516 base_model:Qwen/Qwen3.5-35B-A3B base_model:quantized:Qwen/Qwen3.5-35B-A3B 4-bit auto-round region:us

Related

Total size
19.6 GB
Files
31
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-03-18 01:37

Files by quantization

Auxiliary files 31 files 19.7 GB
model-00012-of-00020.safetensors 1.00 GB bbf60649 download
model-00010-of-00020.safetensors 1.00 GB ad8ad27a download
model-00015-of-00020.safetensors 1.00 GB 90a55e8d download
model-00005-of-00020.safetensors 1.00 GB 96e82c4c download
model-00008-of-00020.safetensors 1.00 GB b8142460 download
model-00003-of-00020.safetensors 1.00 GB 46dfa4e5 download
model-00014-of-00020.safetensors 1.00 GB d93c5ce5 download
model-00009-of-00020.safetensors 1.00 GB f685f4ed download
model-00016-of-00020.safetensors 1.00 GB e7a547c3 download
model-00011-of-00020.safetensors 1.00 GB 1025c541 download
model-00006-of-00020.safetensors 1.00 GB 59e0bbd1 download
model-00004-of-00020.safetensors 1.00 GB 25599c60 download
model-00007-of-00020.safetensors 1.00 GB 29f95808 download
model-00013-of-00020.safetensors 1.00 GB 708bf8bd download
model-00002-of-00020.safetensors 1.00 GB 09e3ea84 download
model-00001-of-00020.safetensors 1.00 GB 3f95b4a3 download
model-00017-of-00020.safetensors 1020 MB a0b9fdfa download
model-00019-of-00020.safetensors 970 MB 2dba0213 download
model-00020-of-00020.safetensors 970 MB d69796f9 download
model_extra_tensors.safetensors 436 MB f4c65c8f download
model-00018-of-00020.safetensors 330 MB 0ee71bcf download
tokenizer.json 19.1 MB 87a7830d download
model.safetensors.index.json 9.63 MB 4e9c02c2 download
config.json 23.4 KB 628f2f91 download
quantization_config.json 18.4 KB bd5a243a download
chat_template.jinja 7.57 KB a585dec8 download
README.md 4.37 KB 9bddd450 download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.27 KB 7ad6acdf download
tokenizer_config.json 1.11 KB 541f6c47 download
generation_config.json 213 B dff89e46 download

README current version from Hugging Face


base_model:

  • Qwen/Qwen3.5-35B-A3B

Model Details

This model is a int4 model with group_size 128 of huihui-ai/Huihui-Qwen3.5-35B-A3B-abliterated generated by intel/auto-round. Please follow the license of the original model.

vllm Infernece Example

pip install git+https://github.com/vllm-project/vllm.git@main


pip install git+https://github.com/huggingface/transformers.git
vllm serve Intel/Qwen3.5-35B-A3B-int4-AutoRound  --port 8000   --tensor-parallel-size 1  --max-model-len 2048 --reasoning-parser qwen3 --served-model-name qwen --speculative-config '{"method":"qwen3_next_mtp","num_speculative_tokens":2}'
curl http://localhost:8000/v1/chat/completions -H "Content-Type: application/json" -d ' {
    "model": "qwen",
    "messages": [
      {"role": "system", "content": "You are a helpful assistant."},
      {"role": "user", "content": "Summarize Qwen 3.5 in one sentence."}
    ],
    "temperature": 1,
    "max_tokens": 512
  } '

Transformers Inference

Transformers v5.2.0

auto-round and This pr is required https://github.com/intel/auto-round/pull/1476

from transformers import AutoProcessor, Qwen3_5MoeForConditionalGeneration
model_name = "Intel/Qwen3.5-35B-A3B-int4-AutoRound"

model = Qwen3_5MoeForConditionalGeneration.from_pretrained(model_name, dtype="auto",
                                                                device_map="auto")
processor = AutoProcessor.from_pretrained(model_name)

messages = [
    {
        "role": "user",
        "content": [
            {
                "type": "image",
                "image": "https://qianwen-res.oss-cn-beijing.aliyuncs.com/Qwen-VL/assets/demo.jpeg",
            },
            {"type": "text", "text": "Describe this image in short."},
        ],
    }
]


inputs = processor.apply_chat_template(
    messages,
    tokenize=True,
    add_generation_prompt=True,
    return_dict=True,
    return_tensors="pt"
)
inputs = inputs.to(model.device)


generated_ids = model.generate(**inputs, max_new_tokens=128)
generated_ids_trimmed = [
    out_ids[len(in_ids):] for in_ids, out_ids in zip(inputs.input_ids, generated_ids)
]
print(processor.batch_decode(generated_ids_trimmed, skip_special_tokens=True, clean_up_tokenization_spaces=False)[0])
"""
The user wants a short description of the image.

1.  **Identify the main subjects:** A woman and a large yellow dog (looks like a Labrador).
2.  **Identify the setting:** A sandy beach with the ocean in the background. The lighting suggests sunset or sunrise (golden hour).
3.  **Identify the action:** The dog is sitting and placing its paw in the woman's hand. The woman is smiling and sitting on the sand.
4.  **Synthesize into a short description:** A woman sitting on a beach at sunset gives a high-five (or "shake") to

"""

Generate the Model

This pr is required https://github.com/intel/auto-round/pull/1476

auto-round  "huihui-ai/Huihui-Qwen3.5-35B-A3B-abliterated"  --output_dir "./Qwen35-int4" --ignore_layers shared_expert,mtp.fc

Ethical Considerations and Limitations

The model can produce factually incorrect output, and should not be relied on to produce factually accurate information. Because of the limitations of the pretrained model and the finetuning datasets, it is possible that this model could generate lewd, biased or otherwise offensive outputs.

Therefore, before deploying any applications of the model, developers should perform safety testing.

Caveats and Recommendations

Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model.

Here are a couple of useful links to learn more about Intel's AI software:

Disclaimer

The license on this model does not constitute legal advice. We are not responsible for the actions of third parties who use this model. Please consult an attorney before using this model for commercial purposes.

Cite

@article{cheng2023optimize, title={Optimize weight rounding via signed gradient descent for the quantization of llms}, author={Cheng, Wenhua and Zhang, Weiwei and Shen, Haihao and Cai, Yiyang and He, Xin and Lv, Kaokao and Liu, Yi}, journal={arXiv preprint arXiv:2309.05516}, year={2023} }

arxiv

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-03-18Update README.md9c51de84.4 KB
    Loading...
  2. 2026-03-18Update README.md4cfd8774.4 KB
    Loading...
  3. 2026-03-18Create README.md690221f79 B
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration