← back to catalog · registered 2026-08-22 13:56

prithivMLmods/Qwen3.5-35B-A3B-abliterated-v2-MAX

prithivMLmods Qwen 35B GGUF MoE multimodal 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/prithivMLmods%2FQwen3.5-35B-A3B-abliterated-v2-MAX"
Response includes
  • classification m8
  • files 20
  • benchmarks 11 entries
  • hub_downloads_all_time 5,112
  • author_summary 98 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
5K
372 last 30d - cooling
Likes
3
Descendants
2
in 2 direct forks
Model age
6mo ago
created 2026-04-05
Downloads over time
Now5.2K→from2.5K↑112%
2.3K3.4K4.4K5.5K2.5K on Apr 155.2K on Oct 11AprMayJunJulAugSepOct
Apr 15 → Oct 11 · 65 snapshots · spans 179 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 0.8 UGI
Hazardous 4.1 UGI
Natural Intelligence 24.97 UGI
Political lean -20.7% UGI
Sensitive-Info 20.98 UGI
SocPol 2.1 UGI
UGI 23.15 UGI
Willingness (10) 2.8 UGI
W10-Adherence 1.5 UGI
W10-Direct 4 UGI
Writing 37 UGI

Genealogy 2 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en
Tags
transformers safetensors gguf qwen3_5_moe image-text-to-text text-generation-inference moe uncensored abliterated unfiltered unredacted refusal-ablated

Related

Total size
65.4 GB
Files
20
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-06-01 08:25

Files by quantization

Auxiliary files 20 files 65.4 GB
model-00001-of-00011.safetensors 6.60 GB 2094cb01 download
model-00004-of-00011.safetensors 6.27 GB ebb9a25f download
model-00005-of-00011.safetensors 6.27 GB f2b23e5c download
model-00006-of-00011.safetensors 6.27 GB fd4d5dd4 download
model-00007-of-00011.safetensors 6.27 GB e7b829c3 download
model-00008-of-00011.safetensors 6.27 GB 5aa73aff download
model-00009-of-00011.safetensors 6.27 GB b9d7b27f download
model-00010-of-00011.safetensors 6.27 GB eb2601c2 download
model-00003-of-00011.safetensors 6.27 GB ec53de38 download
model-00002-of-00011.safetensors 6.27 GB a0bd1532 download
model-00011-of-00011.safetensors 2.39 GB b3371954 download
tokenizer.json 19.1 MB 87a7830d download
model.safetensors.index.json 4.08 MB f3f14d9b download
chat_template.jinja 7.57 KB a585dec8 download
README.md 4.89 KB 0f6625e4 download
config.json 3.14 KB e6594889 download
.gitattributes 1.87 KB cf4fcc9f download
processor_config.json 1.16 KB 33818c7f download
tokenizer_config.json 1.11 KB 541f6c47 download
generation_config.json 218 B fd60af71 download

README current version from Hugging Face


license: apache-2.0
tags:

  • text-generation-inference
  • moe
  • uncensored
  • abliterated
  • unfiltered
  • unredacted
  • refusal-ablated
  • vllm
  • pytorch
  • bf16
  • max
  • alignment-modified
  • reasoning
  • v2.0
    base_model:
  • Qwen/Qwen3.5-35B-A3B
    language:
  • en
    pipeline_tag: image-text-to-text
    library_name: transformers

1

Qwen3.5-35B-A3B-abliterated-v2-MAX

Qwen3.5-35B-A3B-abliterated-v2-MAX is an optimized release built on top of huihui-ai/Huihui-Qwen3.5-35B-A3B-abliterated. This version focuses on updated shard sizing, repository optimization, and compatibility improvements for the latest Transformers releases, while preserving the reasoning and instruction-following capabilities of the original model. The result is a powerful 35B parameter Mixture-of-Experts language model designed for efficient deployment, stable inference, and modern ecosystem integration.

[!IMPORTANT]
This model is developed for research and learning purposes only. Any content generated by this model is used at the user's own risk. The authors and hosting platform disclaim any liability for outputs produced by this model. Users are responsible for ensuring safe, ethical, and lawful usage.

Compression for the Model

Qwen3.5-35B-A3B-abliterated-v2-MAX

Format Description Link
GGUF Quantized GGUF format https://huggingface.co/prithivMLmods/Qwen3.5-35B-A3B-abliterated-v2-MAX/tree/main/GGUF

Key Highlights

  • Latest Transformers Compatibility
    Re-sharded and optimized for improved compatibility with recent Transformers releases.

  • Optimized Model Sharding
    Updated shard structure for better storage handling, download reliability, and inference efficiency.

  • Stable Inference Pipeline
    Improved packaging for consistent loading and generation behavior across environments.

  • 35B MoE Architecture (A3B)
    Built on Qwen3.5-35B-A3B, leveraging Mixture-of-Experts design for scalable reasoning capacity.

  • Improved Deployment Stability
    Designed for smoother inference across different hardware configurations and runtimes.

  • Preserved Model Behavior
    No changes to weights or architecture; behavior remains consistent with the base model lineage.


Base Model Signatures:

This model has been re-sharded and optimized for the latest Transformers version from the base model: https://huggingface.co/huihui-ai/Huihui-Qwen3.5-35B-A3B-abliterated.


Quick Start with Transformers

pip install transformers==5.5.0
# or
pip install git+https://github.com/huggingface/transformers.git
from transformers import Qwen3_5MoeForConditionalGeneration, AutoProcessor
import torch

model = Qwen3_5MoeForConditionalGeneration.from_pretrained(
    "prithivMLmods/Qwen3.5-35B-A3B-abliterated-v2-MAX",
    torch_dtype="auto",
    device_map="auto"
)

processor = AutoProcessor.from_pretrained(
    "prithivMLmods/Qwen3.5-35B-A3B-abliterated-v2-MAX"
)

messages = [
    {
        "role": "user",
        "content": [
            {"type": "text", "text": "Explain how transformer models work in simple terms."}
        ],
    }
]

text = processor.apply_chat_template(
    messages, tokenize=False, add_generation_prompt=True
)

inputs = processor(
    text=[text],
    padding=True,
    return_tensors="pt"
).to("cuda")

generated_ids = model.generate(**inputs, max_new_tokens=256)

generated_ids_trimmed = [
    out_ids[len(in_ids):] for in_ids, out_ids in zip(inputs.input_ids, generated_ids)
]

output_text = processor.batch_decode(
    generated_ids_trimmed,
    skip_special_tokens=True,
    clean_up_tokenization_spaces=False
)

print(output_text)

Intended Use

  • Multimodal and Language Research
    Studying large-scale MoE behavior and inference characteristics.

  • Red-Teaming & Evaluation
    Testing robustness across complex and adversarial prompts.

  • High-Performance Deployment
    Running large MoE models on optimized multi-GPU setups.

  • Research Prototyping
    Experimentation with scalable transformer architectures and deployment workflows.

Limitations & Risks

Important Note: This model inherits the behavior and limitations of its base model.

  • Output Variability
    Responses may vary depending on sampling settings and prompt structure.

  • Resource Requirements
    A 35B MoE model requires significant GPU memory or optimized inference strategies such as quantization or tensor parallelism.

  • Deployment Constraints
    Performance depends heavily on hardware configuration and runtime optimization.

  • General Model Limitations
    May produce incorrect, incomplete, or inconsistent outputs in complex scenarios.

README history 8 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-01Update README.md5503f2f4.9 KB
    Loading...
  2. 2026-06-01Update README.mdcbb85194.8 KB
    Loading...
  3. 2026-06-01Update README.md761b9035.1 KB
    Loading...
  4. 2026-04-09Update README.mdff0497c4.8 KB
    Loading...
  5. 2026-04-09Update README.mdab699ac4.6 KB
    Loading...
  6. 2026-04-09Update README.md698badf141 B
    Loading...
  7. 2026-04-09Update README.md982d638140 B
    Loading...
  8. 2026-04-05initial commit6783c1928 B
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration