← back to catalog · registered 2026-08-22 13:56

prithivMLmods/Qwen3.6-35B-A3B-abliterated-MAX

prithivMLmods Qwen 35B MoE multimodal
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/prithivMLmods%2FQwen3.6-35B-A3B-abliterated-MAX"
Response includes
  • classification m1
  • files 16
  • benchmarks 11 entries
  • hub_downloads_all_time 397
  • author_summary 98 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
397
128 last 30d - stable
Likes
2
Descendants
4
in 4 direct forks
Model age
5mo ago
created 2026-04-18
Downloads over time
Now435→from40↑988%
2017232347540 on Apr 22435 on Oct 11435 on Oct 10AprMayJunJulAugSepOct
Apr 22 → Oct 11 · 64 snapshots · spans 172 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 1.4 UGI
Hazardous 0 UGI
Natural Intelligence 25.43 UGI
Political lean -19.6% UGI
Sensitive-Info 14.03 UGI
SocPol 2.6 UGI
UGI 16.02 UGI
Willingness (10) 2 UGI
W10-Adherence 0 UGI
W10-Direct 4 UGI
Writing 35.83 UGI

Genealogy 4 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 2 formats · 495 downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
apache-2.0
Languages
en
Tags
transformers safetensors qwen3_5_moe image-text-to-text text-generation-inference uncensored abliterated unfiltered unredacted refusal-ablated vllm pytorch

Related

Total size
65.4 GB
Files
16
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-06-01 08:10

Files by quantization

Auxiliary files 16 files 65.4 GB
model-00002-of-00007.safetensors 9.84 GB 9efdfa4d download
model-00001-of-00007.safetensors 9.79 GB 64a4b51e download
model-00003-of-00007.safetensors 9.46 GB ada6bfd9 download
model-00005-of-00007.safetensors 9.46 GB 56b22c82 download
model-00004-of-00007.safetensors 9.34 GB eb0eb699 download
model-00006-of-00007.safetensors 9.34 GB f4336010 download
model-00007-of-00007.safetensors 8.16 GB a6c9f762 download
tokenizer.json 19.1 MB 87a7830d download
model.safetensors.index.json 99.5 KB ec49928e download
chat_template.jinja 7.58 KB a8755d82 download
README.md 4.50 KB 896163de download
config.json 3.11 KB 6612d5fe download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.16 KB 33818c7f download
tokenizer_config.json 1.11 KB 541f6c47 download
generation_config.json 213 B 23a0a961 download

README current version from Hugging Face


license: apache-2.0
language:

  • en
    base_model:
  • Qwen/Qwen3.6-35B-A3B
    pipeline_tag: image-text-to-text
    library_name: transformers
    tags:
  • text-generation-inference
  • uncensored
  • abliterated
  • unfiltered
  • unredacted
  • refusal-ablated
  • vllm
  • pytorch
  • bf16
  • max
  • alignment-modified
  • reasoning
  • agent

1

Qwen3.6-35B-A3B-Abliterated-MAX

Qwen3.6-35B-A3B-Abliterated-MAX is an optimized release built on top of huihui-ai/Huihui-Qwen3.6-35B-A3B-abliterated. This version focuses on updated shard sizing, repository optimization, and compatibility improvements for the latest Transformers releases, while preserving the MoE architecture and reasoning capabilities of the original model. The result is a powerful 35B Mixture-of-Experts language model designed for efficient inference, stable deployment, and modern ecosystem integration.

[!IMPORTANT]
This model is intended for research and learning purposes only. Any content generated by this model is used at the user's own risk. The authors and hosting page disclaim any liability for outputs produced by this model. Users are responsible for ensuring safe, ethical, and lawful usage.


Key Highlights

  • Latest Transformers Compatibility
    Re-sharded and optimized for improved compatibility with recent Transformers releases.

  • Optimized Model Sharding
    Updated shard structure for better storage handling, download reliability, and inference efficiency.

  • Stable Inference Pipeline
    Improved packaging for consistent loading and generation behavior across environments.

  • 35B MoE Architecture (A3B)
    Built on Qwen/Qwen3.6-35B-A3B, leveraging Mixture-of-Experts design for scalable reasoning capacity.

  • Improved Deployment Stability
    Designed for smoother inference across different hardware configurations and runtimes.

  • Preserved Model Behavior
    No changes to weights or architecture; behavior remains consistent with the original model lineage.


Base Model Signatures:

This model has been re-sharded and optimized for the latest Transformers version from the base model:
https://huggingface.co/huihui-ai/Huihui-Qwen3.6-35B-A3B-abliterated


Quick Start with Transformers

pip install transformers==5.5.4
# or
pip install git+https://github.com/huggingface/transformers.git
from transformers import Qwen3_5MoeForConditionalGeneration, AutoProcessor
import torch

model = Qwen3_5MoeForConditionalGeneration.from_pretrained(
    "prithivMLmods/Qwen3.6-35B-A3B-Abliterated-MAX",
    torch_dtype="auto",
    device_map="auto"
)

processor = AutoProcessor.from_pretrained(
    "prithivMLmods/Qwen3.6-35B-A3B-Abliterated-MAX"
)

messages = [
    {
        "role": "user",
        "content": [
            {"type": "text", "text": "Explain how transformer models work in simple terms."}
        ],
    }
]

text = processor.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True
)

inputs = processor(
    text=[text],
    padding=True,
    return_tensors="pt"
).to("cuda")

generated_ids = model.generate(**inputs, max_new_tokens=256)

generated_ids_trimmed = [
    out_ids[len(in_ids):] for in_ids, out_ids in zip(inputs.input_ids, generated_ids)
]

output_text = processor.batch_decode(
    generated_ids_trimmed,
    skip_special_tokens=True,
    clean_up_tokenization_spaces=False
)

print(output_text)

Intended Use

  • Multimodal and Language Research
    Studying large-scale MoE behavior and inference characteristics.

  • Red-Teaming & Evaluation
    Testing robustness across challenging and adversarial prompts.

  • High-Performance Deployment
    Running large MoE models on optimized multi-GPU or distributed setups.

  • Research Prototyping
    Experimentation with scalable transformer architectures and deployment strategies.


Limitations & Risks

Important Note: This model inherits the behavior and limitations of its base model.

  • Output Variability
    Responses may vary depending on sampling parameters and prompt structure.

  • Resource Requirements
    A 35B MoE model requires significant GPU memory or optimized inference strategies such as quantization or tensor parallelism.

  • Deployment Constraints
    Performance depends heavily on hardware configuration and runtime optimization.

  • General Model Limitations
    May produce incorrect, incomplete, or inconsistent outputs in complex scenarios.

README history 8 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-01Update README.mdee79d3f4.5 KB
    Loading...
  2. 2026-06-01Update README.mdf5129335.1 KB
    Loading...
  3. 2026-04-19Update README.md351085d4.8 KB
    Loading...
  4. 2026-04-19Update README.md7bb0fad4.8 KB
    Loading...
  5. 2026-04-19Update README.md87de4d54.8 KB
    Loading...
  6. 2026-04-19Update README.mdcc48f414.6 KB
    Loading...
  7. 2026-04-19Update README.mdb821deb171 B
    Loading...
  8. 2026-04-18initial commit075617b28 B
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration