← back to catalog · registered 2026-08-22 13:56

morikomorizz/Qwen3.6-35B-A3B-Uncensored-HauhauCS-MTP

morikomorizz Qwen 35B GGUF MoE second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/morikomorizz%2FQwen3.6-35B-A3B-Uncensored-HauhauCS-MTP"
Response includes
  • classification m-uncensored
  • files 9
  • hub_downloads_all_time 68,020
  • author_summary 3 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
68K
10K last 30d - stable
Likes
10
Model age
4mo ago
created 2026-06-05
Downloads over time
Now71.4K→from1.4K↑5,090%
026.1K52.2K78.4K1.4K on Jun 1071.4K on Oct 11JunJulAugSepOct
Jun 10 → Oct 11 · 57 snapshots · spans 123 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en
Quantizations
Q2_K Q3_K Q4_K Q5_K Q6_K Q8_K
Tags
gguf en base_model:HauhauCS/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive base_model:quantized:HauhauCS/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive license:apache-2.0 region:us

Related

Total size
154 GB
Files
9
Quantizations
8
Registered
2026-08-22 13:56
Last updated on HF
2026-06-10 14:52

Files by quantization

Q8_K 1 file 41.4 GB
Qwen3.6-35B-A3B-Uncensored-HauhauCS-MTP-Q8_K_P.gguf 41.4 GB 82be937a download
Q6_K 1 file 29.4 GB
Qwen3.6-35B-A3B-Uncensored-HauhauCS-MTP-Q6_K_P.gguf 29.4 GB 76a0d4c2 download
Q5_K 1 file 26.9 GB
Qwen3.6-35B-A3B-Uncensored-HauhauCS-MTP-Q5_K_P.gguf 26.9 GB 833c833f download
Q4_K 1 file 22.7 GB
Qwen3.6-35B-A3B-Uncensored-HauhauCS-MTP-Q4_K_P.gguf 22.7 GB d4c1bc57 download
Q3_K 1 file 18.6 GB
Qwen3.6-35B-A3B-Uncensored-HauhauCS-MTP-Q3_K_P.gguf 18.6 GB 21b86d5b download
Q2_K 1 file 14.8 GB
Qwen3.6-35B-A3B-Uncensored-HauhauCS-MTP-Q2_K_P.gguf 14.8 GB 717690d9 download
F16 1 file 858 MB
mmproj-Qwen3.6-35B-A3B-Uncensored-HauhauCS-MTP-f16.gguf 858 MB c8e70234 download
Auxiliary files 2 files 5.15 KB
README.md 2.98 KB bf07a045 download
.gitattributes 2.17 KB 3a56b896 download

README current version from Hugging Face


license: apache-2.0
language:

  • en
    base_model:
  • HauhauCS/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive

Qwen3.6-35B-A3B-Uncensored-HauhauCS-MTP

This model is a modified version of Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive grafted with the Multi-Token Prediction (MTP) module using the MTP donor from the Qwen 3.6-35B-A3B-MTP-GGUF series by Unsloth.

This modification aims to provide faster inference speeds via MTP-based speculative decoding without sacrificing the base model's original quality or capabilities.

Specifications

  • Base Model: HauhauCS/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive
  • MTP Donor: unsloth/Qwen3.6-35B-A3B-MTP-GGUF
  • Architecture: Mixture of Experts (MoE) — 35B total parameters / ~3B active per forward pass (256 experts, 8 routed per token)
  • Context Window: 262K (262,144 tokens)
  • Multimodal Capabilities: Supports text, image, and video processing
  • Uncensored Nature: Inherits the Aggressive variant from HauhauCS (0/465 refusals on standard evaluation datasets, removing default refusal behavior while maintaining base performance and model traits).

Key MTP Features

  • Inference Speedup: Offers an estimated speedup of 1.4x to 2.2x faster generation (depending on hardware specifications and the inference backend).
  • Consistent Quality: Retains the same output distribution as the base model, meaning no loss in generation accuracy.

Inference & Usage Guide

To utilize the MTP features, you need an inference engine that supports MTP speculative decoding, such as the latest versions of llama.cpp, Unsloth Studio, or SGLang.

Example via llama.cpp (Server CLI)

Run the server with the following arguments to enable the MTP draft module:

llama-cli -m Qwen3.6-35B-A3B-Uncensored-HauhauCS-MTP-Q8_K_P.gguf.gguf \
  --mmproj mmproj-Qwen3.6-35B-A3B-Uncensored-HauhauCS-MTP-f16.gguf \
  --jinja -c 131072 -ngl 99

Notes:

  • Adjust -ngl (GPU offload layers) based on your system's VRAM capacity.
  • The flags --spec-type draft-mtp and --spec-draft-n-max 2 (can be configured up to 6 on capable systems) enable the MTP drafting mechanism.
  • Currently, llama.cpp's MTP implementation does not fully support multi-user scenarios (-np > 1) or concurrent multimodal inputs (--mmproj).

Recommended Sampling Parameters

  • Temperature: 1.0 (or 0.7–0.8 for guided instruction tasks)
  • Top_P: 0.95
  • Min_P: 0.00 (or 0.05 to filter out low-probability tokens)
  • Repeat Penalty: 1.0
  • Presence Penalty: 1.5 (optional, to minimize repetitive sentences in longer contexts)
  • Jinja Template: Use the --jinja flag in llama.cpp to parse instructions with the correct format. If you prefer to disable the built-in thinking mode, you can pass {"enable_thinking": false} in your template configuration.

README history 5 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-05Update README.mde9ccb4d3 KB
    Loading...
  2. 2026-06-05Update README.mdd6d47ad3 KB
    Loading...
  3. 2026-06-05Update README.md554fac93 KB
    Loading...
  4. 2026-06-05Update README.mdc06731a112 B
    Loading...
  5. 2026-06-05initial commitc563e5b28 B
    Loading...

Discussions 1 thread

  1. 2026-06-10Q6?open3 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration