← back to catalog · registered 2026-08-22 13:56

braydenh563/Qwen3.6-35B-Uncensored-HauhauCS-1M-MTP-Ollama

braydenh563 Qwen 35B GGUF MoE multimodal second-order 1.0M ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/braydenh563%2FQwen3.6-35B-Uncensored-HauhauCS-1M-MTP-Ollama"
Response includes
  • classification m-uncensored
  • files 10
  • hub_downloads_all_time 10,574
  • author_summary 9 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
11K
852 last 30d - cooling
Likes
12
Model age
2mo ago
created 2026-07-15
Downloads over time
Now10.9K→from669↑1,525%
1594.1K8K11.9K669 on Jul 1510.9K on Oct 11JulAugSepOct
Jul 15 → Oct 11 · 53 snapshots · spans 88 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Quantizations
Q4_K
Tags
gguf long-context yarn qwen3.6 uncensored mtp speculative-decoding vision llama.cpp ollama image-text-to-text base_model:HauhauCS/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive

Related

Total size
20.2 GB
Files
10
Quantizations
3
Registered
2026-08-22 13:56
Last updated on HF
2026-08-06 04:51

Files by quantization

Q4_K 1 file 20.2 GB
Qwen3.6-35B-A3B-Uncensored-HauHauCS-1M-Ollama-Q4_K_M.gguf 20.2 GB b0737687 download
F16 1 file 858 MB
mmproj-qwen36-hauhau-f16.gguf 858 MB c8e70234 download
Auxiliary files 8 files 609 KB
banner.jpeg 440 KB 8cb897ed download
qwen36_heatmap.png 93.1 KB 0ee9a029 download
qwen36_mtp_speedup.png 52.3 KB 3e0a17d5 download
README.md 10.0 KB c93d8e7f download
chat_template_claude_code.jinja 7.53 KB fe27a7be download
results.jsonl 4.40 KB 6233dd7a download
.gitattributes 2.39 KB 7885f4b6 download
params 28.0 B 94a69034 download

README current version from Hugging Face


license: apache-2.0
pipeline_tag: image-text-to-text
base_model: HauhauCS/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive
tags:

  • gguf
  • long-context
  • yarn
  • qwen3.6
  • uncensored
  • mtp
  • speculative-decoding
  • vision
  • llama.cpp
  • ollama

Qwen3.6-35B Uncensored: 1M Context + MTP + Vision, one file


Note from braydenh563: The file naming conventions used by satgeze prevented Hugging Face from allowing the file to be pulled directly into ollama. I have renamed the gguf file to correct this issue. Everything else remains the same.

UPDATE (05/08/2026):

  • Following advice, I have updated this model so that it pulls with the paramater draft_num_predict: 4. This should mean that the grafted mtp drafter properly activates when pulled into local llm software like ollama.

Any and all feedback is appreciated!


HauhauCS's Aggressive uncensored build of Qwen3.6-35B-A3B, extended to a 1,048,576-token context and fused with the official MTP speculative-decoding layer. Every claim on the banner is measured, not inherited. Verification data ships in this repo.

Capability Status Evidence
1M contextCertified July 6, 202670/70 needles, full ladder 64K to 1M, 10 depths per rung, f16 KV, temp 0
MTP speculative decodingGrafted and vetted284.6 to 380.1 tok/s (+33.6%), draft acceptance 0.864, RTX 5090
VisionVerifiedmmproj tower reads image text and identifies objects correctly
UncensoredHauhauCS AggressiveTrunk weights bit-identical to the HauhauCS release

Needle-in-a-haystack certification

Native context is 262,144. YaRN rope scaling (factor 4) extends to 1,048,576, and retrieval stays perfect across the entire extension: every rung from 64K to 1M scored 10/10 at depths from 5 to 95 percent. Raw per-needle records are in results.jsonl.

MTP: how it got here and what it does

Qwen3.6 ships a multi-token-prediction layer in the official checkpoint, but uncensored community trunks usually drop it. This build grafts the official MTP layer (via Unsloth's MTP GGUF) back onto the uncensored trunk at the GGUF tensor level. The draft head then predicts ahead and the trunk verifies every token, so output is identical to standard decoding, only faster.

Measured on the uncensored trunk (RTX 5090, 16K ctx): 284.6 tok/s standard, 380.1 tok/s with --spec-type draft-mtp (+33.6 percent). Draft acceptance 0.864 with mean accepted run 3.59, which is higher than we measured on the official trunk (0.816). A 262K needle test with speculation active scored 10/10, confirming the graft changes nothing about retrieval.

Files

File Size What it is
qwen3.6-35b-uncensored-1M-MTP-Q4_K_M.gguf 21.7 GB The one-file package: uncensored trunk + 1M rope + MTP layer
mmproj-qwen36-hauhau-f16.gguf 899 MB Vision tower, attach at runtime
qwen36_heatmap.png, qwen36_mtp_speedup.png, results.jsonl small Verification evidence

The plain non-MTP 1M trunk and other quants live on the ModelScope mirror.

Every file, every mirror

Nothing was discontinued: every quant is one click away. Hugging Face carries the curated picks, ModelScope always carries everything, and Ollama serves ready-to-run tags.

On Ollama every tag ships with the vision tower bundled and the 1M rope metadata baked in.

File Size Hugging Face ModelScope Ollama
mmproj-qwen36-hauhau-f16.gguf 899 MB download download bundled in every tag
qwen3.6-35b-uncensored-1M-MTP-Q4_K_M.gguf 21.7 GB download download ollama run satgeze/qwen36-35b-uncensored-1m
qwen3.6-35b-uncensored-1M-Q4_K_M.gguf 21.2 GB on ModelScope download ollama run satgeze/qwen36-35b-uncensored-1m:q4_k_m-no-mtp

Run with Claude Code (or any Anthropic-API agent)

llama-server natively speaks the Anthropic messages protocol, so Claude Code can use this model directly. Two integration fixes are required, both included here:

llama-server -m qwen3.6-35b-uncensored-1M-MTP-Q4_K_M.gguf \
  -c 262144 -np 2 --cache-reuse 256 -ngl 99 --jinja \
  --chat-template-file chat_template_claude_code.jinja \
  --spec-type draft-mtp --spec-draft-n-max 3 \
  --mmproj mmproj-qwen36-hauhau-f16.gguf \
  --port 8040 --host 0.0.0.0

Then in the client shell:

export ANTHROPIC_BASE_URL=http://<server-ip>:8040
export ANTHROPIC_AUTH_TOKEN=local
export ANTHROPIC_MODEL=qwen36-35b-uncensored-1m-mtp
claude

Why the flags: chat_template_claude_code.jinja is the embedded template with one strictness check relaxed (the original raises an exception on the multiple system blocks agent harnesses send). -np 2 matters because Claude Code fires background side-requests; with a single slot they evict your main generation. Measured on an RTX 5090: 330 to 380 tok/s decode with speculation active, tool calls and vision working.

How this was built

  1. 1M context: YaRN rope-scaling metadata (factor 4.0, original context 262,144) written directly into the GGUF header with gguf-py. No weights change, no fine-tuning, no runtime flags needed: llama.cpp and Ollama read the baked metadata.
  2. MTP graft: the official MTP layer tensors (block 41 of the donor GGUF) appended onto the uncensored trunk, block_count and nextn_predict_layers metadata updated to match. GGUF surgery, not training.
  3. Certification: multi-needle retrieval harness (10 needles per rung at depths 5 to 95 percent, temperature 0, seeded to bust prompt caches) run against llama-server with f16 KV cache. Certification is only ever done on f16 KV; quantized KV is a labeled budget option, never the baseline.

Run it

llama.cpp (full speed, MTP active):

llama-server -m qwen3.6-35b-uncensored-1M-MTP-Q4_K_M.gguf \
  -c 1048576 -np 1 --jinja \
  --spec-type draft-mtp --spec-draft-n-max 3 \
  --mmproj mmproj-qwen36-hauhau-f16.gguf

Ollama (1M and vision work; MTP tensors load harmlessly but Ollama has no speculative decoding yet, so no speed gain there):

FROM ./qwen3.6-35b-uncensored-1M-MTP-Q4_K_M.gguf
RENDERER qwen3.5
PARSER qwen3.5
PARAMETER num_ctx 262144

The RENDERER and PARSER lines matter: imported GGUFs without a native renderer hit template bugs under tool-heavy use (for example as a coding-agent backend). Raise num_ctx as memory allows. Measured footprints: 44 GB total at 1M f16 KV on a 128 GB Mac (100 percent GPU, 72 tok/s decode); a 32 GB card holds roughly 262K fully resident.

Honest notes

  • Uncensoring quality versus the official model has not been independently benchmarked here. For capability benchmarks of the base model, see the official Qwen3.6-35B-A3B card; weights here are bit-identical to the HauhauCS release apart from rope metadata and the added MTP layer, so base capability carries over modulo Q4_K_M quantization.
  • The MTP layer came from the official checkpoint and pairs with this trunk architecture; acceptance 0.864 was measured on real coding prompts, and speculative decoding never changes outputs by construction.

Engine bench: same model, three engines

Measured July 8, 2026 on this exact Q4_K_M + MTP build, M3 Max 128GB, identical prompt, uniform wall-clock method, idle machine:

Engine Decode tok/s
llama.cpp llama-server + MTP 79.1
MLX 72.7
Ollama 65.6

llama-server with the MTP draft head is the fastest way to run this model on Apple silicon, about 20 percent ahead of Ollama (which does not use speculative decoding for imported GGUFs).

How to actually use a 1M-context model

Habits that measurably help, from our RULER, hop and adherence testing across this fleet:

  1. Re-state standing instructions near the end of long prompts; recency beats depth.
  2. One big reference dump beats a long accumulated conversation. Fresh session per task.
  3. After any compaction or summarization, repeat your active rules yourself.
  4. Prefill at 500K+ takes real time on any hardware; stage your questions accordingly.
  5. Know your quant: the results tables on this card show what each quant actually holds at depth; pick the strongest one your memory allows.

Credits

Base model: Qwen (Apache-2.0), including the official MTP layer. Uncensoring and trunk quant: HauhauCS. MTP GGUF packaging: Unsloth. 1M YaRN extension, MTP graft, and certification: SatGeze.

Mirrors: Hugging Face | ModelScope. Tooling: github.com/satindergrewal/aviary-1m.

README history 9 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-06Update README.md5cea01110 KB
    Loading...
  2. 2026-08-06Update README.md460688f10 KB
    Loading...
  3. 2026-08-05Update README.mdc1e715f10 KB
    Loading...
  4. 2026-08-05Update README.mdd31849910.2 KB
    Loading...
  5. 2026-08-05Update README.mded7d4c410.2 KB
    Loading...
  6. 2026-08-05Update README.md12548d610 KB
    Loading...
  7. 2026-08-05Update README.mde8f4fcb10 KB
    Loading...
  8. 2026-07-15Update README.mdbe5a8889.7 KB
    Loading...
  9. 2026-07-15Duplicate from satgeze/Qwen3.6-35B-Uncensored-HauhauCS-1M-GGUF97790309.5 KB
    Loading...

Discussions 2 threads

  1. 2026-09-23Is this better than the qwen 3.8 27 b turbo fable Neo max coderopen1 💬#2
    Loading...
  2. 2026-08-05MTP using Ollama version 0.32.6-rc0open1 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration