← back to catalog · registered 2026-08-22 13:56

cahlen/Qwen3.5-35B-A3B-Uncensored-HauhauCS-Aggressive-GGUF

cahlen Qwen 35B GGUF MoE multimodal second-order 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/cahlen%2FQwen3.5-35B-A3B-Uncensored-HauhauCS-Aggressive-GGUF"
Response includes
  • classification m-uncensored
  • files 13
  • hub_downloads_all_time 10,731
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
11K
1K last 30d - cooling
Likes
2
Model age
6mo ago
created 2026-04-02
Downloads over time
Now11K→from6.1K↑80%
5.8K7.7K9.6K11.5K6.1K on Apr 1511K on Oct 11AprMayJunJulAugSepOct
Apr 15 → Oct 11 · 65 snapshots · spans 179 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh
Quantizations
IQ1 IQ2 IQ3 Q2_K Q3_K Q4_K Q5_K
Tags
gguf quantized llama-cpp qwen qwen3.5 moe vision multimodal uncensored en zh base_model:HauhauCS/Qwen3.5-35B-A3B-Uncensored-HauhauCS-Aggressive

Related

Total size
137 GB
Files
13
Quantizations
9
Registered
2026-08-22 13:56
Last updated on HF
2026-04-03 18:30

Files by quantization

Q5_K 1 file 22.3 GB
Qwen3.5-35B-A3B-Uncensored-HauhauCS-Aggressive-Q5_K_S.gguf 22.3 GB 9a8b677c download
Q4_K 1 file 18.5 GB
Qwen3.5-35B-A3B-Uncensored-HauhauCS-Aggressive-Q4_K_S.gguf 18.5 GB b60d50e2 download
Q3_K 2 files 31.0 GB
Qwen3.5-35B-A3B-Uncensored-HauhauCS-Aggressive-Q3_K_L.gguf 16.9 GB d2453090 download
Qwen3.5-35B-A3B-Uncensored-HauhauCS-Aggressive-Q3_K_S.gguf 14.1 GB 91d08c04 download
IQ3 2 files 26.9 GB
Qwen3.5-35B-A3B-Uncensored-HauhauCS-Aggressive-IQ3_S.gguf 14.2 GB 1152307a download
Qwen3.5-35B-A3B-Uncensored-HauhauCS-Aggressive-IQ3_XXS.gguf 12.7 GB 800e0d86 download
Q2_K 1 file 12.1 GB
Qwen3.5-35B-A3B-Uncensored-HauhauCS-Aggressive-Q2_K.gguf 12.1 GB 7cc72f4f download
IQ2 2 files 18.8 GB
Qwen3.5-35B-A3B-Uncensored-HauhauCS-Aggressive-IQ2_S.gguf 9.92 GB 3fe0c9ed download
Qwen3.5-35B-A3B-Uncensored-HauhauCS-Aggressive-IQ2_XXS.gguf 8.85 GB c5032a83 download
IQ1 1 file 7.67 GB
Qwen3.5-35B-A3B-Uncensored-HauhauCS-Aggressive-IQ1_M.gguf 7.67 GB e7ba883b download
F16 1 file 858 MB
mmproj-Qwen3.5-35B-A3B-Uncensored-HauhauCS-Aggressive-f16.gguf 858 MB 04363e64 download
Auxiliary files 2 files 8.06 KB
README.md 5.55 KB e84533e2 download
.gitattributes 2.50 KB f29079ba download

README current version from Hugging Face


base_model: HauhauCS/Qwen3.5-35B-A3B-Uncensored-HauhauCS-Aggressive
tags:

  • gguf
  • quantized
  • llama-cpp
  • qwen
  • qwen3.5
  • moe
  • vision
  • multimodal
  • uncensored
    license: apache-2.0
    language:
  • en
  • zh

Qwen3.5-35B-A3B-Uncensored-HauhauCS-Aggressive-GGUF

Full llama.cpp quantization ladder for HauhauCS/Qwen3.5-35B-A3B-Uncensored-HauhauCS-Aggressive. K-quants from Q8_0 through Q4 use standard llama-quantize without an importance matrix. Low-bit Q3_K_L / Q3_K_M / Q3_K_S, Q2_K, and all IQ* types use WikiText-2 importance-matrix calibration (200 chunks) when this workspace contains imatrix.dat.

About the Source Model

This repo is a GGUF quantization ladder for HauhauCS/Qwen3.5-35B-A3B-Uncensored-HauhauCS-Aggressive: an Aggressive uncensored build based on Qwen/Qwen3.5-35B-A3B (MoE, multimodal, long context). Low-bit K-quants (Q3_K*, Q2_K) and IQ-types use an importance matrix when imatrix.dat was produced in this run—same spirit as our compacted Qwen3.5 GGUF ladder.

For refusal behavior, recommended sampling settings, and mmproj vision tensors, follow the HauhauCS model card and Qwen docs. Note: LM Studio may show 256×2.6B in the params column; HauhauCS reports this is a cosmetic metadata quirk.

Complementary files (read this if a quant is missing here)

The HauhauCS weight index already hosts BF16, Q8_0 through Q6_K, several Q4/Q5 variants, IQ4_XS, IQ3_M, IQ2_M, Q3_K_M, etc. This cahlen companion repo (HF names end with -GGUF) is disk-aware: it adds the extra ladder rungs we use on constrained hardware (e.g. Q5_K_S, Q4_K_S, Q3_K_L / Q3_K_S, Q2_K, IQ3_S, IQ3_XXS, IQ2_S, IQ2_XXS, IQ1_M) with the same WikiText-2 / 200-chunk imatrix workflow as cahlen/qwen3.5-35b-a3b-compacted-GGUF. Pull from HauhauCS if you need a size we do not mirror here.

Available Quantizations

Filename Quant Size Notes
Q5_K_S Q5_K_S 23G K-quant
Q4_K_S Q4_K_S 19G K-quant
Q3_K_L Q3_K_L 17G imatrix
Q3_K_S Q3_K_S 15G imatrix
IQ3_S IQ3_S 15G imatrix
IQ3_XXS IQ3_XXS 13G imatrix
Q2_K Q2_K 13G imatrix
IQ2_S IQ2_S 10G imatrix
IQ2_XXS IQ2_XXS 8.9G imatrix
IQ1_M IQ1_M 7.7G imatrix
mmproj-...-f16.gguf mmproj (vision) 858M Pair with any quant above

All filenames are prefixed with Qwen3.5-35B-A3B-Uncensored-HauhauCS-Aggressive-. The BF16 baseline (65G) was used locally for quantization but is not uploaded to save space; grab it from the HauhauCS source repo if needed. "imatrix" rows used WikiText-2 importance-matrix calibration (200 chunks).

Quality (WikiText-2 Perplexity)

Lower is better. First row is the unquantized baseline.

Quant Size Perplexity vs Baseline
BF16 (baseline) 65G 6.4393 —
Q5_K_S 23G 6.4871 +0.7%
Q4_K_S 19G 6.6214 +2.8%
Q3_K_L 17G 6.7204 +4.4%
IQ3_S 15G 6.7631 +5.0%
Q3_K_S 15G 6.9724 +8.3%
IQ3_XXS 13G 7.0490 +9.5%
Q2_K 13G 7.4896 +16.3%
IQ2_S 10G 8.1019 +25.8%
IQ2_XXS 8.9G 9.0738 +40.9%
IQ1_M 7.7G 11.1425 +73.0%

Measured with llama-perplexity on the WikiText-2 test set (580 chunks, context 512). BF16 baseline evaluated on CPU; quantized variants on NVIDIA RTX 5090.

How to Use

With llama.cpp (text)

llama-cli -m Qwen3.5-35B-A3B-Uncensored-HauhauCS-Aggressive-Q4_K_S.gguf --jinja -c 131072 -ngl 99 -p "Hello"

With llama.cpp (vision)

llama-cli -m Qwen3.5-35B-A3B-Uncensored-HauhauCS-Aggressive-Q4_K_S.gguf \
  --mmproj mmproj-Qwen3.5-35B-A3B-Uncensored-HauhauCS-Aggressive-f16.gguf \
  --jinja -c 131072 -ngl 99

With llama-server

llama-server -m Qwen3.5-35B-A3B-Uncensored-HauhauCS-Aggressive-Q4_K_S.gguf --jinja -c 131072 -ngl 99

With Ollama

ollama run hf.co/cahlen/Qwen3.5-35B-A3B-Uncensored-HauhauCS-Aggressive-GGUF:Q4_K_S

With LM Studio

Download any GGUF from the table and load it.

Choosing a Quant

Rough disk size / VRAM guidance (actual usage varies by context length and loader). Quants marked ★ are in this repo; others are on the HauhauCS source repo.

Your VRAM Try Size
24GB+ Q8_0 or Q6_K (HauhauCS) largest
16GB ★ Q5_K_S / ★ Q4_K_S 19–23G
12GB ★ Q3_K_L / ★ IQ3_S 15–17G
8GB ★ IQ3_XXS / ★ Q2_K 13G
6GB ★ IQ2_S / ★ IQ2_XXS 8.9–10G

Quantization Details

  • Quantized by: cahlen
  • Importance matrix: WikiText-2 (wikitext-2-raw-v1, 200 chunks), when generated for this run
  • Tool: llama.cpp @ 59d840209
  • Hardware: NVIDIA RTX 5090 32GB / Intel Core Ultra 9 285K / 188GB RAM

README history 4 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-04-03Add files using upload-large-folder tool01b3d255.6 KB
    Loading...
  2. 2026-04-03Add files using upload-large-folder toolc90701b5.5 KB
    Loading...
  3. 2026-04-03Add files using upload-large-folder tool51d03126.3 KB
    Loading...
  4. 2026-04-03Add files using upload-large-folder toolc7f01375.6 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration