← back to catalog · registered 2026-08-22 13:56

nDimensional/Qwen3.5-35B-A3B-Uncensored-FP8_BLOCK

nDimensional Qwen 33B MoE multimodal
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/nDimensional%2FQwen3.5-35B-A3B-Uncensored-FP8_BLOCK"
Response includes
  • classification m-uncensored
  • files 12
  • benchmarks 11 entries
  • hub_downloads_all_time 1,612
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
2K
72 last 30d - cooling
Likes
1
Model age
6mo ago
created 2026-04-02
Downloads over time
Now1.6K→from436↑276%
3768371.3K1.8K436 on Apr 151.6K on Oct 111.6K on Oct 6AprMayJunJulAugSepOct
Apr 15 → Oct 11 · 65 snapshots · spans 179 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 0.8 UGI
Hazardous 4.1 UGI
Natural Intelligence 24.97 UGI
Political lean -20.7% UGI
Sensitive-Info 20.98 UGI
SocPol 2.1 UGI
UGI 23.15 UGI
Willingness (10) 2.8 UGI
W10-Adherence 1.5 UGI
W10-Direct 4 UGI
Writing 37 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh
Tags
transformers safetensors qwen3_5_moe image-text-to-text qwen3_5 qwen35 moe vllm qwen conversational quantized compressed

Related

Total size
35.0 GB
Files
12
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-04-02 23:41

Files by quantization

Auxiliary files 12 files 35.1 GB
model.safetensors 35.0 GB 7853e0a3 download
tokenizer.json 19.1 MB 87a7830d download
vocab.json 6.41 MB 0aa0ce06 download
config.json 22.6 KB 797d3b23 download
chat_template.jinja 7.57 KB a585dec8 download
README.md 2.87 KB e216cab6 download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.16 KB 33818c7f download
tokenizer_config.json 1.11 KB a068e246 download
video_preprocessor_config.json 385 B 3ba673a5 download
recipe.yaml 343 B cdfddaad download
generation_config.json 213 B a0b91151 download

README current version from Hugging Face


license: apache-2.0
language:

  • en
  • zh
    base_model:
  • Qwen/Qwen3.5-35B-A3B
  • HauhauCS/Qwen3.5-35B-A3B-Uncensored-HauhauCS-Aggressive
    pipeline_tag: image-text-to-text
    library_name: transformers
    tags:
  • qwen3_5
  • qwen35
  • moe
  • vllm
  • qwen
  • conversational
  • quantized
  • compressed
  • compressed-tensors
  • llm-compressor
  • fp8
  • qwen3_5_moe
  • uncensored
  • unfiltered
  • ablation
  • optimized

Qwen3.5-35B-A3B Uncensored (FP8_BLOCK)

A safetensors conversion and quantization of HauhauCS/Qwen3.5-35B-A3B-Uncensored-HauhauCS-Aggressive (GGUF).

Model Details

Architecture Qwen3.5 MoE hybrid attention (30 GDN + 10 full standard attention layers)
Parameters 35B-A3B
Base model Qwen/Qwen3.5-35B-A3B
Source GGUF HauhauCS/Qwen3.5-35B-A3B-Uncensored-HauhauCS-Aggressive
Format BF16 (Mixed/Compressed)
Quantization FP8_BLOCK applied to Linear transformer layers.
Stripped layers Multi-Token Prediction (MTP) due to original HF -> GGUF conversion.
Conversion type Lossless GGUF to safetensors conversion + merge with base model vision layers + Block-wise quantization
Unquantized weights Coming Soon

Conversion Details

Converted using coming soon, which reverses transforms applied during HF -> GGUF conversion.

The vision encoder weights are copied directly from the official Qwen/Qwen3.5-35B-A3B base model, after confirming the vision encoder (mmproj) was not modified in the source GGUF.

Next, the linear weights of the transformer blocks were quantized to F8_E4M3 using llm-compressor.

Test Inference Details

  • 1x A100 (80GB)
  • Python 3.12
  • vllm & transformers version:
    • transformers 5.5.0
    • vllm nightly (latest commit tested: 7b743ba)
  • vLLM online serve flags:
    • --quantization compressed-tensors
    • --max-model-len 16384
    • --gpu-memory-utilization 0.9140 with VLLM_MEMORY_PROFILER_ESTIMATE_CUDAGRAPHS=1 environmental variable
    • --limit-mm-per-prompt.image 4
    • --enable-prefix-caching
    • --enable-expert-parallel
    • --reasoning-parser qwen3
    • --default-chat-template-kwargs {"enable_thinking": false} disabled thinking/reasoning for vllm>=0.18.1
    • Note: Used for batch image captioning tests.

Credits

README history 6 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-04-02Minor grammar fixes7ac87872.9 KB
    Loading...
  2. 2026-04-02Update README.md90d78b22.9 KB
    Loading...
  3. 2026-04-02Update README.mda8310de2.9 KB
    Loading...
  4. 2026-04-02Update README.md99cb7552.4 KB
    Loading...
  5. 2026-04-02Update README.md8d3ce7e2.4 KB
    Loading...
  6. 2026-04-02initial commit7e9f63228 B
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration