← back to catalog · registered 2026-08-22 13:56

localslm/Qwen3.5-9B-MTP-Uncensored

localslm Qwen 9B GGUF 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/localslm%2FQwen3.5-9B-MTP-Uncensored"
Response includes
  • classification m-uncensored
  • files 6
  • benchmarks 11 entries
  • hub_downloads_all_time 1,475
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
1K
449 last 30d - stable
Likes
1
Model age
2mo ago
created 2026-08-05
Downloads over time
Now1.6K→from591↑168%
5419211.3K1.7K591 on Aug 51.6K on Oct 11AugSepOct
Aug 5 → Oct 11 · 51 snapshots · spans 67 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 1.4 UGI
Hazardous 2.4 UGI
Natural Intelligence 17.62 UGI
Political lean -12.2% UGI
Sensitive-Info 14.65 UGI
SocPol 0.9 UGI
UGI 17.27 UGI
Willingness (10) 2.2 UGI
W10-Adherence 1.5 UGI
W10-Direct 3 UGI
Writing 33.52 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh multilingual
Tags
gguf uncensored qwen3.5 qwen en zh multilingual base_model:Qwen/Qwen3.5-9B base_model:quantized:Qwen/Qwen3.5-9B license:apache-2.0 endpoints_compatible region:us

Related

Total size
21.9 GB
Files
6
Quantizations
2
Registered
2026-08-22 13:56
Last updated on HF
2026-08-08 09:08

Files by quantization

mmproj 1 file 879 MB
mmproj.gguf 879 MB 05f66250 download
Auxiliary files 5 files 21.9 GB
model_q8_0.gguf 9.33 GB 2f5419f8 download
model_q6_k.gguf 7.09 GB a8a28c75 download
model_q4_k_m.gguf 5.48 GB 2d101962 download
README.md 6.26 KB 8f0a9636 download
.gitattributes 1.73 KB eba8c8d9 download

README current version from Hugging Face


license: apache-2.0
base_model:

  • Qwen/Qwen3.5-9B
    tags:
  • uncensored
  • qwen3.5
  • qwen
  • gguf
    language:
  • en
  • zh
  • multilingual

Uncensored Qwen3.5-9B-MTP GGUF

Example Hybrid Setup: CPU (LLM) + GPU (Vision, ~4GB VRAM)

This setup launches a highly optimized hybrid CPU/GPU vision-enabled server using llama.cpp, specifically tuned for maximum performance, minimal memory footprint, and high-quality generation on budget hardware:

  • Fast CPU Text Generation: Leverages Multi-Token Prediction (MTP) speculative decoding. The process is pinned to physical CPU cores, dramatically increasing cache hits and reducing latency.
  • Instant Image Preprocessing (~1 second): Offloads the multimodal projector to even a low-end GPU (4GB VRAM). Even with a massive budget of 4,096 visual tokens, image processing takes under a second on GPU.
  • Massive 49K Context Window: Achieved within standard RAM limits by utilizing 4-bit KV-cache quantization.
  • Advanced Repetition Control: Uses DRY (Don't Repeat Yourself) sampling instead of static penalties to ensure natural, loop-free text generation.
  • WebUI & API: Features an OpenAI-compatible local API/WebUI server on port 8080.
llama-server.exe ^
    -m "D:\LLMs\Qwen3.5-9B-MTP-Uncensored\model.gguf" ^
    --mmproj "D:\LLMs\Qwen3.5-9B-MTP-Uncensored\mmproj.gguf" ^
    --seed -1 ^
    -np 1 ^
    --image-min-tokens 4096 ^
    --image-max-tokens 4096 ^
    -t 6 --threads-batch 6 ^
    -Cr 0-5 --cpu-strict 1 ^
    --prio 2 --poll 100 ^
    -b 4224 -ub 2048 --mtmd-batch-max-tokens 2048 ^
    -ctk q4_0 -ctv q4_0 ^
    -ctkd q4_0 -ctvd q4_0 ^
    --cache-reuse 256 ^
    --cache-prompt ^
    --slot-prompt-similarity 0.0 ^
    --swa-full ^
    -lm mlock ^
    -ngl 0 ^
    --mmproj-offload ^
    -fa on ^
    -c 49152 ^
    --temp 0.3 ^
    --min-p 0.05 ^
    --top-p 0.95 ^
    --top-k 40 ^
    --dry-multiplier 0.8 ^
    --dry-base 1.75 ^
    --dry-allowed-length 2 ^
    --dry-penalty-last-n 256 ^
    --presence-penalty 0.0 ^
    --repeat-penalty 1.0 ^
    --reasoning off ^
    --jinja ^
    --spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-p-min 0.20 ^
    --metrics^
    --host 127.0.0.1 --port 8080

This command launches llama-server.exe (the server binary provided by llama.cpp) to host a local REST API / web interface for a vision-enabled Qwen3.5-9B on Windows:

1. Model & Multimodal Files

  • -m & --mmproj
    Paths to the main language model and the multimodal projector (vision encoder).
  • --image-min-tokens 4096 & --image-max-tokens 4096
    Forces image processing to use a fixed allocation of 4,096 visual tokens per image input, ensuring maximum visual detail and OCR precision.
  • --mmproj-offload
    Crucial optimization: Offloads the heavy vision encoder to the GPU. Even on a cheap GPU, processing 4,096 visual tokens completes in < 1 second.

2. CPU Threading & Hardware Optimization

  • -t 6 & --threads-batch 6
    Allocates 6 CPU threads for token generation and batch processing.
  • -Cr 0-5 --cpu-strict 1
    Pins the process strictly to physical CPU cores 0 through 5. This prevents the OS from migrating threads across cores, vastly improving CPU cache hit rates and generation speed.
  • --prio 2 & --poll 100
    Sets the server to high process priority (2) and lowers the polling interval (100 ms) for faster API responsiveness.
  • -lm mlock
    Locks model memory directly in physical RAM, preventing the OS from swapping it to the disk/pagefile (replaces older --no-mmap arguments).
  • -ngl 0
    Forces the main LLM to run entirely on the CPU.
  • -fa on
    Enables FlashAttention to significantly reduce RAM usage during context processing.

3. Batching, Context, & KV-Cache (Memory Savings)

  • -c 49152 & -np 1
    Sets a massive 49,152 token context window, optimized for a single user/slot (-np 1).
  • -b 4224 & -ub 2048
    Sets the prompt batch size to 4224 and micro-batch size to 2048, balancing RAM spikes and CPU load.
  • --mtmd-batch-max-tokens 2048
    Limits multimodal chunking to 2048 tokens, preventing Out-Of-Memory (OOM) errors during heavy image ingestion.
  • -ctk q4_0 -ctv q4_0 & -ctkd q4_0 -ctvd q4_0
    Applies 4-bit quantization to the KV-cache (for both main and draft models). This is what makes the 49K context window possible on standard consumer RAM setups without degrading output quality.
  • --cache-prompt, --cache-reuse 256, --slot-prompt-similarity 0.0
    Enables aggressive prompt caching. This skips reprocessing the system prompt and chat history, drastically speeding up multi-turn conversations.
  • --swa-full
    Enables full Sliding Window Attention, further optimizing memory footprint.

4. Sampling & DRY Repetition Control

  • --temp 0.3, --top-p 0.95, --top-k 40, --min-p 0.05
    A precise, low-temperature sampling setup that relies heavily on Min-P to truncate low-probability garbage tokens, leading to highly coherent and logical responses.
  • --seed -1
    Ensures a random seed for generation on every prompt.
  • DRY (Don't Repeat Yourself) Sampler (Replaces standard repetition penalties):
    • --dry-multiplier 0.8, --dry-base 1.75, --dry-allowed-length 2, --dry-penalty-last-n 256
      Instead of statically penalizing words (which breaks formatting/code), DRY dynamically detects sequences of repeating tokens and exponentially penalizes them, effectively curing model loops.
  • --presence-penalty 0.0 & --repeat-penalty 1.0
    Standard penalties are disabled to allow the DRY sampler to work without interference.

5. Speculative Decoding & Template Settings

  • --spec-type draft-mtp
    Enables Multi-Token Prediction (MTP).
  • --spec-draft-n-max 2 & --spec-draft-p-min 0.20
    Generates up to 2 draft tokens per step, but only if the model is at least 20% confident (0.20). This prevents wasting CPU cycles on bad guesses, resulting in a smoother, faster generation rate.
  • --reasoning off
    Natively disables the DeepSeek/Qwen thinking/reasoning steps, guaranteeing immediate final-answer generation and saving compute time.
  • --jinja & --metrics
    Enables Jinja2 parsing for complex chat templates and exposes an endpoint for performance monitoring.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-08upload files2d76f186.3 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration