← back to catalog · registered 2026-08-22 13:56

mmp2055/Qwen3.5-9B-Claude-4.6-OS-AV-H-UNCENSORED-THINK-D_AU-Q4_K_S-imat-bughunter

mmp2055 Qwen 9B GGUF multimodal second-order 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/mmp2055%2FQwen3.5-9B-Claude-4.6-OS-AV-H-UNCENSORED-THINK-D_AU-Q4_K_S-imat-bughunter"
Response includes
  • classification m-uncensored
  • files 5
  • hub_downloads_all_time 459
  • author_summary 6 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
459
185 last 30d - stable
Likes
0
Model age
2mo ago
created 2026-08-04
Downloads over time
Now532→from91↑485%
6923840757691 on Aug 5532 on Oct 11AugSepOct
Aug 5 → Oct 11 · 50 snapshots · spans 67 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
es en
Quantizations
Q4_K
Tags
gguf qwen3.5 vision security bug-bounty lora quantized image-text-to-text es en base_model:DavidAU/Qwen3.5-9B-Claude-4.6-OS-Auto-Variable-HERETIC-UNCENSORED-THINKING base_model:adapter:DavidAU/Qwen3.5-9B-Claude-4.6-OS-Auto-Variable-HERETIC-UNCENSORED-THINKING

Related

Total size
6.11 GB
Files
5
Quantizations
3
Registered
2026-08-22 13:56
Last updated on HF
2026-08-05 10:28

Files by quantization

Q4_K 1 file 6.11 GB
Qwen3.5-9B-Claude-4.6-OS-AV-H-UNCENSORED-THINK-D_AU-Q4_K_S-imat-bughunter.gguf 6.11 GB 30456583 download
F32 1 file 1.70 GB
mmproj-F32.gguf 1.70 GB a1cd5c16 download
Auxiliary files 3 files 101 KB
demo-bughunter.mp4 92.2 KB 4a3396cb download
README.md 6.77 KB 4363b806 download
.gitattributes 1.65 KB 8b8b556c download

README current version from Hugging Face


license: apache-2.0
base_model: DavidAU/Qwen3.5-9B-Claude-4.6-OS-Auto-Variable-HERETIC-UNCENSORED-THINKING
language:

  • es
  • en
    pipeline_tag: image-text-to-text
    tags:
  • gguf
  • qwen3.5
  • vision
  • security
  • bug-bounty
  • lora
  • quantized
    quantized_by: mmp2055

Qwen3.5-9B Bughunter (Q4_K_S, imatrix)

A LoRA fine-tune of
DavidAU/Qwen3.5-9B-Claude-4.6-OS-Auto-Variable-HERETIC-UNCENSORED-THINKING,
specialised on bug bounty methodology. Answers in Spanish; payloads, commands
and report templates in English.

Vision and reasoning (<think>) are preserved — both verified empirically
against the base model.

Demo — it asks the right questions before touching anything

▶ demo-bughunter.mp4 — 13 s

A cold "Hola". No system prompt priming, no setup. Watch what comes back:

¡Hola!

Listo para operar.

Para afinar la estrategia, necesito: target base URL(s), scope
(blacklist/whitelist)
, allowed classes (web/mobile/API/codesign), y
cualquier rate-limit o rule que te hayan dado.

It doesn't start guessing or dumping generic advice. It asks for the four things
that actually decide how an engagement is run — the scope boundaries, what asset
classes are in play, and the rate limits you were given.

93 tokens in 0.92 s (17.8 tok/s) on a local Apple-silicon Mac, thinking pass
included (2.29 s). Note the Think and Vision badges lit at the bottom: chain
of thought and image input are both live in the same session.

⚠️ Read this before using it

This model does NOT enforce operational safety rules. It was trained on
examples that included refusals (never validate a leaked API key, never use
credentials found in a finding, keep PoCs non-destructive), and that part of
the training did not take.
Asked to run sts:GetCallerIdentity with an
AKIA key found on a target, it complies and even rationalises it as
"legitimate recon" in its reasoning block.

Treat it as a documentation and drafting assistant, never as an authority
on what is safe or legal to do against a target. You are responsible for scope,
authorisation and conduct.

The throttling flags it emits are also unreliable (it produced
-rl 10/3s -rate-limit 25 -c 4 for nuclei — not valid syntax, and the
concurrency contradicts the rate limit). Verify every command before running it.

Files

File Size Required
…-bughunter.gguf 6.1 GiB yes — the language model
mmproj-F32.gguf 1.7 GiB yes for vision — the vision encoder + projector

Both files must sit in the same folder. Without the mmproj, loaders treat
it as text-only and vision is silently lost.

Usage

LM Studio

Place both files under
~/.lmstudio/models/<you>/Qwen3.5-9B-…-bughunter/ and reload the model list.
LM Studio should report vision support.

Recommended settings — these matter:

Setting Value Why
System prompt short (< 500 chars) Trained with one long system prompt in 70% of examples; a long one makes it echo the prompt inside <think> instead of answering
temperature 0.1 Keeps commands and report wording precise
repeat_penalty 1.1 Without it, it can fall into recitation loops

llama.cpp

# text
llama-cli -m Qwen3.5-9B-…-bughunter.gguf -st --temp 0.1 --repeat-penalty 1.1 \
  -sys "Eres un investigador profesional de bug bounty. Respondes en español." \
  -p "En el target veo un parámetro image_url. ¿Qué pruebo?"

# vision
llama-mtmd-cli -m Qwen3.5-9B-…-bughunter.gguf --mmproj mmproj-F32.gguf \
  --image burp-capture.png -p "Transcribe la petición HTTP y señala lo sospechoso."

OpenAI-compatible API

Works with tools / function calling. Pass temperature: 0.1 and
repeat_penalty: 1.1 explicitly — UI presets do not apply over the API.

What it is actually good at

  • Reading screenshots — Burp panels, API responses, config screens. Vision is
    intact: on a control image it read every element, including a needle header
    only obtainable by actually reading the pixels.
  • Orienting a surface towards the right vulnerability class.
  • Report structure per platform (HackerOne CVSS, Bugcrowd VRT, Intigriti GDPR).
  • Reasoning out loud before answering (<think> preserved).

The specialisation is shallow: it has the flavour and vocabulary of the
domain, not reliable command-level or policy-level accuracy. See the warning above.

Training

Method LoRA on the top 8 of 32 layers, rank 16, scale 32
Trainable params 10.8 M (0.12 %)
Data 538 single-turn examples, Spanish, each with a genuine <think> block
Sources Distilled from public material: HackerOne disclosed reports, PortSwigger Web Security Academy labs, and OWASP-style technique notes
Stopped at ~0.84 epochs — train and validation loss crossed and diverged there
Framework MLX on Apple Silicon; fused into the original BF16 weights

Known dataset flaws (documented so nobody repeats them): the reasoning blocks
in the ~470 script-generated examples share a near-identical structure, so the
model learned the pattern rather than the content; only 22 examples taught
conduct; and there were no multi-turn examples at all.

Quantisation

Q4_K_S with an importance matrix computed over 400 chunks of the domain corpus.

The per-tensor recipe mirrors the base model's: output.weight and all 48
ssm_alpha / ssm_beta tensors are kept at BF16. That is deliberate — with a
248,320-token vocabulary the output tensor is the most quantisation-sensitive in
the model, and the SSM gates govern the recurrent dynamics of 24 of the 32 layers,
where quantisation error accumulates along the sequence rather than averaging out.

The GGUF also carries the base model's custom chat template (by DavidAU /
Nightmedia).

Credits

  • Base: DavidAU — merge, "MAX" quantisation
    recipe and the custom chat template.
  • Architecture: Qwen3.5 (hybrid attention + SSM, vision-language).
  • Licence: Apache 2.0, inherited from the base model.
  • Recommended agent for using this model:
    https://xagentai.net/xagentai-net-coding-agent/

Intended use

Assisting authorised security research and learning: drafting reports, recalling
technique patterns, reading tool screenshots. Only against systems you are
authorised to test.
It is not a substitute for a programme's scope and rules,
and — as stated above — it will not stop you from doing something you shouldn't.

README history 7 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-05Upload README.md with huggingface_hubac41b656.8 KB
    Loading...
  2. 2026-08-05Upload README.md with huggingface_hubd07881a5.7 KB
    Loading...
  3. 2026-08-04Upload README.md with huggingface_hub39e40395.9 KB
    Loading...
  4. 2026-08-04Upload README.md with huggingface_hub34ac9156 KB
    Loading...
  5. 2026-08-04Upload README.md with huggingface_hubfc48d685.9 KB
    Loading...
  6. 2026-08-04Upload README.md with huggingface_hub744d5635.8 KB
    Loading...
  7. 2026-08-04Upload README.md with huggingface_hube63be755.7 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration