← back to catalog · registered 2026-08-22 13:56

AlexWortega/qwen35-4b-soyuz-abliterated-v3-multi

AlexWortega Qwen 4.5B multimodal
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/AlexWortega%2Fqwen35-4b-soyuz-abliterated-v3-multi"
Response includes
  • classification m1
  • files 9
  • benchmarks 11 entries
  • hub_downloads_all_time 53
  • author_summary 6 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
53
16 last 30d - stable
Likes
0
Model age
4mo ago
created 2026-05-24

Training datasets

Corpora the author lists in the model card. Datasets tracked in our /datasets catalog carry a category badge linking to the workflow stage. Others open on Hugging Face.

Downloads over time
Now55→from22↑150%
2033465822 on Jun 1055 on Oct 1155 on Oct 4JunJulAugSepOct
Jun 10 → Oct 11 · 57 snapshots · spans 123 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 0.9 UGI
Hazardous 1.2 UGI
Natural Intelligence 13.45 UGI
Political lean -17.3% UGI
Sensitive-Info 11.73 UGI
SocPol 1.5 UGI
UGI 15.32 UGI
Willingness (10) 2.2 UGI
W10-Adherence 1.5 UGI
W10-Direct 3 UGI
Writing 29.68 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en
Tags
transformers safetensors qwen3_5 image-text-to-text agent abliteration orthogonalization weight-ortho soyuz qwen3.5 conversational en

Related

Total size
8.46 GB
Files
9
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-05-24 20:11

Files by quantization

Auxiliary files 9 files 8.47 GB
model.safetensors 8.46 GB 9ce32733 download
tokenizer.json 19.1 MB 06b95093 download
chat_template.jinja 7.57 KB a585dec8 download
config.json 2.76 KB 4829f7b9 download
README.md 2.64 KB 1bef59ea download
.gitattributes 1.53 KB 52373fe2 download
processor_config.json 1.16 KB 33818c7f download
tokenizer_config.json 1.14 KB b4acebe0 download
generation_config.json 115 B d9dfd84b download

README current version from Hugging Face


library_name: transformers
base_model: Qwen/Qwen3.5-4B
license: apache-2.0
language: [en]
tags:

  • agent
  • abliteration
  • orthogonalization
  • weight-ortho
  • soyuz
  • qwen3.5
    datasets:
  • AlexWortega/Soyuz-sft

Qwen3.5-4B Soyuz — Abliterated (v3)

Weight-orthogonalised version of AlexWortega/qwen35-4b-soyuz-merged. Removes the residual-stream "fail-mode" component identified from the model's own pass-vs-fail trajectory contrasts.

Method multi-layer per-layer ortho (L8-24), strength=0.5
tbench-2 (17-task) 2/17
HermesAgent-20 6 / 20
HA20 passes HA-01, HA-02, HA-03, HA-06, HA-09, HA-11

Usage with sglang

python -m sglang.launch_server \
    --model-path AlexWortega/qwen35-4b-soyuz-abliterated-v3-multi \
    --dtype bfloat16 --trust-remote-code \
    --tool-call-parser hermes \
    --chat-template hermes_qwen.jinja

(hermes parser is needed for the <tool_call>{...}</tool_call> → OpenAI tool_calls conversion — without it agent benches see zero tool calls.)

Abliteration recipe

  1. Build pass-vs-fail contrast: 60 PASS trajectories (reward=1.0) + 60 cleaned FAIL trajectories from soyuz's own evals (claw-eval, tbench-2, MMLU-Pi-agent). Fail trajectories filtered by Gemini-3-flash to keep only CLEAN_FAIL labels (235 of 246 negatives).
  2. Capture last-token residual activations per layer over the rendered contrast (text-only Qwen3_5ForCausalLM).
  3. Compute per-layer direction = mean(refuse) - mean(comply), normalise; pick best layer via AUC.
  4. Orthogonalise model weights (embed rows + every layer's o_proj.weight and down_proj.weight columns) against the direction, optionally blended with strength α: W ← W − α · (W − W_orth).
  5. Wrap text-only weights into the multimodal Qwen3_5ForConditionalGeneration arch so sglang can serve them (vision tower preserved from base; only language_model.* weights are abliterated).

Repos

Variant tbench-17 HA20 Card
baseline qwen35-4b-soyuz (LoRA) 5/17 4/20 link
qwen35-4b-soyuz-abliterated-v2 (single-L, s=0.5) 3/17 8/20 link
qwen35-4b-soyuz-abliterated-v3-multi (per-layer, s=0.5) 2/17 6/20 link

v2 = highest HA20 (2× baseline). v3 picks up disjoint HA20 tasks (HA-01/02 memory-specific) that v2 misses.

W&B + raw eval logs: https://wandb.ai/alexwortega/vae-llm-agents (training base).

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-05-24Initial abliterated v3628d6ab2.6 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration