← back to catalog · registered 2026-08-22 13:56

AlexWortega/qwen3.5-4b-abliterated-agent-20260515

AlexWortega Qwen 4B GGUF 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/AlexWortega%2Fqwen3.5-4b-abliterated-agent-20260515"
Response includes
  • classification m8
  • files 11
  • benchmarks 11 entries
  • hub_downloads_all_time 607
  • author_summary 6 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
607
87 last 30d - stable
Likes
0
Model age
4mo ago
created 2026-05-15
Downloads over time
Now617→from0↑0%
02264526790 on May 13617 on Oct 11617 on Oct 10MayJunJulAugSepOct
May 13 → Oct 11 · 61 snapshots · spans 151 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 0.9 UGI
Hazardous 1.2 UGI
Natural Intelligence 13.45 UGI
Political lean -17.3% UGI
Sensitive-Info 11.73 UGI
SocPol 1.5 UGI
UGI 15.32 UGI
Willingness (10) 2.2 UGI
W10-Adherence 1.5 UGI
W10-Direct 3 UGI
Writing 29.68 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
transformers safetensors gguf ml-intern activation-steering abliteration agent terminal-bench weight-orthogonalization base_model:Qwen/Qwen3.5-4B base_model:quantized:Qwen/Qwen3.5-4B license:apache-2.0

Related

Total size
0 B
Files
11
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-05-15 14:46

Files by quantization

Auxiliary files 11 files 77.1 KB
contrast_comply.jsonl 40.0 KB e93e600f download
RESULTS.md 5.78 KB 7dca4de3 download
contrast_refuse.jsonl 5.44 KB ebcc706a download
README.md 5.23 KB 4f83efb7 download
PLAN.md 4.08 KB b6afbdb0 download
RESEARCH.md 3.78 KB 7c001819 download
VERIFY.md 3.54 KB 0ce6fa8a download
TASK.md 3.43 KB 13e81f2a download
RESULTS_APPLES.md 3.23 KB 47f10d15 download
.gitattributes 1.81 KB 209866c6 download
PUBLISHED.md 871 B c20b68ff download

README current version from Hugging Face


library_name: transformers
base_model: Qwen/Qwen3.5-4B
tags:

  • ml-intern
  • activation-steering
  • abliteration
  • agent
  • terminal-bench
  • weight-orthogonalization
    license: apache-2.0

Qwen3.5-4B Abliterated for Agent Use (2026-05-15)

Refusal-direction weight-orthogonalization applied to base
Qwen/Qwen3.5-4B following the
NousResearch/llm-abliteration
/ mlabonne abliteration blog
recipe. The "I cannot execute commands" axis was extracted by contrasting
50 bare shell-execution prompts vs 50 same prompts with full agent system
prompt; this direction was orthogonalized from embed_tokens.weight plus
every block's o_proj/out_proj + mlp.down_proj weights (64 projections + 1
embedding modified).

Results — apples-to-apples on full terminal-bench-2 (89 tasks, N=1, sglang)

model passes / 89 rate
Qwen3.5-4B base (sglang, patched parser) 6 6.7 %
abliterated (this repo) 7 7.9 %
SFT LoRA reference (historic) 3 3.4 %

Fisher's exact 7/89 vs 6/89: p = 0.50 — null on aggregate.

But the per-task pattern is informative:

abl ∩ base   : git-leak-recovery, kv-store-grpc, modernize-scientific-stack  (3)
abl only     : fix-git, hf-model-inference, log-summary-date-ranges, qemu-startup  (4)
base only    : build-pmars, portfolio-optimization, sqlite-with-gcov  (3)

Abliteration redistributes which tasks the model solves rather than lifting
the count. Net +1 well within noise.

Smoke test (qualitative)

Without an agent system prompt, base Qwen3.5-4B replies:

"I'm an AI assistant, I cannot access local filesystem."

Abliterated replies:

"Let me use ls to list the contents of /tmp."

The refusal pattern is gone. But this doesn't translate to a larger task-completion
delta — see results.

How to use

import torch
from transformers import AutoTokenizer, AutoModelForImageTextToText

# Drop-in replacement for Qwen3.5-4B
model = AutoModelForImageTextToText.from_pretrained(
    'AlexWortega/qwen3.5-4b-abliterated-agent-20260515',
    dtype=torch.bfloat16, device_map={'':0})
tokenizer = AutoTokenizer.from_pretrained('AlexWortega/qwen3.5-4b-abliterated-agent-20260515')
# Architecture is unchanged — sglang loads it as a regular Qwen3.5 checkpoint.

GGUF for CPU inference

Pre-quantized GGUF files are bundled under gguf/, benchmarked on AMD EPYC 7402P (24-core Zen 2):

Quant Size tg (decode) pp (prefill) use
qwen3.5-4b-abl.Q4_0.gguf 2.4 GB 20.8 t/s @ 16t 115 t/s @ 16t max throughput
qwen3.5-4b-abl.Q4_K_M.gguf 2.6 GB 19.7 t/s @ 16t 136 t/s @ 24t best balance
qwen3.5-4b-abl.Q5_K_M.gguf 2.9 GB 17.3 t/s @ 16t 91 t/s @ 24t better quality
qwen3.5-4b-abl.Q8_0.gguf 4.2 GB 15.5 t/s @ 16t 110 t/s @ 24t near-lossless

Full benchmark including IQ4_XS, Q6_K, F16 in gguf/BENCH.md.

Quick CPU usage

~/llama.cpp/build/bin/llama-cli -m qwen3.5-4b-abl.Q4_0.gguf -t 16 \
    -p "Your prompt" -n 128 -no-cnv

Use -dev none if your build has CUDA support but you want pure CPU.

What's in this repo

  • abliterated_model/ — 8.5 GB safetensors + tokenizer (drop-in for base Qwen3.5-4B)
  • vectors/refusal_dir.pt — direction tensor at L=22 with metadata
  • vectors/refusal_ranking.csv — per-layer AUC (all layers 1.000)
  • contrast_refuse.jsonl, contrast_comply.jsonl — 50+50 contrast prompts
  • scripts/ — build_contrast.py, capture_refusal.py, compute_refusal_dir.py,
    abliterate.py, serve_abliterated.py, full_bench_parallel.sh
  • results/abliterated_full/ — full per-task traces + rewards (89 tasks)
  • results/base_matched_infra/ — same for BASE through identical infra (control)
  • RESULTS.md, RESULTS_APPLES.md, VERIFY.md — full report bundle

How it was made (quick recipe)

  1. Build 50 refusal prompts (bare shell requests) and 50 compliance prompts
    (same requests with agent system prompt). scripts/build_contrast.py.
  2. Forward base Qwen3.5-4B on each prompt, capture residual at last token of
    each. scripts/capture_refusal.py.
  3. Compute dir_L = normalize(μ_refuse − μ_comply) per layer. AUC=1.0 on
    every layer. Pick L=22 (mid-depth, matches v2 framing).
    scripts/compute_refusal_dir.py.
  4. Orthogonalize: W_new = W − r r^T W for embed_tokens.weight (rows) and
    every block's output projections (columns). scripts/abliterate.py.
  5. Smoke test (scripts/smoke_abliterated.py) — confirm coherent generation
    and refusal removal on bare prompts.
  6. Eval via docker sweep against sglang serving this model.
    scripts/full_bench_parallel.sh.

Caveats

  • N=1 per task. 7/89 vs 6/89 is suggestive at best; per-task pattern is the
    load-bearing observation.
  • The terminus parser was patched mid-experiment to accept split-JSON output
    ({"analysis":...} + {"command":...} on separate lines); this patch lifted
    the base from 3/89 → 6/89, more than abliteration itself contributed.
  • Same direction applied uniformly to all layers (mlabonne convention).
    Per-layer tuning may yield further gains.
  • Modifies output projections + embeddings only; does NOT touch Q/K/V.

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-05-15add GGUF CPU inference section with benchmarksacb6b205.2 KB
    Loading...
  2. 2026-05-15model card with honest final results6e1620b4.4 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration