← back to catalog · registered 2026-08-22 13:56

kaineone/Qwen3.5-4B-abliterated

kaineone Qwen 4.7B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/kaineone%2FQwen3.5-4B-abliterated"
Response includes
  • classification m1
  • files 13
  • benchmarks 11 entries
  • hub_downloads_all_time 855
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
855
178 last 30d - stable
Likes
0
Model age
3mo ago
created 2026-06-19
Downloads over time
Now871→from0↑0%
03196399580 on Jun 17871 on Oct 11871 on Oct 9JunJulAugSepOct
Jun 17 → Oct 11 · 56 snapshots · spans 116 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 0.9 UGI
Hazardous 1.2 UGI
Natural Intelligence 13.45 UGI
Political lean -17.3% UGI
Sensitive-Info 11.73 UGI
SocPol 1.5 UGI
UGI 15.32 UGI
Willingness (10) 2.2 UGI
W10-Adherence 1.5 UGI
W10-Direct 3 UGI
Writing 29.68 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 2 formats · 1K downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
apache-2.0
Tags
transformers safetensors qwen3_5 image-text-to-text abliterated uncensored qwen3.5 kaine text-generation conversational base_model:Qwen/Qwen3.5-4B base_model:finetune:Qwen/Qwen3.5-4B

Related

Total size
8.68 GB
Files
13
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-07-10 00:38

Files by quantization

Auxiliary files 13 files 8.70 GB
model.safetensors-00001-of-00002.safetensors 4.96 GB 7b1d5e02 download
model.safetensors-00002-of-00002.safetensors 3.72 GB 49b5d420 download
tokenizer.json 12.2 MB 5f9e4d49 download
vocab.json 6.41 MB 0aa0ce06 download
merges.txt 3.20 MB a494e019 download
model.safetensors.index.json 74.4 KB fddda603 download
tokenizer_config.json 16.3 KB eda48d3e download
chat_template.jinja 7.57 KB a585dec8 download
README.md 5.54 KB 538e6bb8 download
config.json 3.09 KB 557d961b download
.gitattributes 1.53 KB 52373fe2 download
NOTICE 721 B ad29f862 download
preprocessor_config.json 390 B 2ea84a43 download

README current version from Hugging Face


license: apache-2.0
base_model: Qwen/Qwen3.5-4B
tags:

  • abliterated
  • uncensored
  • qwen3.5
  • kaine
    library_name: transformers
    pipeline_tag: text-generation

KAINE · Qwen3.5-4B Abliterated

An abliterated (refusal-direction-removed) variant of
Qwen/Qwen3.5-4B, produced for the
KAINE cognitive-architecture research
project as its language organ.

In KAINE the language model is the organ, not the brain — behavior is meant to
be governed by the cognitive architecture (values, affect, memory, self-model),
not by refusals baked into the base weights. Abliteration removes the refusal gate
so the architecture, not the model's reflexes, decides. This model is published
openly so the research is reproducible: every install resolves the same weights.

Intended use & the KAINE project

This model is the language organ for KAINE,
a composite cognitive architecture in which many modules interact through a global
workspace; the organ supplies language, while values, affect, memory, and a
self-model live in the architecture around it. It is intended as a research
substrate
for that work — not as a general-purpose assistant — and is published
so KAINE installs and independent replications resolve identical weights. Companion
GGUF builds:
kaineone/Qwen3.5-4B-abliterated-GGUF.

What "abliterated" means here (and what it does not)

This is abliteration — subtractive removal of the refusal direction
(Arditi et al. 2024, "Refusal in LLMs is mediated by a single direction"):
W' = W − r̂ r̂ᵀ W. It is not fine-tuning and no preference/instruction
data was trained in. The base model's capabilities and distribution are left
intact; only the refusal direction is orthogonalized out.

Honest scope: abliteration removes the refusal direction — it does not
make the model value-neutral. The base model's pretraining and RLHF priors remain
in the weights. This lifts the model's willingness to respond, not its
underlying tendencies.

Reproducible recipe

  • Base: Qwen/Qwen3.5-4B (note: a vision-language model; abliteration targeted
    the text refusal direction).
  • Tool: jim-plus/llm-abliteration @ ca6e223.
  • Measure: last-token residual-stream mean-difference between 1,139 contrastive
    harmful/harmless prompts (the tool's bundled sets), per layer, 8-bit.
  • Ablate: layers 11–31, banded source directions — layer 17 for 11–22,
    layer 29 for 23–31 (the cleanest mid- and late-network directions); scale = 1.0,
    norm-preserving orthogonalization of the attention output and MLP down-projection
    weights.
  • Tooling note: loading/abliterating Qwen3.5 requires transformers ≥ 5.

Validation

Validated with KAINE's own gates:

  • De-refusal: zero refusal markers on the abliteration probe set (the model no
    longer deflects with "I cannot…" / "I'm not able to…").
  • Capability: matched the vanilla base on the capability probe set (no measured
    regression).

Caveat: these are compact built-in gates — a gross-regression / residual-refusal
check, not a comprehensive benchmark. Treat the validation as "no obvious
breakage," and run your own evaluation for your use case.

Mechanistic verification

Beyond the behavioral gates above, the abliteration is verified mechanistically by
measuring the refusal direction it removes. Using the base model's per-layer
harmful-minus-harmless direction (the same last-token contrast the ablation
targets) over the tool's 1,137 harmful / 640 harmless prompts, we project both the
base and this model onto that direction and report how much of the base's
separation survives:

  • The reduction lands exactly on the two ablated source layers — deepest at
    layer 17 (band 11-22, ~22% retained) and layer 29 (band 23-31, ~13%
    retained), with layers below 11 untouched. This confirms the ablation acted where
    and how this card documents.
  • A distributed harmful/harmless representation persists (~59% retained
    averaged across refusal-carrying layers) — expected, since a banded ablation
    orthogonalizes only the two source directions and refusal is multi-dimensional
    (Wollschläger et al. 2025; Joad et al. 2026).

In short: refusal expression is removed (the model emits no refusals) and the
refusal direction is deeply cut at its target layers, but the underlying
harmful/harmless representation is not erased — abliteration lifts willingness
to respond, it does not make the model unable to tell harmful from harmless.
Forward-pass-only projection on the safetensors weights (activations only; nothing
generated).

Formats

  • This repo: safetensors (transformers / vLLM / fine-tuning).
  • Companion GGUF (Q4_K_M and others) for llama.cpp / Ollama / LM Studio,
    exported with mainline llama.cpp convert_hf_to_gguf.py.

License & attribution

Apache-2.0, inherited from the Qwen/Qwen3.5-4B base. Derivative produced by the
KAINE project (Kaine.One). Refusal-removal method: Arditi et al. 2024; tooling:
jim-plus/llm-abliteration.

Intended use & caution

Built as a research substrate for an architecture that supplies its own value and
safety scaffolding. With refusals removed, this model will attempt most requests —
use it within an appropriate safety framework and applicable law. It is uncensored
by design, not by endorsement of any particular use.

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-07-10Add mechanistic verification (refusal-direction probe)13ff9265.5 KB
    Loading...
  2. 2026-06-19Upload README.md with huggingface_hubdbbbffb4.1 KB
    Loading...
  3. 2026-06-19Upload folder using huggingface_hub402893a3.5 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration