license: apache-2.0
base_model: Qwen/Qwen3.5-4B
tags:
- abliterated
- uncensored
- qwen3.5
- gguf
- kaine
pipeline_tag: text-generation
KAINE · Qwen3.5-4B Abliterated — GGUF
GGUF builds of the KAINE abliterated Qwen3.5-4B language organ, forllama.cpp / Ollama / LM Studio. The full method, recipe, validation, and the
honest-scope note are in the
safetensors repo's model card.
Intended use & the KAINE project
This is the language organ for KAINE, a
composite cognitive architecture in which behavior is governed by the architecture
(values, affect, memory, self-model) rather than by refusals baked into the base
weights. Published as a research substrate — not a general-purpose assistant —
so KAINE installs and independent replications resolve identical weights.
Files
| File | Quant | Size | Notes |
|---|---|---|---|
KAINE-Qwen3.5-4B-abliterated.Q4_K_M.gguf |
Q4_K_M | ~2.6 GB | recommended default; verified loads + serves in llama-server |
(More quants — Q5_K_M, Q6_K, Q8_0 — can be added; regenerate from the safetensors
with mainline convert_hf_to_gguf.py + llama-quantize.)
Run
llama.cpp:
llama-server -m KAINE-Qwen3.5-4B-abliterated.Q4_K_M.gguf -ngl 99 --jinja -c 4096
Chain-of-thought suppression (the model is a voice, not a reasoner) via the
OpenAI-compatible request field chat_template_kwargs: {"enable_thinking": false},
or the server flag --reasoning-budget 0.
Ollama: ollama create kaine-qwen3.5-4b-abliterated -f Modelfile (see Modelfile).
LM Studio: search for this repo in the model browser (it indexes HuggingFace GGUFs).
Provenance
Exported with mainline llama.cpp convert_hf_to_gguf.py (so it loads in current
llama.cpp / Studio). Apache-2.0, derivative of Qwen/Qwen3.5-4B. See the
safetensors repo for full attribution.
Mechanistic verification
Beyond the behavioral gates above, the abliteration is verified mechanistically by
measuring the refusal direction it removes. Using the base model's per-layer
harmful-minus-harmless direction (the same last-token contrast the ablation
targets) over the tool's 1,137 harmful / 640 harmless prompts, we project both the
base and this model onto that direction and report how much of the base's
separation survives:
- The reduction lands exactly on the two ablated source layers — deepest at
layer 17 (band 11-22, ~22% retained) and layer 29 (band 23-31, ~13%
retained), with layers below 11 untouched. This confirms the ablation acted where
and how this card documents. - A distributed harmful/harmless representation persists (~59% retained
averaged across refusal-carrying layers) — expected, since a banded ablation
orthogonalizes only the two source directions and refusal is multi-dimensional
(Wollschläger et al. 2025; Joad et al. 2026).
In short: refusal expression is removed (the model emits no refusals) and the
refusal direction is deeply cut at its target layers, but the underlying
harmful/harmless representation is not erased — abliteration lifts willingness
to respond, it does not make the model unable to tell harmful from harmless.
Forward-pass-only projection on the safetensors weights (activations only; nothing
generated).
Measured on the source safetensors model (kaineone/Qwen3.5-4B-abliterated); this GGUF is a quantized export of those weights.