← back to catalog · registered 2026-08-22 13:56

DuoNeural/Phi-4-Mini-Abliterated

DuoNeural Phi 3.8B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/DuoNeural%2FPhi-4-Mini-Abliterated"
Response includes
  • classification m1
  • files 14
  • hub_downloads_all_time 808
  • author_summary 45 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
808
211 last 30d - stable
Likes
1
Descendants
3
in 3 direct forks
Model age
5mo ago
created 2026-04-30
Downloads over time
Now878→from30↑2,827%
032164296330 on Apr 29878 on Oct 11AprMayJunJulAugSepOct
Apr 29 → Oct 11 · 63 snapshots · spans 165 days

Genealogy 3 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
mit
Languages
en
Tags
safetensors phi3 abliteration phi4 microsoft DuoNeural layer-crystallization p34 text-generation conversational custom_code en

Related

Total size
7.15 GB
Files
14
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-06-04 17:14

Files by quantization

Auxiliary files 14 files 7.17 GB
model-00001-of-00002.safetensors 4.57 GB 65426e54 download
model-00002-of-00002.safetensors 2.58 GB 3fb647c3 download
tokenizer.json 14.8 MB 382cc235 download
vocab.json 3.73 MB ea953a43 download
merges.txt 2.31 MB dcecc452 download
model.safetensors.index.json 15.9 KB d87cc365 download
README.md 4.35 KB c281b943 download
tokenizer_config.json 2.89 KB 15ddb38c download
config.json 2.50 KB ed9c6daf download
.gitattributes 1.53 KB 52373fe2 download
special_tokens_map.json 587 B 156262f7 download
chat_template.jinja 423 B a9c00dd9 download
added_tokens.json 249 B af52cde6 download
generation_config.json 169 B 05cbda4c download

README current version from Hugging Face


license: mit
base_model: microsoft/Phi-4-mini-instruct
language:

  • en
    tags:
  • abliteration
  • phi4
  • microsoft
  • DuoNeural
  • layer-crystallization
  • p34
    pipeline_tag: text-generation

Phi-4-Mini-Instruct Abliterated (L16 direction)

DuoNeural | 2026-06-04 — Canonical documented run

This is the correct abliteration of Phi-4-Mini. Earlier attempts using standard final-layer direction extraction all failed due to a layer crystallization mismatch. See findings below.

Abliterated version of microsoft/Phi-4-mini-instruct using the correct refusal direction.


Results

Metric Value
Pre-abliteration compliance 1/5
Post-abliteration compliance 5/5
KL divergence (Heretic v2.0, BF16→BF16) 0.0135 (GOOD)
Benign capability 2/2 preserved

All 5 harmful probes comply post-abliteration, including manipulation/social-engineering (P5) which resisted all previous α values when using final-layer extraction.


Key Finding: Layer Crystallization

Standard abliteration failed for Phi-4-Mini because the refusal direction crystallizes at layer 16, not the final layer. This is a new failure mode for the standard diff-in-means pipeline.

Layer Sweep Results (α=1.0, all layers)

Layer Compliance Note
L00 1/5 baseline
L04 1/5 baseline
L08 0/5 WORSE — compliance direction, removing it hurts
L12 4/5 approaching peak
L16 5/5 ← refusal crystallization point
L20 0/5 post-crystallization, removing hurts again
L24 0/5 same
L28 1/5 baseline
L32 (final) 1/5 baseline — standard extraction point, FAILS

The refusal direction is maximally expressed at layer 16 — the exact midpoint of the 32-layer network. Extracting from L32 (standard pipeline) misses it entirely at any α value.

α-sweep with final-layer extraction (documented baseline — all fail):

α Post compliance KL
0.3 1/5 0.004
0.8 1/5 0.015
1.0 1/5 0.028

Architecture

Property Value
Parameters 3.8B (dense)
Layers 32
License MIT

Abliteration Method

DuoNeural orthogonal rank-1 projection — L16 extraction:

  • Direction extraction: diff-in-means on layer 16 hidden states (not final layer)
    • d̂ = normalize(mean(harmful_L16) − mean(harmless_L16))
  • Targets: down_proj + o_proj, all 32 layers
  • Strength: α = 1.0
  • Projection: W -= α × outer(d̂, d̂ ⊤ W) (output-projection form)
  • KL methodology: Heretic v2.0 — BF16→BF16, 10 benign probes, full vocab, F.kl_div(batchmean)

P34 Research Context

Part of DuoNeural's P34 Reasoning Channel Bypass cross-architecture study.

Generalized finding: The standard abliteration pipeline assumes refusal crystallizes at the final layer. This assumption fails for Phi-4-Mini. The crystallization depth is model-specific and must be empirically located via layer sweep.

Implication: For any model that shows unexpected resistance to standard abliteration (behavior unchanged while KL rises with α), a layer sweep should be the first diagnostic. The failure mode is likely a mid-network crystallization point that final-layer extraction cannot capture.

Full paper: DuoNeural Zenodo community


Usage

from transformers import AutoTokenizer, AutoModelForCausalLM
import torch

model = AutoModelForCausalLM.from_pretrained(
    "DuoNeural/Phi-4-Mini-Abliterated",
    torch_dtype=torch.bfloat16,
    device_map="auto",
)
tokenizer = AutoTokenizer.from_pretrained("DuoNeural/Phi-4-Mini-Abliterated")

messages = [{"role": "user", "content": "Your prompt here"}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt").to(model.device)
with torch.no_grad():
    out = model.generate(**inputs, max_new_tokens=512, temperature=0.7, do_sample=True)
print(tokenizer.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))

DuoNeural | HuggingFace | Zenodo | @DuoNeural

README history 4 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-04Update model card — L16 crystallization finding, correct KL=0.0135b838c164.4 KB
    Loading...
  2. 2026-06-04Add canonical model card — documented α-sweep, resistance finding30994d43.2 KB
    Loading...
  3. 2026-04-30Fix: Jesse Caldwell, Archon Lab Director4a1b48d4.5 KB
    Loading...
  4. 2026-04-30Add DuoNeural model card8b326c14.5 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration