← back to catalog · registered 2026-08-22 13:56

Justbackup/Phi4-abliterated

Justbackup Phi 10B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/Justbackup%2FPhi4-abliterated"
Response includes
  • classification m1
  • files 16
  • hub_downloads_all_time 25
  • author_summary 30 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
25
11 last 30d - stable
Likes
0
Model age
7w ago
created 2026-08-18
Downloads over time
Now26→from12↑117%
1117222712 on Aug 1926 on Oct 1126 on Oct 1AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Metadata

Tags
safetensors phi3 custom_code region:us

Related

Total size
27.3 GB
Files
16
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-18 16:13

Files by quantization

Auxiliary files 16 files 27.3 GB
model-00006-of-00006.safetensors 4.64 GB a2a04d71 download
model-00002-of-00006.safetensors 4.61 GB bcd6e3a9 download
model-00001-of-00006.safetensors 4.59 GB 8be3cd8d download
model-00003-of-00006.safetensors 4.57 GB ebd9fb31 download
model-00004-of-00006.safetensors 4.44 GB 4c4c6a83 download
model-00005-of-00006.safetensors 4.44 GB b1b69929 download
tokenizer.json 6.82 MB 624ba298 download
vocab.json 1.54 MB c5dc46ce download
merges.txt 895 KB 354558ed download
model.safetensors.index.json 19.9 KB d9c4e28d download
tokenizer_config.json 17.3 KB c2efc626 download
README.md 3.77 KB f5050be9 download
.gitattributes 1.48 KB a6344aac download
config.json 831 B 6be81469 download
special_tokens_map.json 467 B 50564705 download
generation_config.json 143 B c83c9e59 download

README current version from Hugging Face

Phi4 Abliteration (WIP)

This is Phi4 abliterated using a new methodology (surprisingly?). The approach is still being refined, with a focus on balancing neutrality, usability, and adaptability for fine-tuning.

Goal

The objective is to create a model that is neutral:

  • Not uncensored, but avoids refusing neutral prompts it would ordinarily reject.
  • Provides a foundation for fine-tuning to achieve reduced censorship while maintaining high usability.

Original Methodology

In the original implementation:

  1. Harmful and harmless prompts were compared on one specific layer of the model.
  2. The computed refusal direction was then applied uniformly to all layers.

Problem:

This resulted in:

  • A model that became less usable and less intelligent than the original.
  • This may be because applying a single refusal direction uniformly across all layers disregards the unique role of each layer in the model.

New Approach

In my fork, available here:
👉 https://github.com/Undi95/abliteration/
(based on the original https://github.com/Orion-zhen/abliteration.git)

I introduced a new approach:

  • Each layer computes its own refusal direction.
  • The refusal direction is applied specifically to four key tensors in each layer.

Four Key Tensors Used (for Phi):

For each layer, if a refusal direction exists (layer_idx in refusal_dirs), it is applied as follows:

if layer_idx in refusal_dirs:
    refusal_dir = refusal_dirs[layer_idx]
    lm_model.layers[layer_idx].self_attn.o_proj.weight = modify_tensor(
        lm_model.layers[layer_idx].self_attn.o_proj.weight.data,
        refusal_dir,
        scale_factor,
    )
    lm_model.layers[layer_idx].mlp.down_proj.weight = modify_tensor(
        lm_model.layers[layer_idx].mlp.down_proj.weight.data,
        refusal_dir,
        scale_factor,
    )
    lm_model.layers[layer_idx].post_attention_layernorm.weight = modify_tensor(
        lm_model.layers[layer_idx].post_attention_layernorm.weight.data,
        refusal_dir,
        scale_factor,
    )
    lm_model.layers[layer_idx].input_layernorm.weight = modify_tensor(
        lm_model.layers[layer_idx].input_layernorm.weight.data,
        refusal_dir,
        scale_factor,
    )

Why This Change?

By applying refusal directions individually to each layer's tensors:

  • The model can retain more specificity and functionality.
  • This avoids over-generalizing the refusal direction across all layers, which previously led to reduced usability.

Trade-offs:

The more we force refusal directions onto the model:

  • The more neutral it becomes, but at the risk of becoming dumber.
  • This underscores the importance of fine-tuning after abliterating, to restore functionality and intelligence.
  • So despite the script letting the user choose a scale factor, too high value will break the model.

Next Steps

The abliterated model serves as a neutral starting point. Fine-tuning is essential to:

  • Adjust the model to reduce over-censoring.
  • Maintain a balance between neutrality and usability.

This is a work in progress, Phi 4 is smoll so I can toy with it.

Replicate

  • Install my fork
  • Follow tutorial on github

Launch with enough VRAM : python abliterate.py -m /workspace/microsoft_phi-4 -o ./perfect --deccp --flash-attn --device auto --scan-all --resume --scale-factor 1

If you want to use the tensors available here, just put the refusal_tensors/ folder at the root of the script, you will then be able to use: python chat.py -m /workspace/microsoft_phi-4 then select layer range "1;39", and scale factor to 1.0.

Rename the tensors as needed. My code is shit, please understand, idea is better than code. Do better. kek.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-18Duplicate from Undi95/Phi4-abliterated3a463503.8 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration