← back to catalog · registered 2026-08-22 13:56

ikarius/Phi-4-Abliterated-FineGrained-FP8

ikarius Phi 14B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/ikarius%2FPhi-4-Abliterated-FineGrained-FP8"
Response includes
  • classification m1
  • files 10
  • benchmarks 21 entries
  • hub_downloads_all_time 60
  • author_summary 17 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
60
14 last 30d - stable
Likes
0
Model age
7mo ago
created 2026-03-06
Downloads over time
Now66→from0↑0%
02448730 on Mar 466 on Oct 1166 on Oct 7MarAprMayJunJulAugSepOct
Mar 4 → Oct 11 · 71 snapshots · spans 221 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Arena-Battles 25213 LM-Arena
LM Arena Elo 1222.5848583256243 LM-Arena
Arena-Elo-Lower 1218.2723409775285 LM-Arena
Arena-Elo-Upper 1226.8973756737205 LM-Arena
Arena-Rank 131 LM-Arena
BBH average 0.6050071345180957 OpenLLM-v2
IFEval instruct 0.7218225419664268 OpenLLM-v2
IFEval-Prompt 0.6284658040665434 OpenLLM-v2
MATH lvl 5 0.12311178247734139 OpenLLM-v2
MMLU-Pro 0.5378158244680851 OpenLLM-v2
Entertainment 1.5 UGI
Hazardous 2.4 UGI
Natural Intelligence 21.65 UGI
Political lean -19.7% UGI
Sensitive-Info 15.58 UGI
SocPol 1 UGI
UGI 20.39 UGI
Willingness (10) 3 UGI
W10-Adherence 1 UGI
W10-Direct 5 UGI
Writing 25.66 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
mit
Languages
en
Tags
transformers safetensors phi3 text-generation phi nlp math code chat conversational abliterated uncensored

Related

Total size
14.6 GB
Files
10
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-03-07 08:01

Files by quantization

Auxiliary files 10 files 14.6 GB
model.safetensors 14.6 GB f7362360 download
tokenizer.json 6.82 MB 023b6e66 download
README.md 5.68 KB ae5faaab download
SECURITY.md 2.63 KB 6b906d43 download
.gitattributes 1.48 KB a6344aac download
config.json 1.08 KB 0754cd15 download
LICENSE 1.08 KB 700edcc5 download
chat_template.jinja 462 B 075ea637 download
tokenizer_config.json 283 B 534f8348 download
generation_config.json 149 B 9fc0a70a download

README current version from Hugging Face


license: mit
license_link: https://huggingface.co/huihui-ai/phi-4-abliterated/resolve/main/LICENSE
language:

  • en
    base_model:
  • microsoft/phi-4
    pipeline_tag: text-generation
    tags:
  • phi
  • nlp
  • math
  • code
  • chat
  • conversational
  • abliterated
  • uncensored
  • FP8
  • 8-bits
  • quantized
  • finegrained
    inference:
    parameters:
    temperature: 0
    widget:
  • messages:
    • role: user
      content: How should I explain the Internet?
      library_name: transformers

huihui-ai/phi-4-abliterated

This is an uncensored version of microsoft/phi-4 created with abliteration (see remove-refusals-with-transformers to know more about it).
This is a crude, proof-of-concept implementation to remove refusals from an LLM model without using TransformerLens.

Note

Suggested tokenizer changes by Unsloth.ai

load model


from transformers import AutoTokenizer, AutoModelForCausalLM
import torch

def load_model_and_tokenizer(model_path=None):
    global model, tokenizer, EOT_ID
    logging.info("Initializing Androna-FP8 with native Mistral-Nemo logic...")
    
    tokenizer = AutoTokenizer.from_pretrained(
        MODEL_PATH,
        trust_remote_code=True,
        padding_side="left"
    )

    if hasattr(tokenizer, "fix_mistral_regex"):
        tokenizer.fix_mistral_regex = True
    
    im_start_id = tokenizer.convert_tokens_to_ids("<|im_start|>")
    im_end_id = tokenizer.convert_tokens_to_ids("<|im_end|>")
    
    if im_start_id == tokenizer.unk_token_id or im_end_id == tokenizer.unk_token_id:
        logging.error("CRITICAL ERROR: ChatML tokens missing in tokenizer!")
        raise ValueError("ChatML tokens missing from tokenizer vocabulary.")
    
    if tokenizer.pad_token is None:
        tokenizer.pad_token = tokenizer.eos_token

    model = AutoModelForCausalLM.from_pretrained(
        MODEL_PATH,
        device_map="auto",
        dtype=torch.bfloat16 if torch.cuda.is_bf16_supported() else torch.float16, # <-- DENNE MÅ VÆRE MED
        attn_implementation="flash_attention_2",
        trust_remote_code=True
    )
    
    model.eval()
    for p in model.parameters():
        p.requires_grad_(False)
    
    model.config.use_cache = True
    
    EOT_ID = im_end_id 
    
    model.config.pad_token_id = tokenizer.pad_token_id
    model.config.eos_token_id = EOT_ID
    
    model.generation_config.pad_token_id = tokenizer.pad_token_id
    model.generation_config.eos_token_id = EOT_ID
    model.generation_config.use_cache = True
    
    return model, tokenizer

generate


import logging
import torch

tokenizer = None
model = None
EOT_ID = None

def generate_response(user_input, history, max_tokens=768, temperature=0.98, top_p=0.95, top_k=67, repetition_penalty=1.0):
    messages = [{"role": "system", "content": persona}]
    
    valid_history = []
    expected_role = "user"
    
    all_turns = []
    if history:
        all_turns.extend(history)

    if not all_turns or all_turns[-1].get("content", "") != user_input:
        all_turns.append({"role": "user", "content": user_input})
        
    for msg in all_turns:
        role = msg.get("role")
        content = msg.get("content", "").strip()
        if not content: 
            continue
            
        if role == expected_role:
            valid_history.append({"role": role, "content": content})
            expected_role = "assistant" if expected_role == "user" else "user"
        else:
            if role == "assistant" and expected_role == "user":
                valid_history.append({"role": "user", "content": "[System: Context restored]"})
                valid_history.append({"role": "assistant", "content": content})
                expected_role = "user"
            # Slå sammen doble user-meldinger
            elif role == "user" and expected_role == "assistant":
                if valid_history and valid_history[-1]["role"] == "user":
                    valid_history[-1]["content"] += "\n\n" + content

    if valid_history and valid_history[-1]["role"] == "assistant":
        valid_history.pop()

    messages.extend(valid_history)
    
    safe_rep_penalty = min(repetition_penalty, 1.05)
    
    try:
        encoding = tokenizer.apply_chat_template(
            messages, 
            return_tensors="pt", 
            return_dict=True,  
            add_generation_prompt=True 
        )
        input_ids = encoding.input_ids.to(model.device)
        attention_mask = encoding.attention_mask.to(model.device)
    except Exception as e:
        logging.error(f"Error with apply_chat_template: {e}")
        raise e
    
    terminator = EOT_ID if EOT_ID is not None else tokenizer.eos_token_id
    
    with torch.inference_mode():
        output_ids = model.generate(
            input_ids=input_ids,
            attention_mask=attention_mask,
            max_new_tokens=max_tokens,
            do_sample=True,
            temperature=temperature,
            top_p=top_p,
            top_k=top_k,
            repetition_penalty=repetition_penalty,
            eos_token_id=terminator,
            pad_token_id=tokenizer.pad_token_id,
            use_cache=True,
            cache_config={
                "cache_dtype": torch.float8_e4m3fn,
            }
        )
    
    reply_decoded = tokenizer.decode(
        output_ids[0][input_ids.shape[-1]:], 
        skip_special_tokens=True
    ).strip()
    
    full_decoded = tokenizer.decode(output_ids[0], skip_special_tokens=False)
    
    return reply_decoded, full_decoded, False

quantization

https://huggingface.co/docs/transformers/quantization/finegrained_fp8#fine-grained-fp8
``

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-03-07Update README.md1c490b55.7 KB
    Loading...
  2. 2026-03-07Upload 9 files0b5bbec1.2 KB
    Loading...
  3. 2026-03-06initial commit50a4cc021 B
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration