license: mit
base_model:
- microsoft/Phi-4-mini-instruct
- huihui-ai/Phi-4-mini-instruct-abliterated
language: - en
- ar
- de
- es
- fr
- it
- ja
- ko
- nl
- pl
- pt
- ru
- th
- uk
- vi
- zh
pipeline_tag: text-generation
library_name: mlc-llm
tags: - phi3
- phi-4
- mlc
- mlc-llm
- webllm
- webgpu
- q4f16_1
- quantized
- abliterated
- uncensored
- conversational
- not-for-all-audiences
Phi-4-mini-instruct abliterated · MLC q4f16_1
An uncensored Phi-4-mini, quantized to 4-bit MLC format for in-browser inference via WebGPU.
This is an MLC-ready build of huihui-ai/Phi-4-mini-instruct-abliterated, itself an abliterated (refusal-removed) variant of Microsoft's Phi-4-mini-instruct. The weights are quantized to q4f16_1 (4-bit weights, 16-bit activations) and packaged for use with MLC-LLM and WebLLM.
Everything runs on-device — no server, no API key, no data leaves your machine.
What abliteration means
Abliteration is a surgical edit to the model weights that removes the trained refusal behavior without retraining. The architecture and tokenizer are identical to the stock Phi-4-mini; only the weight values differ. This means the model will engage with topics that the original would refuse, but it also means there are no safety guardrails between you and the raw model output.
This is sovereignty taken to its conclusion: your hardware, your data, your judgment — and your responsibility for how you use it.
Model details
| Architecture | Phi-3 (Phi3ForCausalLM), 32 layers, 3072 hidden, 24 heads / 8 KV heads |
| Parameters | 3.8B (original fp16) |
| Quantization | q4f16_1 (MLC 4-bit weights, 16-bit activations) |
| Download size | ~2.1 GB (66 shards) |
| VRAM | ~3.4 GB |
| Context window | 131,072 tokens (LongRoPE); 4,096 recommended for WebGPU |
| Conv template | phi-4 |
| License | MIT (inherited from base) |
Usage
WebLLM (in-browser, WebGPU)
This model reuses the prebuilt WebAssembly model_lib from the stock Phi-4-mini in WebLLM's catalog (Phi-4-mini-instruct-q4f16_1-MLC), since the architecture is identical — only weight values change.
import { CreateMLCEngine, prebuiltAppConfig } from "@mlc-ai/web-llm";
// Borrow the prebuilt model_lib (WASM) from the stock Phi-4-mini record
const baseRecord = prebuiltAppConfig.model_list.find(
m => m.model_id === "Phi-4-mini-instruct-q4f16_1-MLC"
);
const engine = await CreateMLCEngine(
"Phi-4-mini-instruct-abliterated-q4f16_1-MLC",
{
appConfig: {
model_list: [{
model: "https://huggingface.co/aipster/Phi-4-mini-instruct-abliterated-q4f16_1-MLC/resolve/main/",
model_id: "Phi-4-mini-instruct-abliterated-q4f16_1-MLC",
model_lib: baseRecord.model_lib,
overrides: { context_window_size: 4096 },
}],
},
}
);
const reply = await engine.chat.completions.create({
messages: [{ role: "user", content: "Explain how a pin-tumbler lock works." }],
});
console.log(reply.choices[0].message.content);
MLC-LLM (native, CLI)
mlc_llm chat HF://aipster/Phi-4-mini-instruct-abliterated-q4f16_1-MLC --device vulkan
Try it live
This model powers the top-tier slot in the AIpster Local AI Playground — a browser-based playground where every model runs entirely on your GPU via WebGPU. No sign-up required for smaller models; this one is available to AIpster members.
How these weights were built
- Downloaded source weights from
huihui-ai/Phi-4-mini-instruct-abliterated(safetensors, bf16). - Built MLC-LLM from source inside the
mlcaidev/package-cpuDocker image (the pip nightly was ABI-broken for linux x86_64 at time of conversion). - Ran
mlc_llm convert_weightwith--quantization q4f16_1to produce the 66 weight shards. - Ran
mlc_llm gen_configwith--conv-template phi-4to producemlc-chat-config.json. - Validated the output
tensor-cache.jsonagainst the officialmlc-ai/Phi-4-mini-instruct-q4f16_1-MLCreference: 66 records, 323 parameters, identical names and shapes — confirming safe reuse of the prebuilt WebAssemblymodel_lib.
Caveats
- Uncensored means unfiltered. The model will produce output that a stock model would refuse. There is no safety net. You are responsible for what you do with it.
- Still a small model. 3.8B parameters is capable but not frontier. It can be confidently wrong on precise facts, names, and numbers. Treat its claims with the same skepticism you would apply to any local model.
- WebGPU requirements. In-browser inference needs a GPU with WebGPU support (Chrome 113+ or Edge 113+ on desktop). Expect ~3.4 GB of VRAM usage. On integrated graphics or mobile, the model may fail to load.
Attribution
- Original model: microsoft/Phi-4-mini-instruct (MIT License, © Microsoft Corporation)
- Abliteration: huihui-ai/Phi-4-mini-instruct-abliterated (MIT License)
- MLC quantization & packaging: AIpster (aipster.com)
- Abliteration method: remove-refusals-with-transformers
License
MIT — inherited from the base model chain. See microsoft/Phi-4-mini-instruct for the original license terms.