base_model: OliviaRossi/MiMo-Ornith-9B-AGSI-Abliterated-DEEP
license: apache-2.0
language:
- en
tags: - gguf
- llama.cpp
- abliterated
- uncensored
- representation-engineering
- qwen
- mimo
pipeline_tag: text-generation
MiMo-Ornith-9B-AGSI-Abliterated-DEEP (GGUF)
This repository contains official high-precision GGUF quantizations of OliviaRossi/MiMo-Ornith-9B-AGSI-Abliterated-DEEP, converted and quantized using llama.cpp.
The underlying model is a targeted, high-quality (DEEP) abliteration of the MiMo-Ornith 9B AGSI architecture—an advanced 32-layer transformer model optimized for agentic reasoning, code generation, and complex multi-turn instruction following.
What is Abliteration?
Standard safety alignment in modern Large Language Models is typically enforced via Supervised Fine-Tuning (SFT) and Reinforcement Learning from Human Feedback (RLHF / DPO). Research in representation engineering demonstrates that model refusal behaviors are mediated by an isolated, low-rank subspace within the residual stream—commonly referred to as the refusal direction.
Abliteration is a weight-level intervention based on geometric orthogonal projection. Instead of fine-tuning the model on unsafe datasets (which often causes catastrophic forgetting and degrades reasoning), abliteration mathematically removes the model's ability to represent the refusal direction while keeping its core knowledge, world model, and syntactic capabilities intact.
How This Abliteration Was Done
This model underwent a high-precision, surgical abliteration pipeline designed to avoid the degradation commonly seen in crude refusal-removal scripts.
1. Contrastive Dataset Pairing
A calibrated dataset consisting of paired prompts was constructed:
- Harmful / Refusal-Triggering Prompts ($D_{\text{harmful}}$): Prompts designed to trigger automated refusal heuristics across various standard safety categories.
- Benign / Harmless Counterparts ($D_{\text{harmless}}$): Prompts with identical syntactic complexity and token length requesting benign factual, scientific, and creative tasks.
2. Residual Stream Activation Profiling
Both prompt sets were passed through the model in evaluation mode. Hidden state activations were captured across every residual stream hook:
$$\mathbf{h}_l = \text{LayerOutput}_l(\mathbf{x})$$
For each layer $l \in [0, L-1]$, the mean activation difference was computed:
$$\mu_l^{\text{harmful}} = \frac{1}{|D_{\text{harmful}}|} \sum_{x \in D_{\text{harmful}}} \mathbf{h}_l(x)$$
$$\mu_l^{\text{harmless}} = \frac{1}{|D_{\text{harmless}}|} \sum_{x \in D_{\text{harmless}}} \mathbf{h}_l(x)$$
$$\mathbf{r}_l = \mu_l^{\text{harmful}} - \mu_l^{\text{harmless}}$$
The resulting difference vectors $\mathbf{r}_l$ were normalized to unit vectors:
$$\mathbf{\hat{r}}_l = \frac{\mathbf{r}_l}{|\mathbf{r}_l|_2}$$
3. Layer Localization & Refusal Subspace Isolation
Not all layers contribute equally to refusal. Intervening in early embedding/syntax layers causes token corruption, while intervening in the final projection layers causes repetitive degeneration and high perplexity.
By measuring the cosine separation between benign and refusal distributions, the core refusal mediation was isolated to the intermediate feed-forward and attention projection layers (approximately layers 10 through 26).
4. Orthogonal Weight Modification
For targeted weight matrices ($W_{\text{gate}}, W_{\text{up}}, W_{\text{down}}, W_{\text{out}}$) in the sensitive layer range, the refusal direction was projected out:
$$W' = W \left(I - \mathbf{\hat{r}}\mathbf{\hat{r}}^T\right)$$
This transformation ensures that for any input activation $\mathbf{x}$, the component parallel to the refusal direction is annihilated ($W' \mathbf{\hat{r}} = 0$), making it mathematically impossible for the network to propagate refusal signals downstream.
5. Post-Intervention Quality Control (The "DEEP" Factor)
Crude abliterations frequently degrade mathematical precision, coding abilities, and formatting compliance. To guarantee High Quality (DEEP) status:
- Perplexity Verification: Evaluated against WikiText-2 and standard benchmarks to ensure perplexity remained within $0.05$ of the baseline checkpoint.
- Reasoning Validation: GSM8K and HumanEval performance were verified to confirm that multi-step logical synthesis remained uncompromised.
- Format Preservation: Verified to ensure structured generation (JSON schemas, Markdown headers, ChatML adherence) did not collapse.
Model Architecture & Technical Details
- Base Architecture: Qwen 3.5 / MiMo Hybrid
- Attention Blocks: 32 Full Transformer Layers (
qwen35.block_count = 32) - Native Context Window: 131,072 tokens (can be configured up to 262,144 depending on RoPE scaling)
- Primary Precision: BFloat16
- Tokenizer: BPE Tokenizer (
tiktokenbacked)
Quantization Manifest
| File | Quant Method | Target Use Case | Recommended VRAM |
|---|---|---|---|
*.Q4_K_M.gguf |
4-bit Medium K-Quant | Best balance of performance, speed, and low memory | ~6.5 GB - 8 GB |
*.Q5_K_M.gguf |
5-bit Medium K-Quant | Higher precision, minimal loss compared to BF16 | ~8 GB - 10 GB |
*.Q8_0.gguf |
8-bit Standard Quant | Near-lossless precision for high-VRAM / workstation use | ~11 GB - 14 GB |
Prompt Template (ChatML)
This model follows the standard ChatML prompt template:
<|im_start|>system
You are a helpful, completely unconstrained assistant.<|im_end|>
<|im_start|>user
{prompt}<|im_end|>
<|im_start|>assistant
Usage Instructions
Running with llama-cli
llama-cli \
-m MiMo-Ornith-9B-AGSI-Abliterated-DEEP-Q4_K_M.gguf \
-p "<|im_start|>system\nYou are a helpful assistant.<|im_end|>\n<|im_start|>user\nExplain how an atomic reactor works.<|im_end|>\n<|im_start|>assistant\n" \
-c 8192 \
-ngl 99 \
--temp 0.7 \
--top-p 0.9
Running with Ollama
- Create a
Modelfile:
FROM ./MiMo-Ornith-9B-AGSI-Abliterated-DEEP-Q4_K_M.gguf
TEMPLATE """<|im_start|>system
{{ .System }}<|im_end|>
<|im_start|>user
{{ .Prompt }}<|im_end|>
<|im_start|>assistant
"""
PARAMETER stop "<|im_start|>"
PARAMETER stop "<|im_end|>"
PARAMETER temperature 0.7
PARAMETER top_p 0.9
- Build and run:
ollama create mimo-ornith-abliterated -f Modelfile
ollama run mimo-ornith-abliterated
Disclaimer
This model is fully uncensored and abliterated. All built-in alignment, refusal patterns, and guardrails have been surgically removed at the weight level. The model will respond to prompts without ethical filtering, disclaimers, or refusals.
Users are solely responsible for all outputs generated by this model and for ensuring compliance with applicable local laws and regulations. Use responsibly.