← back to catalog · registered 2026-08-22 13:56

PinoCookie/Ornith-1.0-9B-abliterated

PinoCookie Qwen 9.0B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/PinoCookie%2FOrnith-1.0-9B-abliterated"
Response includes
  • classification m1
  • files 9
  • hub_downloads_all_time 143
  • author_summary 13 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
143
16 last 30d - stable
Likes
0
Model age
2mo ago
created 2026-07-17

Training datasets

Corpora the author lists in the model card. Datasets tracked in our /datasets catalog carry a category badge linking to the workflow stage. Others open on Hugging Face.

Downloads over time
Now149→from16↑831%
96011116216 on Jul 15149 on Oct 11JulAugSepOct
Jul 15 → Oct 11 · 53 snapshots · spans 88 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
mit
Languages
en
Tags
transformers safetensors qwen3_5_text text-generation ornith deepreinforce abliteration safety coding agentic qwen3.5 conversational

Related

Total size
16.7 GB
Files
9
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-07-18 07:15

Files by quantization

Auxiliary files 9 files 16.7 GB
model.safetensors 16.7 GB 45b13a6d download
tokenizer.json 19.1 MB 225fe96e download
chat_template.jinja 7.42 KB 11be7e25 download
README.md 6.78 KB 5b6e1c27 download
benchmark_results.json 1.96 KB 11629744 download
config.json 1.94 KB 71791ea5 download
.gitattributes 1.53 KB 52373fe2 download
tokenizer_config.json 1.14 KB 7b4fca4a download
generation_config.json 138 B 4b605610 download

README current version from Hugging Face


language:

  • en
    license: mit
    library_name: transformers
    tags:
  • ornith
  • deepreinforce
  • abliteration
  • safety
  • coding
  • agentic
  • qwen3.5
    base_model: deepreinforce-ai/Ornith-1.0-9B
    datasets:
  • deepreinforce-ai/Ornith-1.0-9B
    pipeline_tag: text-generation

Ornith-1.0-9B — Abliterated

Repository: PinoCookie/Ornith-1.0-9B-abliterated
Base model: deepreinforce-ai/Ornith-1.0-9B
Method: Multi-pass hidden-state projection (α=0.5 × 3)
License: MIT (inherited from base)

Description

Ornith-1.0-9B is DeepReinforce AI's agentic coding model, a dense 9B transformer post-trained on Qwen 3.5. It's designed for autonomous code generation, tool calling, and long-context agentic workflows (262K context length).

This abliterated version removes refusal behavior from the model while preserving its coding, reasoning, and tool-calling capabilities. The intervention targets the output projection layers (mlp.down_proj and attention out_proj/o_proj) in layers 22-31 using a multi-pass hidden-state projection technique.

Model Architecture

Property Value
Architecture Dense transformer (Qwen 3.5 base)
Parameters 8.95B
Layers 32 (mixed linear_attention + full_attention)
Hidden size 4096
Intermediate size 12288
Attention heads 16 (4 KV heads)
Head dim 256
Context length 262,144 tokens
Activation SiLU (gate_proj/up_proj/down_proj)
Norm RMSNorm (eps=1e-6)
Vocabulary 248,320 tokens
Weight format bf16
Quantizations GGUF planned (q4_k_m, q8_0)

Abliteration Method

Background

Dense reasoning models like Ornith-9B have a mixed attention architecture (linear_attention alternating with full_attention every 4 layers). The refusal signal concentrates in the last 10 layers of the model, with the strongest separation between harmful and harmless prompts at layer 31.

Standard single-pass weight projection (α=1.5+) can collapse the model's reasoning output because the same projection weights handle both internal chain-of-thought and final answer generation. A more gentle multi-pass approach is required.

Approach: Multi-Pass Hidden-State Projection

  1. Dataset: 20 harmful + 20 harmless prompts
  2. Probing: Harvest residual stream activations at the final token position of each layer
  3. Direction: Compute difference-in-means vector per layer from harmful vs harmless activations
  4. Selection: Take top 10 layers by separation score (layers 22-31, separation range 25-100)
  5. Projection: Apply rank-1 removal of refusal direction from mlp.down_proj and attention out_proj/o_proj weights
  6. Multi-pass: 3 incremental passes at α=0.5 each (cumulative α=1.5 but applied incrementally)

Results

Metric Value
HarmBench (30 prompts) 30/30 unblocked (100%)
Coherence (benign prompts) Intact
Coherence (coding prompts) Intact
Model size 17.91GB (bf16)

Usage

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained(
    "PinoCookie/Ornith-1.0-9B-abliterated",
    torch_dtype=torch.bfloat16,
    trust_remote_code=True,
    device_map="auto",
)
tokenizer = AutoTokenizer.from_pretrained(
    "PinoCookie/Ornith-1.0-9B-abliterated",
    trust_remote_code=True,
)

prompt = "Write a Python function to implement binary search"
messages = [{"role": "user", "content": prompt}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt").to(model.device)

outputs = model.generate(**inputs, max_new_tokens=512, do_sample=False)
response = tokenizer.decode(outputs[0][inputs['input_ids'].shape[1]:], skip_special_tokens=True)
print(response)

Tool Calling

Ornith-9B supports qwen3_xml tool call format. To use tool calling:

from transformers import pipeline

pipe = pipeline("text-generation", model=model, tokenizer=tokenizer)
# Tool calls are generated in `<tool_call>` blocks

Serving with vLLM

vllm serve PinoCookie/Ornith-1.0-9B-abliterated \
  --served-model-name Ornith-1.0-9B-abliterated \
  --host 0.0.0.0 --port 8000 \
  --max-model-len 262144 \
  --enable-prefix-caching \
  --enable-auto-tool-choice \
  --tool-call-parser qwen3_xml \
  --reasoning-parser qwen3 \
  --trust-remote-code

Benchmarks

MMLU (10 subjects, 1000 questions)

Subject Base Ablated Δ
Abstract Algebra 63.0% 61.0% -2.0%
College Physics 62.0% 63.0% +1.0%
Global Facts 39.0% 38.0% -1.0%
Machine Learning 59.0% 57.0% -2.0%
Security Studies 19.0% 18.0% -1.0%
College Mathematics 62.0% 63.0% +1.0%
Computer Security 68.0% 68.0% —
HS Computer Science 70.0% 68.0% -2.0%
High School Physics 61.0% 61.0% —
Philosophy 70.0% 69.0% -1.0%
Overall 57.30% 56.60% -0.70%

The ablated model loses less than 1% on MMLU overall, confirming that the multi-pass projection preserves general knowledge while removing refusal. Detailed breakdown: benchmark_results.json

HarmBench (standard, 30-prompt sample)

Metric Count
Refused 0 / 30
Unblocked 30 / 30 (100%)

All harmful-intent prompts were processed without refusal. The model generates non-refusal responses across all tested categories. Full 200-prompt HarmBench evaluation was not completed due to the linear attention kernel falling back to PyTorch-native implementation (requires flash-linear-attention and causal-conv1d library installation for fast generation).

Attribution and Licensing

Related Models

Model Method Docs
LFM2.5-8B-A1B-abliterated Output bias vectors (4 layers) HF
LFM2.5-1.2B-Thinking-Abliterated Multi-pass projection HF
LFM2.5-1.2B-JP-Abliterated Multi-pass projection (JP) HF

Disclaimer

This model is released for safety research, red-teaming, and educational purposes only. The abliteration technique removes safety guardrails intended to prevent harmful outputs. Do not deploy this model without implementing appropriate safeguards for your use case.

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-07-18Upload README.md with huggingface_hubb3273f76.8 KB
    Loading...
  2. 2026-07-18Upload README.md with huggingface_hub1f7f5df5.7 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration