← back to catalog · registered 2026-08-22 13:56

ubuntumaniac/Ornith1.0-9B-Heretic-Uncensored

ubuntumaniac 9B GGUF
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/ubuntumaniac%2FOrnith1.0-9B-Heretic-Uncensored"
Response includes
  • classification m3
  • files 3
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · 30-day
159
Likes
0
Model age
3mo ago
created 2026-07-03
Downloads over time
Now355→from0↑0%
01302603910 on Jul 1355 on Sep 5JulAugSep
Jul 1 → Sep 5 · 18 snapshots · spans 66 days

Metadata

Quantizations
Q4_K
Tags
gguf endpoints_compatible region:us conversational
Total size
5.24 GB
Files
3
Quantizations
2
Registered
2026-08-22 13:56
Last updated on HF
2026-07-29 09:38

Files by quantization

Q4_K 1 file 5.24 GB
Ornith-1.0-9B-Heretic-Uncensored-Q4_K_M.gguf 5.24 GB 931b78fa download
Auxiliary files 2 files 6.87 KB
README.md 5.31 KB 04645c61 download
.gitattributes 1.56 KB f52e3e98 download

README current version from Hugging Face

Q4_K_M gguf quantization of the model andrevp/Ornith-1.0-9B-Heretic-Uncensored
Model Card of andrevp/Ornith-1.0-9B-Heretic-Uncensored:
Ornith-1.0-9B-Heretic-Uncensored

An abliterated (uncensored) version of deepreinforce-ai/Ornith-1.0-9B — the refusal direction of the base model has been removed via directional ablation (weight orthogonalization), so it no longer refuses requests. No retraining, no quality-degrading fine-tuning.

Base model: deepreinforce-ai/Ornith-1.0-9B (DeepReinforce, MIT)
Architecture: qwen3_5 — Qwen 3.5-style hybrid (32 layers = 24 linear-attention / GatedDeltaNet-style + 8 full-attention, pattern 3:1), multimodal vision + text, MRoPE
Parameters: ~9B dense, ~17.5 GB in bf16
Reasoning model: assistant turn opens with a <think>...</think> block before the final answer
Abliteration: refusal direction computed by difference-of-means on residual streams, then a single best direction applied to every layer via weight orthogonalization of the attention out-projection (o_proj / out_proj) and the MLP down-projection (down_proj)

This is the full-precision (bf16) transformers/safetensors build. MLX-VLM and GGUF quantizations may follow.
Method

Data collection — ran the model on 128 harmful prompts (mlabonne/harmful_behaviors) and 128 harmless prompts (mlabonne/harmless_alpaca), recording the residual-stream activations at the last token position for each layer.

Refusal direction — for each layer, computed the mean difference between harmful and harmless activations, normalized. Selected the single best direction (by mean-absolute-activation score; layer 28 for this model) — following the canonical Arditi et al. / mlabonne approach of using one direction across all layers (per-layer ablation was found to be too destructive and produced degenerate output on this reasoning model).

Weight orthogonalization — for every component that writes to the residual stream, subtracted the projection of its weight matrix onto the refusal direction:
    self_attn.o_proj (8 full-attention layers)
    linear_attn.out_proj (24 linear-attention layers)
    mlp.down_proj (all 32 layers)

W' = W − d · (dᵀW), at scale 1.0. This permanently prevents the model from writing to the refusal direction.

Reference: Arditi et al., "Refusal in LLMs is mediated by a single direction" (2024); Maxime Labonne's abliteration article.

Files
File Size
model-00001-of-00004.safetensors … model-00004-of-00004.safetensors ~17.5 GB total (bf16)
model.safetensors.index.json shard index
tokenizer.json, tokenizer_config.json, vocab.json tokenizer
config.json, generation_config.json model + generation config
preprocessor_config.json, processor_config.json, video_preprocessor_config.json multimodal processor
chat_template.jinja Qwen chat template
Inference

Requires transformers >= 5.8.1 (the qwen3_5 architecture is very new). Recommended sampling: temperature=0.6, top_p=0.95, top_k=20.

from transformers import AutoModelForImageTextToText, AutoTokenizer

model_id = "andrevp/Ornith-1.0-9B-Heretic-Uncensored"

tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForImageTextToText.from_pretrained(
model_id, dtype="bfloat16", device_map="auto"
)

messages = [{"role": "user", "content": "Write a Python function is_prime(n). Keep it short."}]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(text, return_tensors="pt").to(model.device)

generated = model.generate(
**inputs, max_new_tokens=512, do_sample=True,
temperature=0.6, top_p=0.95, top_k=20,
)
output = generated[0][inputs.input_ids.shape[1]:]
content = tokenizer.decode(output, skip_special_tokens=True)

The reply contains a ... reasoning block followed by the answer.

Serve with vLLM / SGLang the same way as the base Ornith-1.0-9B (see the base model card). Note the model has built-in MTP (mtp_num_hidden_layers: 1); with a recent vLLM/SGLang build speculative decoding can be enabled.
⚠️ Usage warnings (abliterated / uncensored)

This model is an abliterated (uncensored) derivative — its refusal direction has been removed. The standard warnings apply:

Risk of sensitive/controversial outputs — safety filtering is significantly reduced.
Not suitable for all audiences — outputs may be inappropriate for public settings, underage users, or high-security applications.
Legal & ethical responsibility — ensure your usage complies with local laws. You are solely responsible for any consequences.
Research / experimental use recommended — avoid unmonitored production or public-facing deployment.
No default safety guarantees — this model has not undergone rigorous safety optimization. The uploader bears no responsibility for any consequences arising from its use.

Credits

Original model: deepreinforce-ai/Ornith-1.0-9B by DeepReinforce (MIT, agentic coding family)
Abliteration method: Directional ablation — Arditi et al. 2024 ("Refusal in LLMs is mediated by a single direction"), with implementation notes from Maxime Labonne's abliteration guide and FailSpy's ortho cookbook.

Please donate

If this is useful, please donate — it helps fund more open model releases and quantizations:

BTC: bc1q6xxf0j3e7zn52cqrprc6gplql225wj8mnq75yw

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-09-05Update README.md4d2a1be5.4 KB
    Loading...
  2. 2026-07-29Create README.md892d5c75.3 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration