← back to catalog · registered 2026-08-22 20:02

cognitivers/Ornith-1.5-35B-A3B-Abliterated-12GB-GGUF

cognitivers 35B GGUF MoE multimodal second-order 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/cognitivers%2FOrnith-1.5-35B-A3B-Abliterated-12GB-GGUF"
Response includes
  • classification m8
  • files 7
  • hub_downloads_all_time 1,149
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
1K
573 last 30d - stable
Likes
1
Model age
7w ago
created 2026-08-22
Downloads over time
Now1.4K→from348↑296%
2976911.1K1.5K348 on Aug 261.4K on Oct 11AugSepOct
Aug 26 → Oct 11 · 47 snapshots · spans 46 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
mit
Quantizations
IQ2
Tags
gguf imatrix qwen3.5-moe qwen3_5_moe moe abliterated uncensored low-vram iq2 vision multimodal image-text-to-text

Related

Total size
31.6 GB
Files
7
Quantizations
3
Registered
2026-08-22 20:02
Last updated on HF
2026-08-22 19:12

Files by quantization

IQ2 3 files 31.4 GB
Ornith-1.5-35B-A3B-Abliterated-IQ2_M.gguf 11.5 GB 4aafa819 download
Ornith-1.5-35B-A3B-Abliterated-IQ2_S.gguf 10.9 GB 894607b7 download
Ornith-1.5-35B-A3B-Abliterated-IQ2_XXS.gguf 9.02 GB bc79296c download
F16 1 file 858 MB
mmproj-Ornith-1.5-35B-A3B-Abliterated-F16.gguf 858 MB f7cfbb1e download
Auxiliary files 3 files 183 MB
imatrix.gguf 183 MB 1646e297 download
README.md 5.01 KB 1bfb0bb3 download
.gitattributes 1.96 KB 6a833729 download

README current version from Hugging Face


license: mit
base_model: PocketAiHub/Ornith-1.5-35B-A3B-Abliterated-GGUF
base_model_relation: quantized
library_name: gguf
pipeline_tag: image-text-to-text
tags:

  • gguf
  • imatrix
  • qwen3.5-moe
  • qwen3_5_moe
  • moe
  • abliterated
  • uncensored
  • low-vram
  • iq2
  • vision
  • multimodal
    quantized_by: cognitivers

Ornith-1.5-35B-A3B-OBLITERATED — low-bit imatrix GGUFs (runs on 12 GB VRAM)

First working IQ2 imatrix quantizations of the abliterated Ornith-1.5-35B-A3B, sized so a 12 GB GPU runs this 35B mixture-of-experts fully on the card. It is an qwen3_5_moe MoE (35B total, only ~3B active per token) — so it is fast — and it keeps vision (image input via the shared mmproj).

Until now the abliterated model only existed as Q4_K_M (19.7 GB) and larger — nothing that fits a mainstream 12 GB card. These are the first IQ2-class GGUFs of it.

Size vs 12 GB budget

Why these didn't exist

MoE models at 2-bit need a complete importance matrix over all experts or llama-quantize refuses ("the result will be garbage"). This model has 256 experts, top-8 routing — rarely-activated experts are easy to miss. We computed a fresh imatrix on the abliterated weights (bartowski calibration_datav3, -c 512 --parse-special, ~100% executed-tensor coverage, imatrix.gguf ships in this repo) and tuned the per-tensor mix so 2-bit doesn't collapse: ffn_down_exps pinned to iq3_xxs (the sensitive projection), output tensor q6_k, token embeddings q4_k.

Benchmarks (measured, not estimated)

Perplexity & KL-divergence on wikitext-2 (ctx 512), HellaSwag over 400 tasks, against the abliterated Q8_0 as baseline. Speeds via llama-bench on an L40S; on a 12 GB card (RTX 3060 class) tok/s is lower but still high thanks to the 3B active path.

File Size PPL ΔPPL vs Q8_0 Mean KLD Same-top-p HellaSwag (400) Target cards
IQ2_M 11.53 GB 9.92 +20.5 % 0.320 75.8 % 77.75 % (−3.0) 12 GB (best quality, headless / ctx ≤4k with desktop)
IQ2_S 10.88 GB 10.25 +24.5 % 0.372 74.3 % 76.25 % (−4.5) 12 GB, comfortable + context headroom
IQ2_XXS 9.02 GB 12.34 +49.9 % 0.593 68.2 % 71.75 % (−9.0) 8–10 GB — usable, clearly degraded
Q8_0 (reference) 34.4 GB 8.23 — — — 80.75 % not in this repo

Quality kept at 2-bit

IQ2_M drops just 3 points of HellaSwag vs the full model (77.75 vs 80.75) in 11.5 GB. That is the pick on a 12 GB card. IQ2_S trades a little quality for context/desktop headroom; IQ2_XXS exists so 8–10 GB cards can run it at all (visibly degraded).

How to run

# llama.cpp — text
llama-server -hf cognitivers/Ornith-1.5-35B-A3B-Abliterated-12GB-GGUF:IQ2_M \
  -ngl 999 -c 8192 -fa on -ctk q8_0 -ctv q8_0

# Ollama
ollama run hf.co/cognitivers/Ornith-1.5-35B-A3B-Abliterated-12GB-GGUF:IQ2_M

Vision (image input): download the language file and mmproj-Ornith-1.5-35B-A3B-Abliterated-F16.gguf (~0.9 GB extra VRAM), then:

llama-mtmd-cli -m Ornith-1.5-35B-A3B-Abliterated-IQ2_M.gguf \
  --mmproj mmproj-Ornith-1.5-35B-A3B-Abliterated-F16.gguf \
  --image photo.jpg -p "Describe this image." -ngl 999

Also works in LM Studio, Jan and koboldcpp. The native MTP speculative head is not included (it wasn't in the source GGUF).

Provenance & reproducibility

  • Source: Q8_0 from PocketAiHub/Ornith-1.5-35B-A3B-Abliterated-GGUF (Q8_0 ≈ lossless; requantized with --allow-requantize). Original model: ornith-ai/Ornith-1.5-35B-A3B; refusal-direction edit by PocketAI Model Lab.
  • Importance matrix: imatrix.gguf in this repo — computed by us on the abliterated weights (llama.cpp, calibration_datav3, -c 512 -b 512 --parse-special).
  • Quantized with llama-quantize: base type + --tensor-type ffn_down_exps=iq3_xxs --output-tensor-type q6_k --token-embedding-type q4_k --imatrix imatrix.gguf.

At 2-bit the loss is real (see the table). These exist to make a 35B MoE runnable on mainstream GPUs, not to replace Q4+ if your hardware fits it.

Safety / uncensored

This is an abliterated (uncensored) model: its learned refusals were suppressed, so it will comply with requests an instruct model would decline, and can produce harmful, illegal, or dangerously wrong output more readily. Quantization partially restores some refusals; this is not a safety or truthfulness guarantee. Evaluate and constrain it for your use case.

Credits

  • ornith-ai — original Ornith-1.5-35B-A3B.
  • PocketAiHub / Pliny — abliteration + reference GGUF conversion.
  • bartowski — calibration_datav3.
  • llama.cpp — quantization + qwen3_5_moe support.

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-22align HellaSwag framing (-3 pts)6a27a4b5 KB
    Loading...
  2. 2026-08-22add model card347d11b5 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration