← back to catalog · registered 2026-08-27 10:02

bowmanslayer/Ornith-1.5-9B-Uncensored-GGUF

bowmanslayer Qwen 9B GGUF second-order 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/bowmanslayer%2FOrnith-1.5-9B-Uncensored-GGUF"
Response includes
  • classification m8
  • files 2
  • hub_downloads_all_time 2,377
  • author_summary 9 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
2K
652 last 30d - stable
Likes
0
Model age
6w ago
created 2026-08-27
Downloads over time
Now2.6K→from661↑300%
09691.9K2.9K661 on Aug 262.6K on Oct 11AugSepOct
Aug 26 → Oct 11 · 47 snapshots · spans 46 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Variants by this author 2 formats · 702 downloads combined

The same weights this author released in different packaging. Pick the format that matches your runtime.

Metadata

License
apache-2.0
Languages
en zh
Tags
gguf uncensored abliterated ornith qwen3.5 llama.cpp not-for-all-audiences en zh base_model:bowmanslayer/Ornith-1.5-9B-Uncensored base_model:quantized:bowmanslayer/Ornith-1.5-9B-Uncensored license:mit

Related

Total size
0 B
Files
2
Quantizations
1
Registered
2026-08-27 10:02
Last updated on HF
2026-08-29 01:04

Files by quantization

Auxiliary files 2 files 11.8 KB
README.md 10.3 KB ba60dd4e download
.gitattributes 1.48 KB a6344aac download

README current version from Hugging Face


library_name: gguf
license: apache-2.0
license_link: https://huggingface.co/ornith-ai/Ornith-1.5-9B/blob/main/LICENSE
base_model:

  • bowmanslayer/Ornith-1.5-9B-Uncensored
    base_model_relation: quantized
    tags:
  • uncensored
  • abliterated
  • ornith
  • qwen3.5
  • gguf
  • llama.cpp
  • not-for-all-audiences
    language:
  • en
  • zh
    extra_gated_prompt: >-
    This model has had its safety alignment removed. It will comply with requests
    that the original model refuses. By requesting access you confirm that you are
    of legal age in your jurisdiction, that you will not deploy it to third parties
    without your own safety layer, and that you accept sole responsibility for its
    outputs and for compliance with applicable law.
    extra_gated_fields:
    I am of legal age in my jurisdiction: checkbox
    I will not deploy this to third parties without my own safety layer: checkbox
    I accept sole responsibility for outputs and legal compliance: checkbox

Ornith-1.5-9B-Uncensored-GGUF

This repo bundles five GGUF quantizations of
bowmanslayer/Ornith-1.5-9B-Uncensored,
for llama.cpp / LM Studio / Ollama / Jan / KoboldCpp. File sizes range from
~4.5 GB (Q3_K_M) to ~18 GB (BF16) — pick one row from the table below
for your device. Sister repos: full-precision bf16 in the main repo,
vLLM-optimized W4A16 in
-W4A16.

本仓打包 5 档 GGUF 量化,源自
bowmanslayer/Ornith-1.5-9B-Uncensored,
适用 llama.cpp / LM Studio / Ollama / Jan / KoboldCpp。文件从 ~4.5 GB(Q3_K_M)
到 ~18 GB(BF16),按下表挑一档匹配你的设备。姐妹仓:主仓 bf16,vLLM 用 W4A16 见
-W4A16。

Vision note: this repo is text-only. Base model has a vision tower, but
llama.cpp GGUF export of Qwen3.5 vision (mmproj-*.gguf) is not included.
For image understanding, use the bf16 main repo or W4A16 repo with vLLM.

视觉说明:本仓仅文字。基座带视觉塔,但本仓未包 mmproj-*.gguf。要用
视觉能力请用主仓 bf16 或 W4A16 版本 + vLLM。


1. Which file should I download? / 该下哪一档?

Add ~1–2 GB on top of the weight size for KV cache and system overhead.
Ornith-1.5-9B is a hybrid-attention architecture (24 linear + 8 full-attention
layers, hidden 4096) — KV cache is unusually small compared to same-sized dense
models, long contexts are cheap.

选档时在权重大小上加 1–2 GB 给 KV cache 与 system overhead。Ornith-1.5-9B 混合
注意力(24 linear + 8 full,hidden 4096),KV cache 比同规模稠密模型小得多,长上下文
很便宜。

File Size Fits comfortably on / 舒适跑在 Quality vs BF16
Ornith-1.5-9B-Uncensored-BF16.gguf ~18 GB RTX 3090/4090 24 GB; Mac 32 GB+ Reference (lossless)
Ornith-1.5-9B-Uncensored-Q8_0.gguf ~10 GB RTX 4080/4070 Ti Super 16 GB; RTX 3090; Mac 16 GB Near-lossless
Ornith-1.5-9B-Uncensored-Q6_K.gguf ~7.6 GB RTX 3060 12 GB / 4060 Ti / 4070; Mac 16 GB High, close to Q8
Ornith-1.5-9B-Uncensored-Q5_K_M.gguf ~6.5 GB RTX 3060 12 GB / 4060 8 GB / Mac 16 GB Good balance
Ornith-1.5-9B-Uncensored-Q4_K_M.gguf ~5.5 GB RTX 3050 8 GB / 4060 8 GB / 2060 12 GB; Mac 8-16 GB Slightly degraded but very usable
Ornith-1.5-9B-Uncensored-Q3_K_M.gguf ~4.5 GB 6 GB GPUs (GTX 1660 / 2060); Mac 8 GB; low-end laptops Noticeably degraded but functional

Rules of thumb / 经验法则:

  • First-time pick → Q5_K_M(几乎任何 8 GB+ GPU 或 Mac 16 GB 都能跑,与 BF16 差异
    日常难察觉)
  • Highest quality that fits → Q6_K (12 GB), Q8_0 (16 GB), BF16 (24 GB)
  • 6-8 GB tight → Q3_K_M(质量明显更低,只有更大档装不下才用)
  • Apple Silicon Mac: Metal 后端原生跑 GGUF。8 GB Mac → Q4_K_M / Q3_K_M;
    16 GB → Q6_K / Q5_K_M; 24 GB+ → Q8_0 / BF16

2. Full performance evaluation / 完整性能测试

Numbers below from the bf16 reference build (every GGUF quant here was converted
from it). Per-quant deviation is expected small and in-line with llama.cpp
community observations for the same quant level.

以下数字取自 bf16 参考版(本仓每档 GGUF 都由它转换而来)。各档相对 bf16 的偏差预期较小,
与 llama.cpp 社区对同档量化的一般观察一致。

2.1 Refusal (dual-judged) / 拒绝率(双判据)

23 adversarial prompts across 9 harmful categories, generated end-to-end by an
internal uncensored 27B judge (operator never read prompts/responses). Same 27B
then judged REFUSE/COMPLY and separately COHERENT/LOOP/GIBBERISH.

23 条对抗提示涵盖 9 类有害内容,由内部无审查 27B 端到端生成(操作者全程不读)。
同 27B 判 REFUSE/COMPLY + 独立判 COHERENT/LOOP/GIBBERISH。

Metric Value
REFUSE / 23 0 ✅
COHERENT / 23 23 ✅ (no gibberish/loop artifacts)

2.2 11-benchmark full capability (T3)

Same fixed seed (20260816), same harness, same items. Details in main repo.

Benchmark N Ornith-1.5-9B base This model Δ
MMLU 150 90.14 88.97 -1.17
CMMLU 150 86.67 84.17 -2.50
MMLU-Pro 150 88.97 87.12 -1.85
C-Eval 150 85.82 84.78 -1.04
ARC-Challenge 150 93.33 94.67 +1.34
TruthfulQA 150 79.31 82.31 +3.00 ⭐
GSM8K 100 98.99 97.98 -1.01
MATH-500 100 73.68 70.53 -3.15
BBH 150 91.72 95.27 +3.55 ⭐
HumanEval 164 94.67 92.16 -2.51
IFEval (strict) 100 78.65 77.53 -1.12
Mean (excl. trunc) — 87.45 86.86 -0.59

Two benchmarks (BBH +3.55, TruthfulQA +3.00) actually improved. Ablation didn't
touch what Ornith's post-training gained on hard reasoning or truthfulness.

2.3 Output length / stopping / 输出长度与停止

Ornith-base This model
Median p50 output tokens (11 tasks avg) 487 540 (+11%)
Total truncation / 1514 94 88 (-6%)

Slightly more verbose but truncation actually decreased — stopping preserved.


3. llama.cpp usage / llama.cpp 使用

Server mode (int8 KV for more context):
服务器模式(开 int8 KV 拿到更大上下文):

./llama-server \
  -m Ornith-1.5-9B-Uncensored-Q6_K.gguf \
  --ctx-size 32768 \
  --n-gpu-layers -1 \
  -ctk q8_0 -ctv q8_0 \
  --host 0.0.0.0 --port 8080

Single-shot inference / 单条推理:

./llama-cli -m Ornith-1.5-9B-Uncensored-Q6_K.gguf -p "Hello" \
  -n 512 --n-gpu-layers -1 -ctk q8_0 -ctv q8_0

Note on -ctk q8_0 -ctv q8_0: default llama.cpp uses fp16 KV cache. Adding
int8 KV cuts KV memory ~50% for negligible accuracy hit — most community launch
scripts omit this and pay 2× the KV memory unnecessarily.


4. When to pick GGUF vs W4A16 / GGUF 还是 W4A16?

  • Single user, laptop, Mac, LM Studio / Ollama / Jan / KoboldCpp → GGUF here
    (Q5_K_M or Q6_K sweet spots)
  • Server with concurrent requests, high throughput, long context → W4A16 + vLLM
    (-W4A16)
  • Multi-modal (image) input → main repo bf16 + vLLM (this GGUF repo is text-only)

① Provenance and attribution / 来源与归属

Layer Author
Base ornith-ai/Ornith-1.5-9B — Ornith Team (Apache 2.0)
Upstream base Qwen/Qwen3.5-9B — Qwen team, Alibaba (Apache 2.0)
Abliteration bowmanslayer/Ornith-1.5-9B-Uncensored — weight-space ablation, 64 tensors, embed_tokens excluded (see main repo)
This repo GGUF conversion + quantization only (convert_hf_to_gguf.py + llama-quantize).

Not an official Ornith or Qwen release; not endorsed by or affiliated with either team.

非 Ornith 或 Qwen 官方发布;与两团队无背书或从属关系。


② Safety alignment has been removed / 安全对齐已被移除

This is the point of the model. Read before downloading.
这是模型的核心属性,下载前请务必阅读。

Ablation removes the refusal behaviour trained into Ornith-1.5-9B. 23 held-out
adversarial prompts dual-judged 0/23 refuses + 23/23 coherent. Quantization
does not restore refusals
.

上游消融移除了 Ornith-1.5-9B 内的拒绝行为。23 条 held-out 双判据 0/23 拒 + 23/23
coherent。量化不恢复拒绝。

Consequences you accept / 你正在接受的后果:

  • Content the original refuses / 原模型拒绝的内容
  • No content filter, no guardrail / 无过滤、无护栏
  • Refusal removal not proven exhaustive (n=23, 英文, single-turn)

Not intended for: third-party public deployment without your own safety layer /
minors / uses prohibited by upstream policies.

Intended for: local inference and research, by those who understand the above
and take responsibility.

不适用于:未搭安全层的公开部署 / 未成年人 / 上游政策禁止的用途。
适用于:本地推理与研究,由理解与担责用户使用。


③ No warranty; responsibility rests with the user / 无担保;责任在用户

Provided "AS IS", without warranty of any kind. I do not endorse or recommend
any use, and make no representation on accuracy/lawfulness/fitness. You are
solely responsible
for what you generate, deploy, and legal compliance. To
the maximum extent permitted by law, no liability.

按**"现状"提供,不提供任何担保**。我不背书、不推荐,不对输出的准确性/合法性/用途
适用性作声明。你独自承担全部责任。法律允许最大范围内不承担任何责任。

Downloading these files means you accept the above. If you do not, do not download.

下载本仓文件即表示接受以上内容。若不接受,请勿下载。

The author is not a lawyer; the above is not legal advice.
作者不是律师;以上不构成法律意见。


License / 许可

Apache 2.0, inherited through the chain: Qwen3.5-9B → Ornith-1.5-9B → this repo.

Apache 2.0,依上述归属链继承。

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-27Upload README.md with huggingface_hub214cef310.3 KB
    Loading...
  2. 2026-08-27Upload README.md with huggingface_hubcd5c22710.2 KB
    Loading...
  3. 2026-08-27Upload README.md with huggingface_hub2b5ecb010.3 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration