license: mit
base_model: huihui-ai/Huihui-Ornith-1.5-35B-A3B-abliterated
tags:
- qwen35moe
- abliterated
- uncensored
- rocmfp4
- halofpx
- gguf
- vision
library_name: gguf
Ornith-1.5-35B-A3B-abliterated-Q4_ROCmFPX_FAST
ROCmFPX quantized version of huihui-ai/Huihui-Ornith-1.5-35B-A3B-abliterated (locally requantized).
- Base model: Ornith-1.5-35B-A3B (Claude-style 35B MoE, 3B active) abliterated (refusal behavior removed, uncensored)
- Quantization: Q8_0 → ROCmFPX (
Q4_0_ROCMFP4_FAST, 4.25 bpw),--allow-requantize - Vision: ✅ bundled mmproj (clip + qwen3vl_merger)
⚠ Important: Format Notice
ROCmFPX is a proprietary quantization format of the halofpx (ROCmFPX) engine — upstream llama.cpp CANNOT load this file!
Load it with halofpx (register this GGUF in the halofpx registry, then POST /api/v1/load).
Files
| File | Size | sha256 |
|---|---|---|
| Ornith-1.5-35B-A3B-abliterated-Q4_ROCmFPX_FAST.gguf | ~17.6 GiB | see sha256.txt |
| mmproj-Ornith-1.5-35B-BF16.gguf | ~861 MiB | see sha256.txt |
Quantization Benchmarks
Measured on AMD Strix Halo (gfx1151), llama-bench, -p 256 -n 256 -t 16 -fa on -ngl 99, same prompt per row.
| Variant | tg256 (t/s) | pp256 (t/s) | Size | bpw |
|---|---|---|---|---|
| Ornith-1.5-35B-A3B-abliterated-Q4_ROCMFP4_FAST (this repo) | 73.0 | 829.7 | 17.6 GiB | 4.25 |
| Ornith-1.5-35B-A3B-abliterated-Q4_ROCMFP4 (base) | 37.8 | 661.7 | 21.7 GiB | 4.50 |
| Ornith-1.5-35B-A3B-abliterated-Q8_ROCMFPX | 47.6 | 813.5 | 34.2 GiB | 8.0 |
| Ornith-1.5-35B-A3B-abliterated-Q4_ROCMFP4_STRIX_LEAN | 42.3 | 719.2 | 17.7 GiB | 4.27 |
Why only Q4_ROCmFPX_FAST is published (other variants not uploaded):
Ornith-1.5-35B-A3B-abliterated-Q4_ROCMFP4(base, dual-scale): 48% slower than FAST with only marginal precision gain — rejectedOrnith-1.5-35B-A3B-abliterated-Q4_ROCMFP4_STRIX_LEAN: 72% slower than FAST — rejectedOrnith-1.5-35B-A3B-abliterated-Q8_ROCMFPX: near-lossless but 2× size and 35% slower — to be published later as the quality tier
Usage
# After registering in halofpx registry (example):
curl -X POST http://127.0.0.1:8010/api/v1/load \
-H "Authorization: Bearer ***" \
-d '{"model_id":"ornith-1.5-35b-abliterated","reasoning_mode":"off"}'
- Reasoning model:
reasoning_mode: offoutputs directly; MTP is a net loss for this model, keep it off - Context: up to 256K supported (measured same speed as 131K)
Disclaimer
This is an abliterated (uncensored) version and may produce outputs that do not conform to safety policies. Use at your own risk. No warranty is provided.
Ornith-1.5-35B-A3B-abliterated-Q4_ROCmFPX_FAST(ROCmFPX)
基于 huihui-ai/Huihui-Ornith-1.5-35B-A3B-abliterated 的 ROCmFPX 量化版(本地 requantize 产物)。
- 原模型:Ornith-1.5-35B-A3B(Claude 风格 35B MoE,3B 激活)的 abliterated 消融版(去拒绝行为,无审查)
- 量化:Q8_0 → ROCmFPX(
Q4_0_ROCMFP4_FAST,4.25 bpw),--allow-requantize - 视觉:✅ 附带 mmproj(clip + qwen3vl_merger)
⚠ 重要:格式说明
ROCmFPX 是 halofpx(ROCmFPX)引擎专属量化格式——上游 llama.cpp 无法加载此文件!
请使用 halofpx 加载(halofpx registry 添加此 GGUF 后 POST /api/v1/load)。
文件
| 文件 | 大小 | sha256 |
|---|---|---|
| Ornith-1.5-35B-A3B-abliterated-Q4_ROCmFPX_FAST.gguf | ~17.6 GiB | 见仓库 sha256.txt |
| mmproj-Ornith-1.5-35B-BF16.gguf | ~861 MiB | 见仓库 sha256.txt |
量化基准测试
测试环境:AMD Strix Halo(gfx1151),llama-bench,-p 256 -n 256 -t 16 -fa on -ngl 99,每行同一 prompt。
| 变体 | tg256 (t/s) | pp256 (t/s) | 大小 | bpw |
|---|---|---|---|---|
| Ornith-1.5-35B-A3B-abliterated-Q4_ROCMFP4_FAST(本仓库) | 73.0 | 829.7 | 17.6 GiB | 4.25 |
| Ornith-1.5-35B-A3B-abliterated-Q4_ROCMFP4(基础版) | 37.8 | 661.7 | 21.7 GiB | 4.50 |
| Ornith-1.5-35B-A3B-abliterated-Q8_ROCMFPX | 47.6 | 813.5 | 34.2 GiB | 8.0 |
| Ornith-1.5-35B-A3B-abliterated-Q4_ROCMFP4_STRIX_LEAN | 42.3 | 719.2 | 17.7 GiB | 4.27 |
仅发布 Q4_ROCmFPX_FAST 的原因(其他量化不上传):
Ornith-1.5-35B-A3B-abliterated-Q4_ROCMFP4(基础版,双 scale):比 FAST 慢 48%,精度提升有限——弃用Ornith-1.5-35B-A3B-abliterated-Q4_ROCMFP4_STRIX_LEAN:比 FAST 慢 72%——弃用Ornith-1.5-35B-A3B-abliterated-Q8_ROCMFPX:近无损但体积 2 倍、慢 35%——待后续上传(高质量档)
使用
# halofpx registry 注册后(例):
curl -X POST http://127.0.0.1:8010/api/v1/load \
-H "Authorization: Bearer ***" \
-d '{"model_id":"ornith-1.5-35b-abliterated","reasoning_mode":"off"}'
- 思考模型:
reasoning_mode: off直接输出;MTP 对该模型为净亏损,勿开 - 上下文:支持 256K(实测与 131K 速度持平)
免责声明
本模型为 abliterated(消融去审查)版本,可能生成不符合安全政策的输出。使用者自行承担全部责任。本仓库不提供任何保证。