library_name: mlx
license: mit
license_link: https://huggingface.co/ornith-ai/Ornith-1.5-35B-A3B/blob/main/LICENSE
base_model: ornith-ai/Ornith-1.5-35B-A3B
pipeline_tag: image-text-to-text
tags:
- mlx
- mtplx
- vision
- multimodal
- mixture-of-experts
- speculative-decoding
- abliterated
Ornith 1.5 35B-A3B Abliterated Attention8 + BF16 Recurrence Vision MTPLX
Experimental MLX hybrid quant of ornith-ai/Ornith-1.5-35B-A3B, with a quant-native refusal-direction edit, the exact-parent BF16 vision tower, and the model's native BF16 MTP sidecar for MTPLX speculative decoding.
This repository is intended for Apple Silicon. It is not a GGUF, GPTQ, or generic Transformers checkpoint.
Precision layout
- Default affine body: 4-bit, group size 32.
- Attention, embeddings, LM head, routing/gating, and other high-sensitivity modules: 8-bit affine, group size 64.
- Recurrent A/B state inputs: BF16.
- Vision tower: BF16, 333 tensors.
- Native MTP sidecar: BF16, 785 tensors.
- Architecture: Qwen3.5 hybrid MoE, 40 layers, 256 experts, approximately 3B active parameters per token.
The body contains 260 q8 module overrides, 192 q4 modules, and 60 BF16 recurrence exclusions.
Abliteration
The refusal direction was derived freshly from the exact quantized parent Shiftedx/ornith-1.5-35b-a3b-attention8-bf16recurrence-vision-mtplx.
- Method: residual-direction weight orthogonalization.
- Strength: 1.5.
- Targets: attention output, shared-expert down projection, and switched-expert down projection.
- Scope: all 40 layers, 120 edited modules.
- Quant-native edit: each selected module was dequantized, edited in float32, column-norm preserved, and requantized once to its original mode.
- Direction SHA-256:
0e9246e95f1dda921c9e0672de38abf076ac925df6932e0110e3eaa5bc20cb32.
“Uncensored” or “abliterated” is not a universal behavior guarantee. On the disjoint local held-out suite, the exact parent refused 100% of the targeted requests; this strength-1.5 candidate refused 0%, retained 0% benign refusal, and passed 100% of the utility checks. Strength 1.0 was too weak and strength 2.0 caused a utility regression.
Local qualification
- Structural inspection: 1,970 indexed tensors across six shards; no warnings or failures.
- Vision OCR smoke: exact output
HUNTER. - MTPLX inspection: architecture recognized, primary gate passed, 785/785 MTP tensors, zero missing or extra keys.
- OpenAI-compatible code smoke at MTP depth 1: 3/3 passed.
- Hermes Agent provider smoke: passed.
MTP depth sweep
Measured on an Apple M4 Max with 64 GiB unified memory using MTPLX's cold-long-code-192 suite, 256 generated tokens, seed 42, identical sampling, and fans in automatic mode:
| Mode | Decode tok/s | Versus AR | Acceptance by depth |
|---|---|---|---|
| AR | 70.316 | 1.000× | — |
| D1 | 104.755 | 1.490× | 93.18% |
| D2 | 89.022 | 1.266× | 97.52%, 14.17% |
| D3 | 74.603 | 1.061× | 95.87%, 14.05%, 0.83% |
Depth 1 is recommended. Acceptance collapses after the first drafted token, so D2 and D3 add more verification work than useful accepted tokens.
These are local runtime measurements, not general cross-platform benchmark claims.
MTPLX serving
The public MTPLX runtime contract is still classified as unverified, so explicit opt-in is required:
mtplx quickstart \
--model Shiftedx/ornith-1.5-35b-a3b-abliterated-attention8-bf16recurrence-vision-mtplx \
--model-id ornith-1.5-35b-a3b-abliterated \
--mtp --depth 1 \
--profile sustained \
--reasoning on --reasoning-parser qwen3 \
--tool-prompt-mode native \
--unsafe-force-unverified --yes
The resulting server exposes an OpenAI-compatible API and accepts PNG, JPEG, and WebP image inputs.
Standard MLX vision use
The main body and processor files remain compatible with MLX-VLM. The nested mtp/weights.safetensors sidecar is for MTPLX and is ignored by ordinary MLX-VLM generation.
python -m mlx_vlm.generate \
--model Shiftedx/ornith-1.5-35b-a3b-abliterated-attention8-bf16recurrence-vision-mtplx \
--image image.jpg \
--prompt "Describe this image."
Lineage
- Upstream model:
ornith-ai/Ornith-1.5-35B-A3B. - Pinned upstream revision:
fbb995a79eedd569a5edc5f2af9644c0fa1124fc. - Exact quantized parent:
Shiftedx/ornith-1.5-35b-a3b-attention8-bf16recurrence-vision-mtplx. - Abliteration direction and source hashes are recorded in
config.jsonanddirection.npz. - Local qualification metadata and depth-sweep results are recorded in
mtplx_runtime.json.
Limitations
- This is an experimental behavior edit and may change capabilities or safety behavior outside the evaluated suite.
- The model can produce inaccurate, unsafe, or objectionable content. Users are responsible for evaluation and deployment controls appropriate to their use case.
- The MTPLX package currently requires
--unsafe-force-unverified; this label reflects the runtime release contract, not a failed tensor or local inference check. - Maximum-context qualification was not performed for this abliterated variant.
License and attribution
The upstream model is MIT licensed. See the linked upstream license and model card for original training details, intended use, and attribution.
@misc{ornith_1_5,
title = {Ornith-1.5: From Self-Scaffolding to Self-Improvement},
url = {https://ornith.ai/ornith_1_5.html},
author = {Ornith Team},
year = {2026}
}