← back to catalog · registered 2026-08-22 13:56

lemuralabs/Keye-VL-2.0-30B-A3B-uncensored-mlx-mxfp4

lemuralabs 31B MoE multimodal
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/lemuralabs%2FKeye-VL-2.0-30B-A3B-uncensored-mlx-mxfp4"
Response includes
  • classification m1
  • files 23
  • hub_downloads_all_time 548
  • author_summary 31 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
548
69 last 30d - stable
Likes
0
Model age
4mo ago
created 2026-06-01
Downloads over time
Now572→from377↑52%
367442517592377 on Aug 5572 on Oct 11AugSepOct
Aug 5 → Oct 11 · 50 snapshots · spans 67 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
mlx safetensors KeyeVL2 keye vision-language moe abliterated mxfp4 image-text-to-text conversational custom_code base_model:Kwai-Keye/Keye-VL-2.0-30B-A3B

Related

Total size
16.0 GB
Files
23
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-05 19:13

Files by quantization

Auxiliary files 23 files 16.0 GB
model-00001-of-00004.safetensors 4.96 GB dc882d39 download
model-00003-of-00004.safetensors 4.93 GB e033e28b download
model-00002-of-00004.safetensors 4.93 GB 1006fb14 download
model-00004-of-00004.safetensors 1.18 GB 7cf821f7 download
tokenizer.json 10.9 MB 85911670 download
vocab.json 2.65 MB 4783fe10 download
model.safetensors.index.json 145 KB 70b42236 download
modeling_keye_topk_mask_30ba3b.py 64.3 KB 4c2dc920 download
modeling_keye_vl_2.py 64.3 KB 4a6f114e download
image_processing_keye_vl_2.py 24.0 KB c490bfe1 download
processing_keye_vl_2.py 20.8 KB b2f5776f download
logo.png 18.6 KB a9400259 download
configuration_keye_vl_2.py 12.5 KB 03a6272a download
README.md 3.23 KB 297dfd4d download
config.json 3.04 KB 94036366 download
chat_template.jinja 2.77 KB 3c9bb5a0 download
.gitattributes 1.53 KB 52373fe2 download
added_tokens.json 1.03 KB dc0f3f34 download
processor_config.json 855 B befbc099 download
preprocessor_config.json 686 B 343a25dd download
tokenizer_config.json 604 B 99ca5332 download
special_tokens_map.json 477 B ca57839e download
generation_config.json 121 B a811a763 download

README current version from Hugging Face


license: apache-2.0
base_model: Kwai-Keye/Keye-VL-2.0-30B-A3B
tags:

  • keye
  • vision-language
  • moe
  • abliterated
  • mlx
  • mxfp4
    pipeline_tag: image-text-to-text
    library_name: mlx

Lemura Labs

Keye-VL-2.0-30B-A3B — Abliterated — MLX (MXFP4)

Format Task Params Type Quant License

MXFP4 (≈4.43 bpw) MLX build of the abliterated
Kwai-Keye/Keye-VL-2.0-30B-A3B,
for Apple Silicon. Quantized from
lemuralabs/Keye-VL-2.0-30B-A3B-uncensored
with mlx-vlm. ~16 GB on disk; runs in ~17 GB; ≈107 tok/s on an M4 Max.

Yes — This MLX build runs Keye coherently on a Mac. The original model uses a
CUDA-only sparse-attention indexer (SALightningIndexer) that is unstable on MPS — so
the stock model generates garbage via Transformers on Apple Silicon. This port runs the
mathematically-equivalent dense attention, which is coherent and fast on MLX.

Requirements — custom mlx-vlm model class

Keye is not yet in mainline mlx-vlm. This repo bundles the support module under
mlx_vlm_keye_support/keyevl2/. Install it:

pip install mlx-vlm
# copy the bundled module into your mlx-vlm install:
python - <<'PY'
import mlx_vlm, os, shutil
dst = os.path.join(os.path.dirname(mlx_vlm.__file__), "models", "keyevl2")
shutil.copytree("mlx_vlm_keye_support/keyevl2", dst, dirs_exist_ok=True)
print("installed keyevl2 ->", dst)
PY

Also register the prompt format (one line in mlx_vlm/prompt_utils.py): add
"keye_vl2": MessageFormat.LIST_WITH_IMAGE_FIRST, to the format map.

Usage

python -m mlx_vlm generate --model lemuralabs/Keye-VL-2.0-30B-A3B-uncensored-mlx-mxfp4 \
 --prompt "Describe this image." --image path/to/img.jpg --trust-remote-code

Notes

  • Quant: MXFP4, group size 32, 4.432 bpw (whole model, incl. vision tower).
  • Vision: the SigLIP tower + mlp_AR projector are included (quantized). Text gen is
    verified coherent; image understanding is functional but the packed-vision forward in
    this port is a first cut — report issues.
  • Abliterated (refusals reduced; see the base abliterated card for method/limits).
  • The text backbone reuses mlx-vlm's qwen3_vl_moe; the sparse sa_indexer is dropped.

Abliteration removes safety alignment; you are responsible for use.

Other variants of this model (public on Lemura Labs)

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-05Initial commitc3999aa3.2 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration