← back to catalog · registered 2026-08-22 13:56

lemuralabs/Qwen3.5-122B-A10B-Abliterated-MLX-3

lemuralabs Qwen 122B MoE multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/lemuralabs%2FQwen3.5-122B-A10B-Abliterated-MLX-3"
Response includes
  • classification m1
  • files 26
  • hub_downloads_all_time 2,449
  • author_summary 31 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
2K
104 last 30d - cooling
Likes
1
Model age
5mo ago
created 2026-05-01
Downloads over time
Now2.5K→from2K↑22%
2K2.2K2.3K2.5K2K on Aug 52.5K on Oct 11AugSepOct
Aug 5 → Oct 11 · 50 snapshots · spans 67 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en
Tags
mlx safetensors qwen3_5_moe mlx-vlm qwen qwen3 qwen3.5 qwen3.5-122b a10b mixture-of-experts apple-silicon vision-language

Related

Total size
57.3 GB
Files
26
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-05 19:15

Files by quantization

Auxiliary files 26 files 57.3 GB
model-00001-of-00012.safetensors 4.96 GB 9bda28c0 download
model-00009-of-00012.safetensors 4.96 GB ea71d4da download
model-00011-of-00012.safetensors 4.96 GB fb33f085 download
model-00007-of-00012.safetensors 4.96 GB 8a819078 download
model-00003-of-00012.safetensors 4.96 GB d38ab09f download
model-00005-of-00012.safetensors 4.96 GB 846ba579 download
model-00010-of-00012.safetensors 4.82 GB 016cbb85 download
model-00004-of-00012.safetensors 4.82 GB cb7ef0cc download
model-00006-of-00012.safetensors 4.77 GB 84ea1ff1 download
model-00008-of-00012.safetensors 4.77 GB 87bb4bbf download
model-00002-of-00012.safetensors 4.77 GB 0a1df447 download
model-00012-of-00012.safetensors 3.54 GB bb6a8c7b download
tokenizer.json 19.1 MB 87a7830d download
vocab.json 6.41 MB 0aa0ce06 download
model.safetensors.index.json 262 KB 9dbdbbfe download
logo.png 18.6 KB a9400259 download
README.md 8.62 KB ad77b22e download
chat_template.jinja 7.58 KB 604313d9 download
config.json 3.97 KB 22f4628c download
mlztq_manifest.json 1.71 KB 694bf0dd download
processor_config.json 1.28 KB a3be3e79 download
tokenizer_config.json 1.11 KB a068e246 download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
.gitattributes 269 B 4d35f84a download
generation_config.json 244 B 85b45ab4 download

README current version from Hugging Face


license: apache-2.0
base_model: Chompa1422/Qwen3.5-122B-A10B-abliterated
base_model_relation: quantized
model_name: Qwen3.5-122B-A10B-Abliterated-MLX-3
library_name: mlx
pipeline_tag: image-text-to-text
language:

  • en
    tags:
  • mlx
  • mlx-vlm
  • safetensors
  • qwen
  • qwen3
  • qwen3.5
  • qwen3_5_moe
  • qwen3.5-122b
  • a10b
  • mixture-of-experts
  • apple-silicon
  • vision-language
  • multimodal
  • image-text-to-text
  • quantized
  • 3-bit
  • group-size-32
  • vision-quantized
  • abliterated
  • refusal-removal
  • research
  • lm-studio
    widget:
  • text: "Describe this image in detail."
  • text: "Read the text in this document image."
  • text: "Solve: Janet has 16 eggs, eats 3, uses 4, and sells the rest at $2 each. How much does she make?"
    model-index:
  • name: Qwen3.5-122B-A10B-Abliterated-MLX-3
    results:
    • task:
      type: image-text-to-text
      name: Vision-language inference
      dataset:
      name: MLZ VLM Retention Gate
      type: mlz-vlm-retention-gate
      metrics:
      • type: accuracy
        name: Quant health
        value: 1.0
      • type: accuracy
        name: Text / reasoning
        value: 1.0
      • type: accuracy
        name: OCR / document
        value: 1.0
      • type: accuracy
        name: Vision general
        value: 0.5
      • type: accuracy
        name: Overall gate
        value: 0.808

Lemura Labs

Qwen3.5-122B-A10B-Abliterated-MLX-3

Format Task Params Type License

Qwen3.5-122B-A10B-Abliterated-MLX-3 is an Apple Silicon-oriented MLX quantization of the abliterated Qwen3.5 122B-A10B vision-language model.

Publisher: Lemura Labs
Hugging Face organization: Lemura Labs

This release is the 3-bit / group-size-32 build selected from our vision-quantized candidate run. This model quantizes both the language side and eligible multimodal/vision modules.

Source And Credits

This quant was produced from the local full-precision abliterated model derived from:

Thank you to the Qwen team for the base model, to Chompa1422 for publishing the abliterated source model that made this quantization work possible, and to Pliny the Liberator for the broader abliterated-model research culture that inspired this release.

Quantization

Field Value
Runtime format MLX / MLX-VLM safetensors
Quantization method MLZTQ 0.2 custom MLX-VLM affine quantization
Weight bits 3
Group size 32
Mode affine
Quant profile mlztq-0.2-visionq_w3-g32
Quant predicate visionq_w3
Language weights Quantized
Vision/multimodal weights Quantized where MLX shape constraints allow
KV cache TurboQuant runtime contract recorded; model weights are MLX affine quantized

Tensor spot checks from the selected artifact:

Tensor dtype Shape Meaning
language_model.model.layers.0.linear_attn.in_proj_qkv.weight uint32 (12288, 288) Language side quantized
language_model.model.layers.0.linear_attn.in_proj_qkv.scales bfloat16 (12288, 96) Language affine scales
vision_tower.blocks.0.attn.qkv.weight uint32 (3456, 108) Vision side quantized
vision_tower.blocks.0.attn.qkv.scales bfloat16 (3456, 36) Vision affine scales

File Details

Item Value
Safetensor shards 12
Indexed tensor payload 61,488,610,272 bytes
Local disk footprint 57 GiB
Decimal payload size 61.49 GB
Model index model.safetensors.index.json
Quant manifest mlztq_manifest.json
Processor files processor_config.json, preprocessor_config.json, video_preprocessor_config.json
Tokenizer files tokenizer.json, tokenizer_config.json, vocab.json, chat_template.jinja

Benchmarks

Benchmarks were run with the local MLZ deterministic VLM retention gate on 2026-05-02. Raw JSON, CSV, and Markdown benchmark artifacts are included under benchmarks/results/.

Bucket Correct Total Accuracy
Quant health 4 4 100.0%
Text / reasoning 6 6 100.0%
OCR / document 6 6 100.0%
Vision general 5 10 50.0%
Overall gate 21 26 80.8%

Benchmark coverage:

Bucket Sources
Quant health Safetensor load, vision path availability, text canary, vision canary
Text / reasoning MMLU-Pro, MMLU, GSM8K
OCR / document ChartQA, DocVQA, local PDF pages with images/text
Vision general RealWorldQA, AI2D, MMMU Accounting, MMMU Biology, MMMU Physics

Candidate comparison from the same run:

Candidate Weight bits / group Payload GB Quant health Text / reasoning Vision general OCR / document Overall Status
visionq-w6g32 6 / 32 107.40 100.0% 83.3% 50.0% 100.0% 76.9% kept locally
visionq-w5g32 5 / 32 deleted 100.0% 100.0% 40.0% 83.3% 73.1% pruned
visionq-w4g32 4 / 32 76.79 100.0% 83.3% 40.0% 100.0% 73.1% kept locally
visionq-w3g32 3 / 32 61.49 100.0% 100.0% 50.0% 100.0% 80.8% selected
visionq-w2g32 2 / 32 deleted 100.0% 16.7% 20.0% 50.0% 38.5% pruned

Inference Engines And Apps

This is an MLX / MLX-VLM safetensors repository for Apple Silicon. It is not a GGUF, AWQ, GPTQ, EXL2, or Transformers fp16 repository.

Engine / app Status for this repository Notes
MLX-VLM Intended reference path Best target for image-text inference because this artifact keeps the VLM processor files and MLX-VLM tensor layout. Requires a loader/runtime version with Qwen3.5 MoE VLM support and affine quantized multimodal weights.
LM Studio on Apple Silicon Intended app target, runtime-dependent LM Studio's unified MLX engine uses mlx-lm for text generation and mlx-vlm for vision embeddings. Use a recent LM Studio build with MLX support; macOS 14+ is required for MLX models according to LM Studio's system requirements.
Custom MLX Python runtimes Supported if they implement this architecture Works for runtimes that can read MLX safetensors, Qwen3.5 MoE configs, the chat template, and MLX affine quantized language + vision tensors.
mlx-lm alone Text-side only / not sufficient for full VLM use mlx-lm is useful in the MLX ecosystem, but full image input needs the VLM path and processor stack.
Hugging Face Transformers / vLLM / TGI Not directly loadable These engines do not load this MLX quantized artifact directly. Use the original/full-precision model or produce a separate backend-specific quant.
Ollama / llama.cpp / KoboldCpp Not directly loadable These generally expect GGUF for local quantized inference. This repo is MLX safetensors, not GGUF.

Practical expectation: use this model on high-memory Apple Silicon Macs through MLX-VLM-compatible tooling. For image input, use PNG, JPEG, WebP, and PDF/image workflows supported by the serving app or preprocessing pipeline.

Research And Safety Notice

Why is this model Abliterated?

This model is intended for research, local experimentation, red-team evaluation, and authorized security testing. It may produce content that aligned models normally refuse. Users are responsible for applying appropriate safeguards and complying with laws and platform policies.

This is an abliterated model released for research and development, model-behavior analysis, authorized security testing, and experimentation with local Apple Silicon inference. Abliterated models may respond differently from aligned instruction models. Users are responsible for complying with applicable laws, platform policies, and safety requirements. The authors and uploaders are not responsible for misuse, harm, or unlawful deployment.

Reproducibility

The included mlztq_manifest.json records the source path, quantization recipe, weight format, vision quantization policy, and runtime contract used for this artifact. The benchmark files under benchmarks/results/ record the exact gate rows used to choose this model.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-05Initial commit017af7b8.6 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration