← back to catalog · registered 2026-08-22 13:56

lemuralabs/Qwen3.6-27B-Claude-Opus-Reasoning-Distill-v2-abliterated-6bit-mlx

lemuralabs Qwen 27B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/lemuralabs%2FQwen3.6-27B-Claude-Opus-Reasoning-Distill-v2-abliterated-6bit-mlx"
Response includes
  • classification m1
  • files 15
  • hub_downloads_all_time 4,526
  • author_summary 31 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
5K
76 last 30d - cooling
Likes
2
Model age
5mo ago
created 2026-05-09
Downloads over time
Now4.5K→from4.2K↑9%
4.2K4.3K4.4K4.6K4.2K on Aug 54.5K on Oct 11AugSepOct
Aug 5 → Oct 11 · 50 snapshots · spans 67 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en multilingual
Tags
mlx safetensors qwen3_5 mlx-lm qwen qwen3 qwen3.6 abliterated refusal-ablated reasoning claude-opus-distill text-generation

Related

Total size
20.4 GB
Files
15
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-05 19:15

Files by quantization

Auxiliary files 15 files 20.4 GB
model-00001-of-00005.safetensors 5.00 GB fcdceb1d download
model-00002-of-00005.safetensors 4.98 GB e67a520f download
model-00003-of-00005.safetensors 4.97 GB 3d76134a download
model-00004-of-00005.safetensors 4.45 GB dfc2dc71 download
model-00005-of-00005.safetensors 985 MB 6662dd58 download
tokenizer.json 19.1 MB 06b95093 download
model.safetensors.index.json 185 KB 833de0b9 download
logo.png 18.6 KB a9400259 download
chat_template.jinja 7.58 KB a8755d82 download
README.md 7.51 KB 8acce5b4 download
config.json 3.27 KB 024a3882 download
DFLASH_SPECULATIVE_DECODING.md 1.83 KB 8eed81f0 download
.gitattributes 1.53 KB 52373fe2 download
tokenizer_config.json 1.36 KB e1a3e847 download
generation_config.json 218 B edcb0e7f download

README current version from Hugging Face


license: apache-2.0
language:

  • en
  • multilingual
    tags:
  • mlx
  • mlx-lm
  • qwen
  • qwen3
  • qwen3.6
  • abliterated
  • refusal-ablated
  • reasoning
  • claude-opus-distill
  • text-generation
  • apple-silicon
    library_name: mlx
    pipeline_tag: text-generation
    base_model: TeichAI/Qwen3.6-27B-Claude-Opus-Reasoning-Distill-v2
    base_model_relation: quantized

Lemura Labs

Qwen3.6-27B-Claude-Opus-Reasoning-Distill-v2-abliterated-6bit-mlx

Format Task Params Type BPW Size Context License

THIS IS A TEXT-ONLY MODEL — NO VISION

The upstream abliteration pass stripped the vision tower. For vision-capable Qwen 3.6 27B Opus-Distill MLX quants, see our parallel repos at huggingface.co/lemuralabs (look for repos without -abliterated in the name).

6-bit affine MLX quantization of an abliterated Qwen 3.6 27B Claude-Opus reasoning distill, by the Lemura Labs team.

Standard 6-bit MLX quantization. Effectively lossless on most reasoning benchmarks vs BF16. The recommended pick when you have the RAM headroom.


TL;DR

Disk size ~20 GB
Effective BPW 6.0
Scheme Affine 6-bit, group size 64 (mlx-lm default)
Recommended RAM 32 GB Apple Silicon (M4 Pro 32 GB, M3/M2 Max base)
Vision No — text-only (the upstream abliteration step stripped the ViT)
Made by Lemura Labs

Lineage

Qwen/Qwen3.6-27B (Qwen Team — base pretrain)
 │
 ▼
TeichAI/Qwen3.6-27B-Claude-Opus-Reasoning-Distill-v2 (TeichAI — Claude-Opus reasoning distill)
 │
 ▼
abliterated (refusal-ablated) via OBLITERATUS v0.1.2 (multi-direction SVD, BF16)
 │
 ▼
this repo — 6-bit affine, MLX format (Lemura Labs team — quantization)

Direct upstream links:


Use it

mlx-lm (recommended)

pip install mlx-lm
from mlx_lm import load, generate

model, tokenizer = load("lemuralabs/Qwen3.6-27B-Claude-Opus-Reasoning-Distill-v2-abliterated-6bit-mlx")
prompt = "Explain the difference between SSM and softmax attention in three sentences."
out = generate(model, tokenizer, prompt=prompt, max_tokens=400)
print(out)

Chat template

messages = [
 {"role": "system", "content": "You are a helpful, candid reasoning assistant."},
 {"role": "user", "content": "Plan a 3-day Tokyo itinerary for a foodie."},
]
prompt = tokenizer.apply_chat_template(messages, add_generation_prompt=True, tokenize=False)
print(generate(model, tokenizer, prompt=prompt, max_tokens=600))

CLI

mlx_lm.generate --model lemuralabs/Qwen3.6-27B-Claude-Opus-Reasoning-Distill-v2-abliterated-6bit-mlx --prompt "Hello" --max-tokens 256

Quantization details

  • Source weights: BF16 abliterated checkpoint (28 shards, ~57 GB) derived from TeichAI/Qwen3.6-27B-Claude-Opus-Reasoning-Distill-v2 via OBLITERATUS multi-direction SVD ablation (preserves coherence; KL drift = 0.149 from base).
  • Quantization scheme: Affine 6-bit, group size 64 (mlx-lm default).
  • Group size: 64.
  • Calibration corpus: mlx-lm calibration_v5 (~427 KB English text, used for OptiQ sensitivity ranking; uniform/affine variants do not require calibration).
  • Sanity check: forward perplexity on held-out calibration text within 1–3% of next-higher-precision sibling.

Architecture notes

The Qwen 3.6 27B family uses a hybrid attention stack — 4 GatedDeltaNet (linear-attention/SSM) layers followed by 1 full-softmax-attention layer, repeated 16× for 64 total layers, 5120 hidden, 248K vocab, 262K context. The SSM kernels lack a VJP path in MLX, so backward-pass-based quant methods (DWQ, dynamic quant) cannot be applied here — OptiQ's forward-only sensitivity approach is the only calibration-aware option that works on this architecture. That's why the OptiQ variants exist.


Behavior caveats

  • Text-only — no vision. The abliteration pipeline (OBLITERATUS) ran on the LM tower and stripped the ViT. For vision-capable quants of the same Opus-Distill v2 lineage, use our parallel non-abliterated repos at huggingface.co/lemuralabs (any repo without -abliterated in the name).
  • This is an abliterated model — refusal directions were surgically removed from the parent. It will answer prompts the parent would refuse. Use responsibly and within applicable law.
  • Quantization preserves abliteration: the refusal rate measured at BF16 (~35% from a 100% baseline) stays in that range across our quants.

Credits

Quantization & release Lemura Labs
Reasoning distill TeichAI (Claude-Opus 4.5/4.6 high-reasoning datasets)
Foundation model Qwen Team
Abliteration toolkit OBLITERATUS by elder-plinius
Quant toolkit mlx-lm, mlx-optiq

License

Apache-2.0, inherited from the foundation and distill upstream.


Need a hosted endpoint, custom quant, or larger-scale inference? — multi-provider LLM routing for the Indian developer ecosystem.

3.3–3.7× faster decoding with DFlash (lossless, MLX)

This MLX build supports lossless block-diffusion speculative decoding via DFlash in mlx_vlm — no requantization, no model changes. On an Apple M4 Max we measured 3.38× (8-bit) and 3.67× (bf16) decode speedups with byte-identical output; other MLX quants of this model should see a similar ~3×.

python3 -m mlx_vlm generate \
 --model lemuralabs/Qwen3.6-27B-Claude-Opus-Reasoning-Distill-v2-abliterated-6bit-mlx \
 --draft-model z-lab/Qwen3.6-27B-DFlash --draft-kind dflash \
 --prompt "Write a merge function for two sorted lists in Python." --max-tokens 256
  • Requires mlx_vlm ≥ 0.5.0 and access to the gated drafter z-lab/Qwen3.6-27B-DFlash (one-click "Agree and access").
  • Accelerates the text path only (vision is unaffected); adds ~3.9 GB for the drafter.
  • Acceptance ≈ 8.95 tokens/round (block size 16); the target runs ~10× fewer forward passes.
  • Full write-up & benchmarks: [] · see also DFLASH_SPECULATIVE_DECODING.md.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-05Initial commitc7579c77.5 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration