← back to catalog · registered 2026-08-22 13:56

lemuralabs/Qwen3.6-27B-Claude-Opus-Reasoning-Distill-v2-abliterated-OptiQ-4.5bpw-mlx

lemuralabs Qwen 27B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/lemuralabs%2FQwen3.6-27B-Claude-Opus-Reasoning-Distill-v2-abliterated-OptiQ-4.5bpw-mlx"
Response includes
  • classification m1
  • files 14
  • hub_downloads_all_time 3,750
  • author_summary 31 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
4K
169 last 30d - cooling
Likes
2
Model age
5mo ago
created 2026-05-09
Downloads over time
Now3.8K→from3.1K↑23%
3K3.3K3.6K3.9K3.1K on Aug 53.8K on Oct 11AugSepOct
Aug 5 → Oct 11 · 50 snapshots · spans 67 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en multilingual
Tags
mlx safetensors qwen3_5 mlx-lm qwen qwen3 qwen3.6 abliterated refusal-ablated reasoning claude-opus-distill text-generation

Related

Total size
16.2 GB
Files
14
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-05 19:16

Files by quantization

Auxiliary files 14 files 16.2 GB
model-00002-of-00004.safetensors 4.98 GB 18e0d3bc download
model-00001-of-00004.safetensors 4.97 GB 11fcf270 download
model-00003-of-00004.safetensors 4.96 GB 86724b06 download
model-00004-of-00004.safetensors 1.26 GB 0deb6e72 download
tokenizer.json 19.1 MB 06b95093 download
model.safetensors.index.json 185 KB 0178b42a download
config.json 105 KB 2e483b8e download
optiq_metadata.json 51.4 KB 26cbf1f2 download
logo.png 18.6 KB a9400259 download
chat_template.jinja 7.58 KB a8755d82 download
README.md 6.55 KB e272be3f download
.gitattributes 1.53 KB 52373fe2 download
tokenizer_config.json 1.36 KB e1a3e847 download
generation_config.json 218 B edcb0e7f download

README current version from Hugging Face


license: apache-2.0
language:

  • en
  • multilingual
    tags:
  • mlx
  • mlx-lm
  • qwen
  • qwen3
  • qwen3.6
  • abliterated
  • refusal-ablated
  • reasoning
  • claude-opus-distill
  • text-generation
  • apple-silicon
    library_name: mlx
    pipeline_tag: text-generation
    base_model: TeichAI/Qwen3.6-27B-Claude-Opus-Reasoning-Distill-v2
    base_model_relation: quantized

Lemura Labs

Qwen3.6-27B-Claude-Opus-Reasoning-Distill-v2-abliterated-OptiQ-4.5bpw-mlx

Format Task Params Type BPW Size Context License

THIS IS A TEXT-ONLY MODEL — NO VISION

The upstream abliteration pass stripped the vision tower. For vision-capable Qwen 3.6 27B Opus-Distill MLX quants, see our parallel repos at huggingface.co/lemuralabs (look for repos without -abliterated in the name).

OptiQ mixed 4.5 BPW MLX quantization of an abliterated Qwen 3.6 27B Claude-Opus reasoning distill, by the Lemura Labs team.

The sweet spot. KL-sensitivity scan from mlx-optiq keeps the most-perturbing projections at 6-bit while pushing the rest to 4-bit. Quality very close to 6-bit at ~80% of the size.


TL;DR

Disk size ~16 GB
Effective BPW 4.5
Scheme OptiQ sensitivity-driven mixed 4.5 BPW (mostly 4-bit with selective 6-bit on critical projections)
Recommended RAM 24 GB Apple Silicon (M4 Pro base, M3/M2 Pro)
Vision No — text-only (the upstream abliteration step stripped the ViT)
Made by Lemura Labs

Lineage

Qwen/Qwen3.6-27B (Qwen Team — base pretrain)
 │
 ▼
TeichAI/Qwen3.6-27B-Claude-Opus-Reasoning-Distill-v2 (TeichAI — Claude-Opus reasoning distill)
 │
 ▼
abliterated (refusal-ablated) via OBLITERATUS v0.1.2 (multi-direction SVD, BF16)
 │
 ▼
this repo — OptiQ mixed 4.5 BPW, MLX format (Lemura Labs team — quantization)

Direct upstream links:


Use it

mlx-lm (recommended)

pip install mlx-lm
from mlx_lm import load, generate

model, tokenizer = load("lemuralabs/Qwen3.6-27B-Claude-Opus-Reasoning-Distill-v2-abliterated-OptiQ-4.5bpw-mlx")
prompt = "Explain the difference between SSM and softmax attention in three sentences."
out = generate(model, tokenizer, prompt=prompt, max_tokens=400)
print(out)

Chat template

messages = [
 {"role": "system", "content": "You are a helpful, candid reasoning assistant."},
 {"role": "user", "content": "Plan a 3-day Tokyo itinerary for a foodie."},
]
prompt = tokenizer.apply_chat_template(messages, add_generation_prompt=True, tokenize=False)
print(generate(model, tokenizer, prompt=prompt, max_tokens=600))

CLI

mlx_lm.generate --model lemuralabs/Qwen3.6-27B-Claude-Opus-Reasoning-Distill-v2-abliterated-OptiQ-4.5bpw-mlx --prompt "Hello" --max-tokens 256

Quantization details

  • Source weights: BF16 abliterated checkpoint (28 shards, ~57 GB) derived from TeichAI/Qwen3.6-27B-Claude-Opus-Reasoning-Distill-v2 via OBLITERATUS multi-direction SVD ablation (preserves coherence; KL drift = 0.149 from base).
  • Quantization scheme: OptiQ sensitivity-driven mixed 4.5 BPW (mostly 4-bit with selective 6-bit on critical projections).
  • Group size: 64.
  • Calibration corpus: mlx-lm calibration_v5 (~427 KB English text, used for OptiQ sensitivity ranking; uniform/affine variants do not require calibration).
  • Sanity check: forward perplexity on held-out calibration text within 1–3% of next-higher-precision sibling.

Architecture notes

The Qwen 3.6 27B family uses a hybrid attention stack — 4 GatedDeltaNet (linear-attention/SSM) layers followed by 1 full-softmax-attention layer, repeated 16× for 64 total layers, 5120 hidden, 248K vocab, 262K context. The SSM kernels lack a VJP path in MLX, so backward-pass-based quant methods (DWQ, dynamic quant) cannot be applied here — OptiQ's forward-only sensitivity approach is the only calibration-aware option that works on this architecture. That's why the OptiQ variants exist.


Behavior caveats

  • Text-only — no vision. The abliteration pipeline (OBLITERATUS) ran on the LM tower and stripped the ViT. For vision-capable quants of the same Opus-Distill v2 lineage, use our parallel non-abliterated repos at huggingface.co/lemuralabs (any repo without -abliterated in the name).
  • This is an abliterated model — refusal directions were surgically removed from the parent. It will answer prompts the parent would refuse. Use responsibly and within applicable law.
  • Quantization preserves abliteration: the refusal rate measured at BF16 (~35% from a 100% baseline) stays in that range across our quants.

Credits

Quantization & release Lemura Labs
Reasoning distill TeichAI (Claude-Opus 4.5/4.6 high-reasoning datasets)
Foundation model Qwen Team
Abliteration toolkit OBLITERATUS by elder-plinius
Quant toolkit mlx-lm, mlx-optiq

License

Apache-2.0, inherited from the foundation and distill upstream.


Need a hosted endpoint, custom quant, or larger-scale inference? — multi-provider LLM routing for the Indian developer ecosystem.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-05Initial commit953839a6.6 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration