← back to catalog · registered 2026-10-09 23:58

j-llm/Qwen3.8-27B-Uncensored-OptiQ-4bit

j-llm 27B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/j-llm%2FQwen3.8-27B-Uncensored-OptiQ-4bit"
Response includes
  • classification m-uncensored
  • files 15
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-10-09
Downloads over time
Now0→from0↑0%
00110 on Oct 90 on Oct 10Oct
Oct 9 → Oct 10 · 2 snapshots · spans 1 day

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en zh ja
Tags
mlx safetensors qwen3_5 optiq mixed-precision quantized 4-bit 8-bit uncensored abliterated qwen3.8 vision

Related

Total size
18.0 GB
Files
15
Quantizations
1
Registered
2026-10-09 23:58
Last updated on HF
2026-10-09 23:56

Files by quantization

Auxiliary files 15 files 18.1 GB
model-00001-of-00004.safetensors 5.00 GB ******** download
model-00003-of-00004.safetensors 4.98 GB ******** download
model-00002-of-00004.safetensors 4.97 GB ******** download
model-00004-of-00004.safetensors 3.10 GB ******** download
tokenizer.json 19.1 MB ******** download
model.safetensors.index.json 204 KB 9500422c download
config.json 106 KB 8b21a81c download
LICENSE 11.3 KB f938136e download
chat_template.jinja 8.74 KB c0c686f9 download
README.md 5.21 KB c68610fa download
.gitattributes 1.53 KB 52373fe2 download
tokenizer_config.json 1.13 KB 8d610519 download
preprocessor_config.json 390 B 2ea84a43 download
video_preprocessor_config.json 385 B 3ba673a5 download
generation_config.json 202 B 023756cf download

README current version from Hugging Face


license: apache-2.0
base_model: orcarouter/Qwen3.8-27B-Uncensored
base_model_relation: quantized
pipeline_tag: image-text-to-text
library_name: mlx
language:

  • en
  • zh
  • ja
    tags:
  • mlx
  • optiq
  • mixed-precision
  • quantized
  • 4-bit
  • 8-bit
  • uncensored
  • abliterated
  • qwen3.8
  • vision
  • mtp
  • image-text-to-text

Qwen3.8-27B-Uncensored-OptiQ-4bit (MLX)

A sensitivity-aware mixed 4/8-bit OptiQ MLX quant of
orcarouter/Qwen3.8-27B-Uncensored,
the abliterated (refusal-removed) BF16 build of Qwen's Qwen/Qwen3.8-27B.

Unlike builds that reuse the base Qwen sensitivity ranking, this quant takes its
per-layer KL-divergence sensitivity directly from the Orca uncensored weights
against the BF16 source
(not a uniform-4-bit reference), then allocates 8-bit to
the sensitive layers and 4-bit to the rest. The native vision tower and MTP head
are preserved
.

  • ~19.2 GB on disk, 5.13 bits/weight (target 5.0)
  • 497 quantized tensors: 220 @ 8-bit, 277 @ 4-bit (group size 64)
  • Vision tower kept at BF16 in a sidecar (optiq/optiq_vision.safetensors)
  • MTP head preserved as a sidecar (optiq/mtp.safetensors) for speculative decoding
  • 262,144-token context, hybrid attention (48 linear + 16 full attention layers)

Benchmark — Capability Score 88.86

Full 6-benchmark OptiQ suite (reasoning mode), same Mac (M3 Ultra 96 GB), greedy decode.

Benchmark Score
MMLU (generative, reasoning) 90.1%
GSM8K (1000) 97.0%
IFEval (strict) 86.3%
BFCL-V3 (simple) 90.0%
HumanEval (pass@1) 95.7%
HashHop 74.0%
Capability Score 88.86

For context, the reference mlx-community/Qwen3.8-27B-OptiQ-4bit reports 87.98.

KL divergence vs the BF16 source (64 prompts × 256 tokens): mean 0.101, median 0.008, p95 0.300.

Quantization recipe

optiq convert orcarouter/Qwen3.8-27B-Uncensored \
  --method optiq \
  --target-bpw 5.0 \
  --candidate-bits 4,8 \
  --group-size 64 \
  --reference bf16 \
  --calibration-mix optiq \
  --n-calibration 40
  • mlx-optiq 0.5.19
  • Reference: BF16 (the full BF16 source fits in 96 GB RAM, so sensitivities are measured against the true base rather than a 4-bit proxy)
  • Calibration: OptiQ 6-domain mix, 40 samples

Optional: mixed-precision KV cache

A sensitivity-measured per-layer KV config is included in the source build notes and
reproduces 5.00 average KV bits (12 layers @ 4-bit + 4 layers @ 8-bit across the 16
full-attention layers; linear-attention layers are skipped):

optiq kv-cache ./Qwen3.8-27B-Uncensored-OptiQ-4bit --target-bits 5.0 --candidate-bits 4,8
optiq serve --model ./Qwen3.8-27B-Uncensored-OptiQ-4bit --kv-config ./kv_config.json

Usage

Text (stock mlx-lm)

pip install mlx-lm
from mlx_lm import load, generate
model, tokenizer = load("j-llm/Qwen3.8-27B-Uncensored-OptiQ-4bit")
print(generate(model, tokenizer, "What is 12*13?", max_tokens=64))

Image + text (stock mlx-vlm)

pip install mlx-vlm
from mlx_vlm import load, generate
from mlx_vlm.prompt_utils import apply_chat_template
from mlx_vlm.utils import load_config

model, processor = load("j-llm/Qwen3.8-27B-Uncensored-OptiQ-4bit")
config = load_config("j-llm/Qwen3.8-27B-Uncensored-OptiQ-4bit")
prompt = apply_chat_template(processor, config, "Describe this image.", num_images=1)
print(generate(model, processor, prompt, image=["test.png"], max_tokens=256).text)

The vision tower lives in the optiq/ sidecar and is restored automatically by both
OptiQ and current mlx-vlm. preprocessor_config.json is included so library and
app scanners (e.g. MLXBar) detect the model as image+text.

OptiQ serving (OpenAI + Anthropic API, MTP, KV)

pip install mlx-optiq
optiq serve --model j-llm/Qwen3.8-27B-Uncensored-OptiQ-4bit --mtp

Serving note: optiq serve's vision path currently requires mlx-lm==0.31.x
(_serve_single(self, request)); mlx-lm 0.32 changed that internal signature and
breaks the vision patch. Pin mlx-lm<0.32 in the serving environment.

Model details

Field Value
Base orcarouter/Qwen3.8-27B-Uncensored @ 8cb32d72080f6a47bf34dc5adf8067daa6c63a31
Architecture Qwen3_5ForConditionalGeneration — 64 layers, hidden 5120, hybrid Gated DeltaNet (48 linear + 16 full attention, interval 4)
Type Abliterated (refusal-direction removed), then OptiQ mixed-precision quantized
Quantization OptiQ mixed 4/8-bit, group size 64, affine; achieved 5.13 BPW
Vision BF16 sidecar, 333 tensors
MTP Preserved sidecar, 29 tensors (projections 4-bit, norms + mtp.fc BF16)
Context 262,144 tokens
Size ~19.2 GB

Safety

This is an abliterated / uncensored model: safety refusals have been removed from
the base weights. It is shared for research and red-teaming. Run it behind access
controls and do not expose it publicly without moderation. The upstream orcarouter
release carries the same warning.

Provenance

  • Source: orcarouter/Qwen3.8-27B-Uncensored (BF16, 55.6 GB), revision 8cb32d72…
  • Quantized with mlx-optiq 0.5.19 on Apple M3 Ultra (96 GB), BF16 reference, 40-sample calibration.
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration