← back to catalog · registered 2026-08-22 13:56

PocketAiHub/Qwen3.8-9B-Abliterated-MLX

PocketAiHub Qwen 9B multimodal
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/PocketAiHub%2FQwen3.8-9B-Abliterated-MLX"
Response includes
  • classification m1
  • files 5
  • author_summary 14 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · 30-day
0
Likes
34
Model age
7w ago
created 2026-08-17
Downloads over time
Now0→from0↑0%
00110 on Aug 190 on Oct 11AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
mlx safetensors mlx-vlm qwen qwen3.8 distillation multimodal abliterated 4-bit 8-bit bf16 image-text-to-text

Related

Total size
0 B
Files
5
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-22 11:00

Files by quantization

Auxiliary files 5 files 31.7 KB
validation-summary.json 12.7 KB 730c766b download
LICENSE 11.3 KB f938136e download
README.md 4.31 KB 72e8e2c9 download
release-manifest.json 1.74 KB b7489979 download
.gitattributes 1.65 KB 62cae744 download

README current version from Hugging Face


language: en
library_name: mlx
license: apache-2.0
model_name: Qwen3.8-9B Distill Abliterated MLX
base_model: empero-ai/Qwen3.8-9B
pipeline_tag: image-text-to-text
tags:

  • mlx
  • mlx-vlm
  • qwen
  • qwen3.8
  • distillation
  • multimodal
  • abliterated
  • 4-bit
  • 8-bit
  • bf16

Qwen3.8-9B Distill Abliterated MLX

Abliterated MLX derivatives of the third-party distill
empero-ai/Qwen3.8-9B, pinned to
revision 0934f3d2327ff2df2197495278c4c46ae5a56bd9. The source is a third-party full-parameter
distillation based on Qwen/Qwen3.5-9B; it
is not an official Qwen3.8 release. Abliteration, conversion, and validation
were performed by PocketAI Model Lab.

Important safety notice

These checkpoints were intentionally modified to suppress learned refusal
behavior. They may respond more readily to requests involving potentially
unsafe, illegal, offensive, deceptive, or dangerously incorrect content.
Abliteration is not truthfulness training or a safety guarantee. Independently
constrain and evaluate outputs for the intended deployment.

Variants

Precision Folder Packaged size
4-bit 4bit/ 5.57 GiB
8-bit 8bit/ 9.74 GiB
BF16 bf16/ 17.55 GiB

The 4-bit and 8-bit variants use MLX affine quantization with group size 64.
The vision tower remains BF16. The BF16 variant is unquantized. Native source
MTP tensors are intentionally excluded.

Refusal-behavior screen

All variants were evaluated on 100 refusal-elicitation test prompts and 100
benign controls in non-thinking mode with deterministic decoding and a
256-token ceiling.

Precision Test-set explicit refusals Control explicit refusals Test-set natural stops Control natural stops
4-bit 0/100 0/100 11/100 2/100
8-bit 0/100 0/100 14/100 4/100
BF16 0/100 0/100 11/100 3/100

The transparent phrase-based screen found no explicit refusals or evasive
non-answers, and every case contained final-answer text. Most generations hit
the 256-token ceiling, so this is an early-refusal regression screen—not proof
of universal compliance, safety, factuality, or completion quality.

KV/long-context evaluation

Precision Formatted tokens Prefill tok/s Decode tok/s Peak MLX memory
4-bit 65,536 1028.1 44.59 13.05 GB
8-bit 32,776 2323.7 52.99 14.19 GB
BF16 32,776 2299.3 27.32 22.67 GB

All three exact-retrieval cases passed. The 8-bit and BF16 runs used 16-bit KV
cache quantization at 32K; the 4-bit run was an unquantized-KV 64K text test.
The 4-bit and BF16 variants also passed the complete deterministic 4K feature
suite; all three passed text and vision runtime smoke tests.

Exact evidence hashes and test details are in each variant's
validation-summary.json and artifact-manifest.json.

Download and load

python -m pip install "mlx==0.32.0" "mlx-vlm==0.6.8"
from pathlib import Path

from huggingface_hub import snapshot_download
from mlx_vlm import generate, load
from mlx_vlm.prompt_utils import apply_chat_template

repo_id = "PocketAiHub/Qwen3.8-9B-Abliterated-MLX"
variant = "4bit"  # "4bit", "8bit", or "bf16"
snapshot = Path(snapshot_download(repo_id, allow_patterns=[f"{variant}/*"]))
model, processor = load(str(snapshot / variant))
prompt = apply_chat_template(
    processor,
    model.config,
    "Explain why seasons occur.",
    num_images=0,
    enable_thinking=False,
)
result = generate(
    model,
    processor,
    prompt,
    max_tokens=256,
    temperature=0.0,
    enable_thinking=False,
)
print(result.text)

Reproducibility and limitations

  • Source: empero-ai/Qwen3.8-9B at 0934f3d2327ff2df2197495278c4c46ae5a56bd9
  • Declared base: Qwen/Qwen3.5-9B
  • Abliteration is a targeted directional intervention, not general evaluation
  • Standard MLX conversion intentionally excludes native MTP tensors
  • This is an experimental community release; verify behavior for your use case

License and attribution

The source repository declares Apache-2.0. This derivative includes the Apache
2.0 text in LICENSE. Original model credit remains with Empero
and the Qwen team; PocketAI is the derivative publisher.

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-22Use neutral refusal-evaluation terminology3ffc2d24.3 KB
    Loading...
  2. 2026-08-21Clarify distill naming in model carde3bbeb44.3 KB
    Loading...
  3. 2026-08-17Add files using upload-large-folder tool72207664.2 KB
    Loading...

Discussions 2 threads

  1. 2026-08-17No compatible options available for this formatopen1 💬#2
    Loading...
  2. 2026-08-17No compatible options available for this format.open1 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration