← back to catalog · registered 2026-08-22 13:56

Foresee/Qwen3.8-9B-heretic-uncensored-5bit-MLX

Foresee Qwen 9.0B second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/Foresee%2FQwen3.8-9B-heretic-uncensored-5bit-MLX"
Response includes
  • classification m3
  • files 10
  • hub_downloads_all_time 5,877
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · lifetime
6K
2K last 30d - stable
Likes
3
Model age
7w ago
created 2026-08-18
Downloads over time
Now6.4K→from1.7K↑272%
1.5K3.3K5.1K6.9K1.7K on Aug 196.4K on Oct 11AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en
Tags
mlx safetensors qwen3_5 qwen3.5 qwen3.8 text-generation conversational en license:apache-2.0 5-bit region:us

Related

Total size
6.26 GB
Files
10
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-18 08:12

Files by quantization

Auxiliary files 10 files 6.27 GB
model-00001-of-00002.safetensors 4.99 GB 72c76a46 download
model-00002-of-00002.safetensors 1.27 GB 297e3724 download
tokenizer.json 19.1 MB 06b95093 download
model.safetensors.index.json 92.9 KB 7c9beda0 download
chat_template.jinja 7.57 KB a585dec8 download
config.json 3.02 KB 0f2c8773 download
.gitattributes 1.53 KB 52373fe2 download
README.md 1.52 KB 4935a065 download
tokenizer_config.json 1.17 KB ffb89ca0 download
generation_config.json 174 B a0850a0f download

README current version from Hugging Face


license: apache-2.0
base_model: rohit267/Qwen3.8-9B-heretic-uncensored
language:

  • en
    library_name: mlx
    pipeline_tag: text-generation
    tags:
  • mlx
  • qwen3_5
  • qwen3.5
  • qwen3.8
  • text-generation

Qwen3.8-9B Heretic Uncensored - 5-bit MLX

This repository contains 5-bit MLX weights for rohit267/Qwen3.8-9B-heretic-uncensored.

The requested reference model was saga404/Qwen3.8-9B-heretic-uncensored-Q5_0-GGUF.
MLX LM cannot convert GGUF weights directly.
This conversion used the original BF16 safetensors weights to avoid an extra dequantization and requantization step.

Quantization

Setting Value
Weight bits 5
Group size 32
Mode affine
MLX LM version 0.31.3
MLX version 0.32.1

The final model uses about 6.001 bits per weight after scales, biases, and unquantized parameters are included.

Use

Install MLX LM:

pip install mlx-lm

Run the downloaded model:

mlx_lm.generate \
  --model ./Qwen3.8-9B-heretic-uncensored-5bit-MLX \
  --prompt "What is 2 + 2?" \
  --max-tokens 128

Use the source model's recommended sampling settings for longer responses:

temperature=0.6
top_p=0.95
top_k=20

Validation

The converted weights loaded and generated a correct response to 2 + 2 on Apple silicon.
Peak memory during this smoke test was 6.859 GB.

License

The source model is licensed under Apache 2.0.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-18Upload folder using huggingface_hub8a3eec71.5 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration