← back to catalog · registered 2026-08-22 13:56

divinetribe/Hermes-4-14B-abliterated-4bit-mlx

divinetribe Qwen 15B second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/divinetribe%2FHermes-4-14B-abliterated-4bit-mlx"
Response includes
  • classification m1
  • files 10
  • hub_downloads_all_time 4,056
  • author_summary 14 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
4K
1K last 30d - stable
Likes
5
Model age
4mo ago
created 2026-05-22
Downloads over time
Now4.5K→from93↑4,788%
01.7K3.3K5K93 on May 204.5K on Oct 11MayJunJulAugSepOct
May 20 → Oct 11 · 61 snapshots · spans 144 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en
Tags
mlx safetensors qwen3 abliterated uncensored hermes hermes-4 mlx-4bit apple-silicon on-device mlx-lm local-llm

Related

Total size
7.74 GB
Files
10
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-22 20:27

Files by quantization

Auxiliary files 10 files 7.75 GB
model-00001-of-00002.safetensors 4.99 GB a6baa3d9 download
model-00002-of-00002.safetensors 2.75 GB a4facab0 download
tokenizer.json 10.9 MB be756060 download
model.safetensors.index.json 84.3 KB 0ca4bf99 download
README.md 5.74 KB 14cb3e87 download
chat_template.jinja 3.80 KB b76787ec download
config.json 1.98 KB e1ea9fc5 download
.gitattributes 1.53 KB 52373fe2 download
tokenizer_config.json 384 B 397f359a download
generation_config.json 186 B 136752c6 download

README current version from Hugging Face


base_model: Babsie/Hermes-4-14B-BF16-abliterated
base_model_relation: quantized
license: apache-2.0
language:

  • en
    library_name: mlx
    pipeline_tag: text-generation
    tags:
  • abliterated
  • uncensored
  • hermes
  • hermes-4
  • qwen3
  • mlx
  • mlx-4bit
  • apple-silicon
  • on-device
  • mlx-lm
  • local-llm
  • macbook
  • 14b
  • 4-bit

Hermes-4-14B-abliterated-4bit-mlx

Uncensored Hermes 4 14B, running locally on a Mac. Abliterated, quantized to 4-bit MLX for Apple Silicon — 8.3 GB, comfortable on a 16 GB Mac. No cloud, no API key, no refusals.

A 4-bit MLX quantization of Babsie/Hermes-4-14B-BF16-abliterated, tuned for fast on-device inference on Apple Silicon.

  • Base model: Babsie/Hermes-4-14B-BF16-abliterated (BF16 abliterated, ~28 GB)
  • Architecture: Qwen3 (Hermes 4 post-training)
  • Quantization: 4-bit affine, group size 64
  • Format: MLX safetensors
  • Footprint: ~8 GB on disk (two safetensors shards: 5.35 GB + 2.95 GB), runs comfortably on a 16 GB Mac and flies on 32 GB+
  • Context: 40,960 tokens (~40 K) — inherited unchanged from upstream Hermes 4 14B

Why this exists

Hermes 4 is NousResearch's instruction-tuned family. The 14B variant is built on Qwen3, so it inherits Qwen3's tokenizer and chat template — but the post-training is pure Hermes, which means it's tuned for tool use, structured output, and not refusing benign-but-edgy questions.

Babsie's abliteration applies refusal-direction projection (Arditi et al., 2024) to the BF16 weights. The model still has all its capabilities — it just doesn't pre-empt with a refusal on questions that the upstream Hermes would politely decline. You become the moderator.

This MLX 4-bit conversion makes it usable on Apple Silicon at the speed the hardware was designed for. 14B at 4-bit fits in ~8 GB, which means it runs even on a 16 GB MacBook — and on a 64+ GB machine you can keep multiple models loaded simultaneously.

Usage

Python (mlx-lm)

from mlx_lm import load, generate

model, tokenizer = load("divinetribe/Hermes-4-14B-abliterated-4bit-mlx")

messages = [{"role": "user", "content": "Write a haiku about local inference."}]
prompt = tokenizer.apply_chat_template(
    messages, add_generation_prompt=True, tokenize=False
)

response = generate(model, tokenizer, prompt=prompt, max_tokens=256)
print(response)

Local OpenAI-compatible server

pip install mlx-lm
mlx_lm.server --model divinetribe/Hermes-4-14B-abliterated-4bit-mlx --port 8080

That gives you a local http://localhost:8080/v1/chat/completions endpoint that any OpenAI SDK client can hit. No tokens, no API bills, no telemetry.

Inside claude-code-local

MLX_MODEL=divinetribe/Hermes-4-14B-abliterated-4bit-mlx \
  bash scripts/start-mlx-server.sh

This routes Claude Code (or any Anthropic-API client) to the local model. See the claude-code-local repo for the full setup.

Where it fits in the lineup

Model (all on divinetribe) Disk Params Best for
Llama-3.3-70B-Instruct-abliterated-8bit-mlx ~75 GB 71 B Hardest reasoning, 96 GB+ Macs
gemma-4-31b-it-abliterated-4bit-mlx ~16 GB 31 B Daily coding, 32 GB+ Macs
Hermes-4-14B-abliterated-4bit-mlx (this) ~8 GB 14 B 16 GB Macs, tool use, instruction-following

Abliteration

"Abliteration" suppresses the model's built-in refusal direction so it doesn't refuse benign-but-edgy requests. It is not a general capability upgrade — use responsibly, and you remain bound by the upstream Hermes 4 / Qwen3 licenses.

Credits

License

Apache 2.0, inherited from the upstream Hermes 4 / Qwen3 family.


About the author

This model was built by Matt Macosko (@nicedreamzapp) for the claude-code-local stack — run Claude Code 100% on-device with local AI on Apple Silicon (⭐ 2,664 on GitHub).

More abliterated MLX models

Part of the Abliterated MLX for Apple Silicon
collection — 11 uncensored models converted for Macs, from Gemma 4 12B up to
Llama 3.3 70B, plus Qwen3, Qwen3-VL, Hermes 4 and Muse Glimmer 30B.


Part of Claude Code Local

This model is one of the fighters in Claude Code Local (3.2k★), which runs Claude Code 100% on-device on Apple Silicon through an MLX-native Anthropic-API server. Not sure which local model to run as an agent? Check the Agent-12 local agent leaderboard: real agent tasks, judged by the filesystem, same hardware for every row.

Built by Matt Macosko in Arcata, CA. Open to work on local-AI and Apple Silicon inference: [email protected].

README history 5 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-22Model card: link to Claude Code Local and the Agent-12 leaderboard637fb2b5.7 KB
    Loading...
  2. 2026-08-12SEO: lead with a searchable summary, cross-link the collectiond9103bf5.1 KB
    Loading...
  3. 2026-08-12SEO: base_model lineage, language, Apple Silicon / MLX tags4644e9c4.6 KB
    Loading...
  4. 2026-05-22fix(README): correct context (40K not 128K) and disk size (~8GB not ~7GB)faf64f64.6 KB
    Loading...
  5. 2026-05-22Initial upload: MLX 4-bit quantization of Babsie/Hermes-4-14B-BF16-abliterated95f34c74.4 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration