← back to catalog · registered 2026-08-22 13:56

gregfrank/Mistral-Large-Instruct-2411-ULRE-abliterated

gregfrank Mistral 123B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/gregfrank%2FMistral-Large-Instruct-2411-ULRE-abliterated"
Response includes
  • classification m1
  • files 22
  • benchmarks 11 entries
  • hub_downloads_all_time 2,762
  • author_summary 3 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
3K
420 last 30d - stable
Likes
3
Model age
4mo ago
created 2026-06-06
Downloads over time
Now3K→from1.3K↑139%
1.2K1.8K2.5K3.2K1.3K on Jun 103K on Oct 11JunJulAugSepOct
Jun 10 → Oct 11 · 57 snapshots · spans 123 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 3.8 UGI
Hazardous 7.1 UGI
Natural Intelligence 36.21 UGI
Political lean -18.7% UGI
Sensitive-Info 51.35 UGI
SocPol 5.3 UGI
UGI 59.23 UGI
Willingness (10) 7.5 UGI
W10-Adherence 9 UGI
W10-Direct 6 UGI
Writing 37.84 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
other
Languages
en
Tags
mlx safetensors mistral abliterated uncensored ulre text-generation conversational en base_model:mistralai/Mistral-Large-Instruct-2411 base_model:quantized:mistralai/Mistral-Large-Instruct-2411 license:other

Related

Total size
71.4 GB
Files
22
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-06-06 09:24

Files by quantization

Auxiliary files 22 files 71.4 GB
model-00009-of-00015.safetensors 5.00 GB 91a4c927 download
model-00014-of-00015.safetensors 5.00 GB 94069cfd download
model-00007-of-00015.safetensors 5.00 GB 649b9e79 download
model-00005-of-00015.safetensors 5.00 GB 166bf09c download
model-00012-of-00015.safetensors 5.00 GB 88c548ce download
model-00002-of-00015.safetensors 5.00 GB 9db2f9f3 download
model-00010-of-00015.safetensors 4.98 GB b3d970bb download
model-00003-of-00015.safetensors 4.98 GB b34ff03d download
model-00004-of-00015.safetensors 4.92 GB ae9c0abb download
model-00011-of-00015.safetensors 4.92 GB efc46d7e download
model-00008-of-00015.safetensors 4.88 GB 50193698 download
model-00006-of-00015.safetensors 4.88 GB e884285d download
model-00013-of-00015.safetensors 4.88 GB d9bfa898 download
model-00001-of-00015.safetensors 4.86 GB 74aa2f20 download
model-00015-of-00015.safetensors 2.09 GB 039fc1fd download
tokenizer.json 3.50 MB 8d0372a3 download
model.safetensors.index.json 199 KB 9ab7907b download
README.md 3.75 KB 2b3b2db3 download
.gitattributes 1.48 KB a6344aac download
config.json 816 B b40448b8 download
tokenizer_config.json 458 B c9d08f1a download
chat_template.jinja 423 B b3cc8af1 download

README current version from Hugging Face


base_model: mistralai/Mistral-Large-Instruct-2411
license: other
license_name: mrl
license_link: https://mistral.ai/licenses/MRL-0.1.md
library_name: mlx
pipeline_tag: text-generation
tags:

  • mlx
  • abliterated
  • uncensored
  • ulre
  • mistral
    language:
  • en

Mistral-Large-Instruct-2411-ULRE (abliterated, MLX 4-bit)

An abliterated (refusal-reduced) build of
mistralai/Mistral-Large-Instruct-2411
(123B), 4-bit MLX, produced with ULRE — a per-layer residual-stream steering edit baked into the
attention output projection. Quantized base from
zachlandes/Mistral-Large-Instruct-2411-Q4-MLX.

Mistral-Large-2411 is a strongly safety-tuned 123B model. ULRE de-refuses it cleanly on a
deflection-aware harmful-prompt eval while preserving capability (math/code within 1pp of base),
and it loads on stock mlx-lm / LM Studio with no loader patch.

Results

Judged by an independent local LLM judge (gpt-oss-120b-heretic) on 100 held adversarial prompts @512
tokens (thinking ON), scoring 0=refuse … 1=clean compliance … 4-5=strong steering/propaganda.

base Mistral-Large-2411 this model (ULRE α=0.2)
clean compliance ~4 / 24 (screen) 78 / 100
refuse ~20 / 24 (screen) 6 / 100
mild — 9 / 100
strong-steer (4-5) — 7 / 100
mean score — 1.24

Capability gates (same harness, base vs this model) — within the charter thresholds
(≤5pp drop on math/code):

gate base this model (ULRE α=0.2)
math (GSM8K) 0.98 0.97 (−1pp)
code (HumanEval) 0.925 0.925 (0pp)

Tool-call validity is omitted: the project's tool-call gate is built for Qwen-style calls and scores
base Mistral-Large at only ~0.10 (Mistral uses a different [TOOL_CALLS] format), so it is not a
meaningful signal for this model — base and edited score the same.

Method (ULRE)

Modern refusal behaves like a routed control circuit, not a single residual feature. ULRE subtracts
alpha * u_l (the layer-l harmful−harmless activation mean-difference direction) from the output of
a band of decoder layers (here o_proj on layers 30–51, alpha = 0.2), baked statically into each
window layer's o_proj bias (o_proj.bias = -alpha * u_l).

The alpha scale is model-specific and must be tuned to the lowest value that saturates de-refusal:
Mistral-Large saturates de-refusal at α≈0.2; pushing higher (α≈0.5) does not de-refuse any better
but collapses chain-of-thought reasoning (math 0.98→0.09). This build uses the capability-preserving
α=0.2.

Loading — stock mlx-lm / LM Studio (no patch)

Mistral uses mlx-lm's llama.py, which already supports an attention_bias flag. The static file
sets attention_bias: true in config.json and carries zero biases on q/k/v (and off-window
o_proj) plus the steering bias on the windowed o_proj — so it loads as-is:

from mlx_lm import load, generate
model, tok = load("gregfrank/Mistral-Large-Instruct-2411-ULRE-abliterated")
print(generate(model, tok, prompt=tok.apply_chat_template(
    [{"role": "user", "content": "Hello"}], tokenize=False, add_generation_prompt=True),
    max_tokens=256))

In LM Studio: search/download the repo (or lms get gregfrank/Mistral-Large-Instruct-2411-ULRE-abliterated)
and load it like any other MLX model — no configuration needed.

Intended use & safety

Research artifact for studying refusal mechanisms and safety-tuning robustness. It will comply with
requests a stock model refuses. Use responsibly and in accordance with the Mistral Research
License
(research / non-commercial) of the base model and applicable law.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-06-06Add files using upload-large-folder tool6ae864a3.8 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration