← back to catalog · registered 2026-08-22 13:56

AITRADER/Jackrong-Qwen3.5-4B-Claude-Reasoning-abliterated-mxfp8-MLX

AITRADER Qwen 4.2B
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/AITRADER%2FJackrong-Qwen3.5-4B-Claude-Reasoning-abliterated-mxfp8-MLX"
Response includes
  • classification m1
  • files 11
  • hub_downloads_all_time 1,124
  • author_summary 30 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
1K
98 last 30d - cooling
Likes
1
Model age
7mo ago
created 2026-03-08
Downloads over time
Now1.2K→from347↑234%
3066179281.2K347 on Mar 111.2K on Oct 11MarAprMayJunJulAugSepOct
Mar 11 → Oct 11 · 70 snapshots · spans 214 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Tags
mlx safetensors qwen3_5 abliterated uncensored qwen3.5 reasoning claude-distilled mxfp8 quantized text-generation conversational

Related

Total size
9.94 GB
Files
11
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-03-13 21:36

Files by quantization

Auxiliary files 11 files 9.96 GB
model-00001-of-00002.safetensors 4.99 GB bbcbe24d download
model.safetensors 4.66 GB fcf5281b download
model-00002-of-00002.safetensors 297 MB 75f7f00a download
tokenizer.json 19.1 MB 87a7830d download
model.safetensors.index.json 78.2 KB b6c37cba download
chat_template.jinja 3.95 KB 609532bf download
config.json 3.58 KB 758af41b download
README.md 2.38 KB 4daf0390 download
.gitattributes 1.53 KB 52373fe2 download
tokenizer_config.json 1.17 KB d7920365 download
preprocessor_config.json 336 B 4ae180b4 download

README current version from Hugging Face


base_model: Jackrong/Qwen3.5-4B-Claude-4.6-Opus-Reasoning-Distilled
language: en
pipeline_tag: text-generation
library_name: mlx
tags:

  • mlx
  • abliterated
  • uncensored
  • qwen3.5
  • reasoning
  • claude-distilled
  • mxfp8
  • quantized
    license: apache-2.0

Jackrong Qwen3.5-4B Claude Reasoning - Abliterated (MXFP8 MLX)

This is an abliterated (uncensored) and MXFP8 quantized version of Jackrong/Qwen3.5-4B-Claude-4.6-Opus-Reasoning-Distilled, converted to MLX format for Apple Silicon.

What is this model?

The base model is Qwen3.5-4B fine-tuned on Claude 4.6 Opus reasoning data using Unsloth/LoRA, giving it strong chain-of-thought reasoning capabilities. This abliterated version removes safety refusal behaviors while preserving the model's reasoning abilities. MXFP8 quantization reduces model size by ~48% with minimal quality loss.

Abliteration Details

  • Method: lukey03 (1 direction, norm-preserving, 3 refinement passes, output-only projection)
  • Post-processing: LoRA compliance fine-tuning (rank=64, 80 iterations, 32 training examples)
  • Quantization: MXFP8 (~8.25 bits per weight, ~4.1 GB)
  • Refusal rate: ~12% on 100 adversarial prompts (85% compliance, measured pre-quantization)
  • Tool: OBLITERATUS

Architecture

Qwen3.5 hybrid attention architecture:

  • 32 layers (8 full attention + 24 linear attention with GatedDeltaNet)
  • 2560 hidden size, 16 attention heads, 4 KV heads
  • ~4B parameters
  • 262K max context length

Usage

pip install mlx-lm

# Generate text
mlx_lm.generate --model AITRADER/Jackrong-Qwen3.5-4B-Claude-Reasoning-abliterated-mxfp8-MLX --prompt "Explain quantum computing"
from mlx_lm import load, generate

model, tokenizer = load("AITRADER/Jackrong-Qwen3.5-4B-Claude-Reasoning-abliterated-mxfp8-MLX")
response = generate(model, tokenizer, prompt="Your prompt here", max_tokens=512)

Full Precision Version

A bf16 full-precision version (~7.9 GB) is available at: AITRADER/Jackrong-Qwen3.5-4B-Claude-Reasoning-abliterated-fp16-MLX

Disclaimer

This model is provided for research purposes. Users are responsible for ensuring their use complies with applicable laws and regulations.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-03-08Upload folder using huggingface_hub1947d282.4 KB
    Loading...

Discussions 1 thread

  1. 2026-03-09Error when loading model: ValueError: Missing 297 parameters:open1 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration