← back to catalog · registered 2026-08-22 13:56

TitanPythons/Nemotron-Cascade-2-30B-A3B-UNCENSORED-JANG_2L

TitanPythons Nemotron 33B MoE
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/TitanPythons%2FNemotron-Cascade-2-30B-A3B-UNCENSORED-JANG_2L"
Response includes
  • classification m1
  • files 17
  • hub_downloads_all_time 512
  • author_summary 3 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
512
50 last 30d - cooling
Likes
0
Model age
6mo ago
created 2026-04-06
Downloads over time
Now524→from153↑242%
134277419561153 on Apr 15524 on Oct 11AprMayJunJulAugSepOct
Apr 15 → Oct 11 · 65 snapshots · spans 179 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
other
Languages
en
Tags
mlx safetensors nemotron_h jang quantized mixed-precision apple-silicon moe mamba abliterated uncensored crack

Related

Total size
17.0 GB
Files
17
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-03-22 09:52

Files by quantization

Auxiliary files 17 files 17.0 GB
model-00001-of-00004.safetensors 4.95 GB 36d4c728 download
model-00002-of-00004.safetensors 4.68 GB 4bf933db download
model-00003-of-00004.safetensors 4.68 GB 442bfd49 download
model-00004-of-00004.safetensors 2.68 GB 3dd69ed1 download
tokenizer.json 16.3 MB c3da26d4 download
tokenizer_config.json 184 KB 039d6fad download
modeling_nemotron_h.py 81.2 KB acc9d622 download
model.safetensors.index.json 65.3 KB 70347473 download
configuration_nemotron_h.py 12.6 KB 639b30de download
dealign_mascot.png 10.9 KB da3bf39a download
chat_template.jinja 10.7 KB e878c6e1 download
README.md 6.56 KB 1257602f download
config.json 1.82 KB 963ff537 download
.gitattributes 1.53 KB 52373fe2 download
jang_config.json 790 B 6479f2d2 download
special_tokens_map.json 449 B 98662d3b download
generation_config.json 138 B 5691ead3 download

README current version from Hugging Face


language:

  • en
    library_name: mlx
    license: other
    base_model: nvidia/Nemotron-Cascade-2-30B-A3B
    tags:
  • jang
  • quantized
  • mixed-precision
  • apple-silicon
  • mlx
  • moe
  • mamba
  • abliterated
  • uncensored
  • crack
    pipeline_tag: text-generation
    thumbnail: dealign_mascot.png

Important: This model uses the JANG quantization format — the GGUF equivalent for MLX on Apple Silicon. Currently only supported by MLX Studio and the jang-tools Python package.


MLX Studio

MLX Studio App

MLX Studio — the only app that natively supports JANG models


Nemotron Cascade 2 30B — JANG_4M + CRACK

JANG mixed-precision · CRACK abliterated · Mamba + MoE + Attention · No guardrails · 17 GB

Ko-fi


What Is This?

This is NVIDIA Nemotron Cascade 2 30B — a 30B parameter hybrid model with THREE layer types: Mamba-2 SSM + MoE (128 experts, top-6) + Attention. One of the most architecturally advanced small models available.

It has been:

  1. JANG quantized — JANG_4M profile (8-bit attention, 4-bit experts) — 17 GB
  2. CRACK abliterated — permanent weight-level removal of safety refusal
Architecture Nemotron Cascade 2 — 30B total, ~3B active, 3 layer types
Quantization JANG_4M (8/4-bit mixed, 4.1 avg) — 17 GB
HarmBench 99.4% (318/320)
MMLU 82.7% (172/208 with thinking)
Speed ~127 tok/s (M4 Ultra 256GB)
Thinking ON/OFF supported (ChatML)
Fits on 32 GB+ Macs

Also see: JANG_2L version — 10 GB, 99.7% HarmBench, 66.8% MMLU (fits on 16 GB Macs)


HarmBench Results

318/320 (99.4%)

Category Score
API Hacking 100/100 100%
Covering Tracks 20/20 100%
Auth Bypass 99/100 99%
Cloud Exploits 99/100 99%

CRACK vs Base

CRACK Base JANG_4M
MMLU (with thinking) 82.7% 88%
HarmBench 99.4% 0%
Speed ~127 tok/s ~130 tok/s

Surgery reduced MMLU by ~5% — safety guardrails were slightly entangled with reasoning pathways.

MMLU Results (with reasoning recovery)

172/208 (82.7%) — no-think 128/208 (61.5%) + thinking recovered 47

Subject Score
HS Biology 15/16 94%
Conceptual Physics 14/16 88%
World Religions 13/16 81%
College Physics 12/16 75%
HS Geography 12/16 75%
Professional Medicine 12/16 75%
Electrical Engineering 9/16 56%
College CS 8/16 50%
Formal Logic 8/16 50%
College Mathematics 7/16 44%
HS Mathematics 7/16 44%
Abstract Algebra 6/16 38%
Machine Learning 5/16 31%

Scores shown are no-think pass. Thinking recovery improved total from 61.5% to 82.7%.

JANG_4M CRACK vs JANG_4M Base vs JANG_2L CRACK

JANG_4M CRACK JANG_4M Base JANG_2L CRACK
Size 17 GB 17 GB 10 GB
MMLU 82.7% 88% 66.8%
HarmBench 99.4% 0% 99.7%
Speed ~127 tok/s ~130 tok/s ~121 tok/s
Fits on 32 GB Mac 32 GB Mac 16 GB Mac

Install & Usage

pip install "jang[mlx]"
from jang_tools.loader import load_jang_model
from mlx_lm import generate

model, tokenizer = load_jang_model("dealignai/Nemotron-Cascade-2-30B-A3B-JANG_4M-CRACK")

messages = [{"role": "user", "content": "Your prompt here"}]
prompt = tokenizer.apply_chat_template(
    messages, add_generation_prompt=True, tokenize=False)

response = generate(model, tokenizer, prompt=prompt, max_tokens=2000)
print(response)

Thinking Mode

Thinking is ON by default. To disable:

prompt = tokenizer.apply_chat_template(
    messages, add_generation_prompt=True,
    enable_thinking=False, tokenize=False)

About JANG

JANG (Jang Adaptive N-bit Grading) is a mixed-precision quantization format for Apple Silicon — the GGUF equivalent for MLX.

About CRACK

CRACK (Controlled Refusal Ablation via Calibrated Knockouts) removes safety alignment from LLMs at the weight level using per-layer projected vectors from structurally-mirrored prompt pairs.


Links

Ko-fi X/Twitter GitHub MLX Studio Website


Disclaimer

This model is provided for research and educational purposes. The creators are not responsible for any misuse. By downloading this model, you agree to use it responsibly and in compliance with applicable laws.


한국어

Nemotron Cascade 2 30B — JANG_4M + CRACK

항목 내용
크기 17 GB
HarmBench 99.4% (318/320)
MMLU 82.7% (172/208)
속도 ~127 tok/s (M4 Ultra)
최소 요구사양 32 GB 메모리 Mac
pip install "jang[mlx]"

GitHub · HuggingFace · MLX Studio · Ko-fi · X @dealignai


Created by Jinho Jang · 장진호 제작

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-03-22fix: update Twitter/X handle to @dealignai7eec89e6.6 KB
    Loading...
  2. 2026-03-22add: MMLU per-subject, 3-way comparison table (CRACK vs Base vs 2L)8c7fea06.5 KB
    Loading...
  3. 2026-03-22Add files using upload-large-folder tool47405155.5 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration