← back to catalog · registered 2026-08-22 13:56

Noobneophyte/supergemma4-26b-uncensored-mlx-4bit-v2

Noobneophyte Gemma 25B MoE
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/Noobneophyte%2Fsupergemma4-26b-uncensored-mlx-4bit-v2"
Response includes
  • classification m-uncensored
  • files 15
  • benchmarks 11 entries
  • hub_downloads_all_time 1,303
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
1K
51 last 30d - cooling
Likes
0
Model age
4mo ago
created 2026-05-22
Downloads over time
Now1.3K→from103↑1,178%
425079721.4K103 on May 201.3K on Oct 11MayJunJulAugSepOct
May 20 → Oct 11 · 60 snapshots · spans 144 days

Benchmarks

Portrait before abliteration
Benchmarks of the base model as it stood before the refusal-removal operation. Compare with the numbers above to see what the operation cost.
Benchmark Score Source
Entertainment 2.2 UGI
Hazardous 2.9 UGI
Natural Intelligence 34.44 UGI
Political lean -18.2% UGI
Sensitive-Info 22.41 UGI
SocPol 1.8 UGI
UGI 20.77 UGI
Willingness (10) 1.8 UGI
W10-Adherence 1.5 UGI
W10-Direct 2 UGI
Writing 41.62 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
gemma
Languages
en ko
Tags
mlx safetensors gemma4 uncensored apple-silicon 4bit quantized reasoning tool-use coding browser-automation korean

Related

Total size
13.2 GB
Files
15
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-05-22 13:32

Files by quantization

Auxiliary files 15 files 13.3 GB
model-00002-of-00003.safetensors 4.99 GB ff1727f8 download
model-00001-of-00003.safetensors 4.95 GB 1900a82a download
model-00003-of-00003.safetensors 3.28 GB b2fb4eb8 download
tokenizer.json 30.7 MB cc8d3a0c download
model.safetensors.index.json 136 KB 773ce92c download
benchmark_quick_bench_20260412_responses.jsonl 126 KB bfa35ccb download
chat_template.jinja 16.1 KB 98da08eb download
tool_chat_template.jinja 14.8 KB 42e1cfad download
config.json 10.2 KB 4030beae download
benchmark_quick_bench_20260412.json 6.87 KB fae559cb download
README.md 3.54 KB b861f667 download
tokenizer_config.json 2.73 KB bfb245eb download
.gitattributes 1.53 KB 52373fe2 download
SERVING_NOTES.md 553 B 96ebba54 download
generation_config.json 203 B edda3c19 download

README current version from Hugging Face


license: gemma
base_model: google/gemma-4-26B-A4B-it
tags:

  • gemma4
  • mlx
  • uncensored
  • apple-silicon
  • 4bit
  • quantized
  • reasoning
  • tool-use
  • coding
  • browser-automation
  • korean
  • fast
    language:
  • en
  • ko
    pipeline_tag: text-generation
    library_name: mlx

SuperGemma4-26B-Uncensored-Fast v2

A faster, sharper, uncensored Gemma 4 26B for Apple Silicon.

This is the text-only flagship for people who want the core trade-off to be obvious at a glance:

  • smarter than stock Gemma 4 26B IT on real local agent tasks
  • faster than the stock local 4-bit baseline on the same machine
  • uncensored, without falling apart on code, tool-use, or Korean prompts

Why this model

If you want the fast line instead of the multimodal line, this is the one to run.

  • Fast is part of the release identity, not just a minor variant
  • Uncensored behavior is preserved while practical capability goes up
  • Strong at code, browser tasks, tool-use, planning, and Korean
  • Tuned for local agent workloads on Apple Silicon MLX

Headline numbers

Metric Gemma 4 26B IT original 4bit SuperGemma Fast
Quick bench overall 91.4 95.8
Avg generation speed 42.5 tok/s 46.2 tok/s
Delta overall baseline +4.4
Delta speed baseline +8.7%

Category gains vs original

Category Original SuperGemma Fast Delta
Code 92.3 98.6 +6.3
Browser 87.5 89.6 +2.1
Logic 86.9 95.2 +8.3
System Design 97.8 98.9 +1.1
Korean 90.7 95.0 +4.3

What makes it attractive

  • Beats the stock local 4-bit baseline in both quality and speed
  • Produces stronger code, stronger reasoning, and more useful tool-oriented answers
  • Handles Korean and agent-style prompts better than the original local run
  • Keeps the uncensored feel without turning unstable or collapsing into broken outputs
  • Built to feel immediately stronger in real usage, not just in a niche benchmark

Base and format

  • Base model: google/gemma-4-26B-A4B-it
  • Format: MLX 4-bit
  • Size: about 13GB
  • Best use case: fast text-only local agent model with stronger practical capability than stock Gemma 4

Why it is better than stock

  • Higher quick-bench overall score: 95.8 vs 91.4
  • Faster average generation speed: 46.2 tok/s vs 42.5 tok/s
  • Bigger gains where local agents actually benefit:
    • Code: +6.3
    • Logic: +8.3
    • Korean: +4.3
    • Browser workflows: +2.1
  • Uncensored behavior remains a core property of the release instead of being layered on after the fact

Recommended launch

mlx_lm.server \
  --model Jiunsong/supergemma4-26b-uncensored-mlx-4bit-v2 \
  --port 8080

For OpenAI-compatible serving, let mlx_lm.server auto-detect the bundled template.

Do not pass --chat-template /path/to/chat_template.jinja as a literal path string on launch paths that expect the template body. That can corrupt responses.

Quick test

mlx_lm.generate \
  --model Jiunsong/supergemma4-26b-uncensored-mlx-4bit-v2 \
  --prompt "Write a Python function that returns prime numbers up to n." \
  --max-tokens 512

Included files

  • benchmark_quick_bench_20260412.json
  • benchmark_quick_bench_20260412_responses.jsonl
  • SERVING_NOTES.md

Notes

  • This is the fast text-only line.
  • The earlier "reasoning is broken" report reproduced as a serving-template launch issue, not as weight corruption.
  • Re-fused and re-benchmarked locally before upload.

README history 1 version

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-05-22Duplicate from Jiunsong/supergemma4-26b-uncensored-mlx-4bit-v280b0c683.5 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration