← back to catalog · registered 2026-08-22 13:56

Lookoff/gemma-4-26b-a4b-it-uncensored-4bit

Lookoff Gemma 25B MoE second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/Lookoff%2Fgemma-4-26b-a4b-it-uncensored-4bit"
Response includes
  • classification m1
  • files 11
  • hub_downloads_all_time 1,128
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
1K
286 last 30d - stable
Likes
1
Model age
7w ago
created 2026-08-16
Downloads over time
Now1.3K→from495↑153%
4577471K1.3K495 on Aug 191.3K on Oct 11AugSepOct
Aug 19 → Oct 11 · 48 snapshots · spans 53 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
gemma
Languages
en
Tags
mlx safetensors gemma4 gemma-4 mixture-of-experts moe abliterated uncensored 4bit quantized apple-silicon macos

Related

Total size
13.2 GB
Files
11
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-08-16 21:20

Files by quantization

Auxiliary files 11 files 13.3 GB
model-00002-of-00003.safetensors 4.99 GB 1aedb165 download
model-00001-of-00003.safetensors 4.95 GB 1900a82a download
model-00003-of-00003.safetensors 3.28 GB 37ed8feb download
tokenizer.json 30.7 MB cc8d3a0c download
model.safetensors.index.json 136 KB 773ce92c download
config.json 12.0 KB 50fdaa82 download
chat_template.jinja 11.8 KB 33c51c2d download
README.md 5.92 KB 8f10c671 download
tokenizer_config.json 2.76 KB fd34b1f4 download
.gitattributes 1.53 KB 52373fe2 download
generation_config.json 203 B edda3c19 download

README current version from Hugging Face


license: gemma
base_model: TrevorJS/gemma-4-26B-A4B-it-uncensored
library_name: mlx
pipeline_tag: text-generation
language:

  • en
    tags:
  • mlx
  • gemma4
  • gemma-4
  • mixture-of-experts
  • moe
  • abliterated
  • uncensored
  • 4bit
  • quantized
  • apple-silicon
  • macos
  • local-llm
  • text-generation
  • turbofieldfare

Gemma 4 26B Uncensored — TurboFieldfare-Compatible 4-bit Checkpoint

Run an uncensored Gemma 4 26B-A4B model locally on Apple Silicon with approximately 2 GB of active RAM using TurboFieldfare Uncensored.

This is a purpose-built and validated 4-bit checkpoint for TurboFieldfare's Swift/Metal expert-streaming runtime. It uses the quantization recipe and tensor layout expected by TurboFieldfare: affine 4-bit weights with group size 64 and an 8-bit router projection.

[!IMPORTANT]
The complete checkpoint occupies approximately 14.3 GB on disk. The approximately 2 GB RAM figure refers to the resident weights and KV cache used by TurboFieldfare while routed experts are streamed from SSD. Loading the complete checkpoint normally with MLX requires substantially more memory.

Why this checkpoint exists

TurboFieldfare does not simply load an arbitrary 4-bit model. Its hand-written Metal kernels and streaming installer expect a specific weight representation.

This checkpoint was produced and validated for:

  • affine 4-bit quantization with group size 64;
  • 4-bit embeddings, attention, and shared and routed expert weights;
  • an 8-bit router.proj;
  • the tensor structure consumed by TurboFieldfare's installer and Metal kernels;
  • SSD-backed routed-expert streaming on Apple Silicon.

Generic MLX or GGUF quantizations should not be assumed to be drop-in TurboFieldfare inputs, even when they use the same base model and nominal bit width.

Quick start with TurboFieldfare

Requirements

  • Apple Silicon Mac
  • macOS 26 with Metal 4
  • Xcode 26 and Swift 6.2 or newer
  • approximately 15 GB available for download and 14.3 GB for the installed checkpoint

Clone and build the uncensored fork:

git clone https://github.com/Lookoff-AIMLAPI/turbo-fieldfare-uncensored.git
cd turbo-fieldfare-uncensored
swift build -c release

Install this checkpoint in TurboFieldfare's streaming format:

.build/release/TurboFieldfareRepack \
  --variant uncensored \
  --output scratch/gemma4-uncensored.gturbo

Run it with the CLI:

.build/release/TurboFieldfareCLI \
  --model scratch/gemma4-uncensored.gturbo \
  --messages-file msgs.json

Or launch the native Mac app with the uncensored variant selected:

TURBOFIELDFARE_VARIANT=uncensored .build/release/TurboFieldfareMac

See the TurboFieldfare Uncensored repository for the runtime, local OpenAI-compatible server, benchmarks, and complete usage documentation.

Use with plain MLX

The checkpoint can also be loaded directly with mlx-lm:

pip install -U mlx-lm
mlx_lm.chat --model Lookoff/gemma-4-26b-a4b-it-uncensored-4bit

Plain MLX loads the checkpoint normally and does not provide TurboFieldfare's approximately 2 GB active-RAM behavior.

Checkpoint details

Property Value
Architecture Gemma 4 26B-A4B Mixture of Experts
Behavior Instruction-tuned, abliterated / reduced refusals
Format MLX Safetensors
Weight quantization Affine 4-bit
Group size 64
Router projection 8-bit
Installed size Approximately 14.3 GB
TurboFieldfare active RAM Approximately 2 GB of resident weights and 4K KV cache
Target runtime TurboFieldfare on Apple Silicon
Secondary runtime mlx-lm with normal full-checkpoint memory behavior

The RAM figure is runtime-specific. Prompt length, selected context size, KV-cache configuration, expert-cache settings, page-cache state, and macOS memory pressure affect observed memory use and performance.

Provenance

  1. Architecture and original weights: google/gemma-4-26b-a4b, distributed under the Gemma Terms of Use.
  2. Instruction tuning and abliteration source: TrevorJS/gemma-4-26B-A4B-it-uncensored. Its model card reports biprojection plus Expert-Granular Abliteration, a 0.7% refusal rate, and KL divergence of 0.09 relative to the base checkpoint.
  3. This repository: custom affine quantization and packaging matching the 4-bit/group-64 layout required by TurboFieldfare, with router.proj retained at 8-bit.

For structural compatibility, the quantization layout follows the official mlx-community/gemma-4-26b-a4b-it-4bit release while applying it to the uncensored source checkpoint.

Scope and limitations

  • This repository contains model weights, not the TurboFieldfare runtime itself.
  • TurboFieldfare's current path is text-only; images, audio, and video are not supported.
  • Reduced refusal behavior does not guarantee correctness, neutrality, or suitability for a particular use.
  • Generated output may be inaccurate, harmful, or otherwise inappropriate. Evaluate important outputs independently.
  • Performance varies by Mac, SSD, prompt, cache state, and runtime settings.

License and responsible use

The checkpoint inherits the Gemma license and remains governed by the Gemma Terms of Use and Gemma Prohibited Use Policy.

The model has reduced refusal behavior compared with the aligned instruction checkpoint and may produce content that the original model would decline. Users are responsible for complying with applicable law, the model terms, and the prohibited-use policy.

TurboFieldfare is an independent project and is not affiliated with, sponsored by, or endorsed by Google.

README history 2 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-08-16docs: clarify TurboFieldfare-compatible checkpoint1c165735.9 KB
    Loading...
  2. 2026-08-16Upload uncensored Gemma 4 26B-A4B 4-bit MLX (TurboFieldfare-compatible)fba4e512.4 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration