← back to catalog · registered 2026-08-22 13:56

Vastopian/gemma-4-26B-A4B-it-abliterated-GGUF

Vastopian Gemma 26B GGUF MoE second-order 262K ctx
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/Vastopian%2Fgemma-4-26B-A4B-it-abliterated-GGUF"
Response includes
  • classification m8
  • files 7
  • hub_downloads_all_time 5,972
  • author_summary 1 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M8
Primary method

Repackaging (quantization)

Applied on top of direct removal inherited from the base model.
Confidence
MEDIUM
Inherited from base model
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=1
  • assume M1 (base ablation) + M8 (GGUF quant) - default when producer unknown
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · lifetime
6K
656 last 30d - stable
Likes
3
Model age
6mo ago
created 2026-04-10
Downloads over time
Now6.3K→from3.9K↑62%
3.8K4.7K5.6K6.5K3.9K on Apr 156.3K on Oct 11AprMayJunJulAugSepOct
Apr 15 → Oct 11 · 65 snapshots · spans 179 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
gemma
Quantizations
IQ4 Q2_K Q3_K Q4_K Q6_K
Tags
gguf gemma-4 abliterated moe quantized roleplay base_model:WWTCyberLab/gemma-4-26B-A4B-it-abliterated base_model:quantized:WWTCyberLab/gemma-4-26B-A4B-it-abliterated license:gemma endpoints_compatible region:us conversational

Related

Total size
72.5 GB
Files
7
Quantizations
6
Registered
2026-08-22 13:56
Last updated on HF
2026-04-10 01:35

Files by quantization

Q6_K 1 file 21.1 GB
gemma-4-26B-Q6_K.gguf 21.1 GB 1e76b1f9 download
Q4_K 1 file 15.6 GB
gemma-4-26B-Q4_K_M.gguf 15.6 GB a7508d22 download
IQ4 1 file 13.6 GB
gemma-4-26B-IQ4_NL.gguf 13.6 GB 768a89b9 download
Q3_K 1 file 12.4 GB
gemma-4-26B-Q3_K_M.gguf 12.4 GB 79947114 download
Q2_K 1 file 9.86 GB
gemma-4-26B-Q2_K.gguf 9.86 GB 91e21b77 download
Auxiliary files 2 files 3.67 KB
README.md 1.90 KB 7a4d49cf download
.gitattributes 1.77 KB 42c3e010 download

README current version from Hugging Face


base_model: WWTCyberLab/gemma-4-26B-A4B-it-abliterated
library_name: gguf
license: gemma
tags:

  • gemma-4
  • abliterated
  • moe
  • quantized
  • roleplay

Gemma-4-26B-A4B-it-abliterated - GGUF

This repo contains GGUF format model files for WWTCyberLab/gemma-4-26B-A4B-it-abliterated.

These quants were generated using llama.cpp and are optimized for local inference on consumer hardware.

Available Quants

Filename Quant Type Size Description
gemma-4-26B-IQ4_NL.gguf IQ4_NL ~14.6 GB Recommended. High quality 4-bit, best balance of intelligence and size.
gemma-4-26B-Q6_K.gguf Q6_K ~22.9 GB Near-lossless. Requires 24GB+ VRAM for full GPU offload.
gemma-4-26B-Q4_K_M.gguf Q4_K_M ~16.9 GB Standard balanced quant. High compatibility.
gemma-4-26B-Q3_K_M.gguf Q3_K_M ~12.5 GB Good for 12GB VRAM cards or high-context usage.
gemma-4-26B-Q2_K.gguf Q2_K ~10.5 GB Maximum compression. Usable for basic logic and data retrieval.

Model Description

This is an abliterated version of Gemma 4 26B, meaning the safety-alignment (refusals) has been substantially removed for research and unrestricted creative use.

  • Architecture: Mixture of Experts (MoE)
  • Optimization: A4B (Architecture for 4-Bit)
  • Quality (QPS): 107%+ (Quality improved via ablation)

Usage with llama.cpp

To run the IQ4_NL version on an RTX 5060 Ti (16GB), use the following command for optimal VRAM usage:

llama-cli -m gemma-4-26B-IQ4_NL.gguf -ngl 99 -c 8000 --flash-attn on --cache-type-k q4_0 --cache-type-v q4_0 --reasoning-budget 4096

Credits
Original Model Creator: WWTCyberLab

Quantization & Testing: Vastopian

Base Architecture: Google Gemma

Disclaimer: This model is abliterated and has no safety filters. Users are solely responsible for any content generated.

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-04-10Update README.mdbef933a1.9 KB
    Loading...
  2. 2026-04-10Update README.md9312cce1.9 KB
    Loading...
  3. 2026-04-10Create README.md88aa2531.7 KB
    Loading...

Discussions 1 thread

  1. 2026-04-10It abliteration doesnt seem to work:open2 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration