← back to catalog · registered 2026-08-22 13:56

Sikaworld1990/gemma-3-12b-qat-abliterated-sikaworld-fp4-ltx2

Sikaworld1990 Gemma 12B second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/Sikaworld1990%2Fgemma-3-12b-qat-abliterated-sikaworld-fp4-ltx2"
Response includes
  • classification m1
  • files 4
  • author_summary 2 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · 30-day
0
Likes
40
Model age
7mo ago
created 2026-03-11
Downloads over time
Now0→from0↑0%
00110 on Mar 110 on Oct 11MarAprMayJunJulAugSepOct
Mar 11 → Oct 11 · 70 snapshots · spans 214 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
apache-2.0
Languages
en
Tags
nvfp4 gemma-3 comfyui abliterated qat ltx-2 ltx-2.3 uncensored feature-extraction en base_model:mlabonne/gemma-3-12b-it-qat-abliterated base_model:finetune:mlabonne/gemma-3-12b-it-qat-abliterated

Related

Total size
19.6 GB
Files
4
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-03-11 17:20

Files by quantization

Auxiliary files 4 files 19.6 GB
Gemma3-12B-NVFP4-Sikaworld-HF.safetensors 11.3 GB 3d5ad04b download
Gemma3-12B-NVFP4-Sikaworld-Pure.safetensors 8.30 GB 973a4a14 download
README.md 5.07 KB e0c3285c download
.gitattributes 1.48 KB a6344aac download

README current version from Hugging Face


license: apache-2.0
pipeline_tag: feature-extraction
base_model: mlabonne/gemma-3-12b-it-qat-abliterated
tags:

  • nvfp4
  • gemma-3
  • comfyui
  • abliterated
  • qat
  • ltx-2
  • ltx-2.3
  • uncensored
    model_type: gemma
    language:
  • en

🌍 Gemma‑3‑12B‑QAT‑Abliterated — Sikaworld FP4 Editions

Blackwell‑optimized FP4 text encoders for LTX‑2 and 2.3, based on mlabonne’s improved Abliteration technique.

🌐 Overview

The NVIDIA Blackwell architecture update introduced first‑class support for FP4/NVFP4 inference, enabling extremely fast and memory‑efficient text encoders. At the same time, the LTX‑2 development team officially recommends Gemma‑QAT‑based encoders for video generation due to their stable activation distributions, strong semantic gradients, and robust temporal behavior.

This repository provides two custom FP4 variants of the uncensored Gemma‑3‑12B‑QAT model created by mlabonne using his improved Abliteration v2 method.

Both models are fully uncensored, explicitly optimized for LTX‑2 and of course LTX-2.3, and designed to deliver strong motion vectors while maintaining spatial coherence.


📦 The Two FP4 Editions

🛡️ FP4 High‑Fidelity Edition (Protected Layers) [Recommended]

This version uses a surgical mixed‑precision stabilizer to preserve facial symmetry and spatial coherence.

  • Layers 0–1 (Input embeddings) kept in BF16.
  • Layers 44–47 (Final output projections) kept in BF16.
  • All LayerNorms and Biases kept in BF16.
  • All mid-transformer layers quantized to FP4.

Best for: Maximum stability, minimal facial drift, consistent anatomy, and strong but mathematically controlled motion vectors. Highly recommended for complex I2V/T2V tasks.

🚀 FP4 Pure Edition (No Protected Layers)

This version is a relentless, flat FP4/NVFP4 quantization of the Abliterated QAT model.

  • All transformer layers (0-47) quantized to FP4.
  • Only LayerNorms and Biases remain in BF16.

Best for: Maximum performance, the absolute lowest VRAM footprint, and the fastest inference on Blackwell GPUs. It trades a tiny amount of spatial stability for raw speed and more intense, aggressive motion vectors.


🧰 Usage in ComfyUI

  1. Download your preferred .safetensors file.
  2. Place the file inside your ComfyUI models folder:
    ComfyUI/models/text_encoders/
  3. Load the model via the standard DualCLIPLoader or LTX‑2 Text Encoder Loader.
  4. Recommended dtype: fp8_e4m3fn (Note: The BF16‑protected layers will automatically be respected and kept in BF16 by ComfyUI's loader).

💡 Prompting Tip: Start your prompts with direct action verbs (e.g., "running", "falling", "embracing", "exploding"). FP4 models respond extremely well to dynamic, upfront phrasing.


🔬 Technical Background

Why Gemma‑QAT for LTX‑2?

The LTX‑2 base model architecture reacts very sensitively to the text encoder's conditioning. The LTX‑team recommends QAT (Quantization-Aware Training) encoders because they provide:

  • Stable activation distributions
  • Smooth residual streams
  • Strong temporal gradients
  • Robust spatial alignment
  • Heavily reduced “frozen video” (motion collapse) behavior

The Abliteration V2 Magic

These models are derived from mlabonne/gemma-3-12b-it-qat-abliterated. Abliteration is a multi‑step orthogonalization process, not just a simple deletion. It compares residual streams from harmful vs. harmless samples, computes a "refusal direction", and subtracts this direction natively from the hidden states of target modules. The result is a fully uncensored, high‑fidelity instruction model with loud and uninhibited semantic gradients — acting as the perfect cure for static/frozen LTX‑2 generations.

Why FP4 for Blackwell GPUs?

NVIDIA's latest Blackwell Tensor Cores are explicitly optimized for FP4/NVFP4 mathematical operations. This format offers:

  • Significantly higher throughput than FP8
  • Extremely low VRAM footprint
  • Faster long‑prompt (prefill) inference
  • Decreased pressure on memory bandwidth

These FP4 editions feature a pure FP4 tensor layout (with appropriate micro-block and global scales) fully compatible with NVFP4 hardware acceleration on RTX 50‑series and data center hardware.


📊 Technical Summary

Component 🛡️ High‑Fidelity Edition 🚀 Pure Edition
Base Model mlabonne/gemma‑3‑12b‑it‑qat‑abliterated mlabonne/gemma‑3‑12b‑it‑qat‑abliterated
Quantization FP4 + BF16 stabilizer Pure FP4
Protected Layers 0–1, 44–47 None
Norms & Biases BF16 BF16
Inference Speed Fast Fastest
Stability Highest Moderate
VRAM Usage Low Lowest

--

🏷️ Credits & Acknowledgments

  • Base Model & Abliteration v2: mlabonne
  • QAT Architecture & Gemma Weights: Google
  • FP4 Optimization, Hybrid Architecture & Stabilization: Sikaworld
  • LTX‑2 & QAT Recommendation: Lightricks / LTX‑Team

README history 14 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-03-11Update README.mdb62d8e95.1 KB
    Loading...
  2. 2026-03-11Update README.mdb1f2fbe5.1 KB
    Loading...
  3. 2026-03-11Update README.md8d831015 KB
    Loading...
  4. 2026-03-11Update README.md47d6e895.1 KB
    Loading...
  5. 2026-03-11Update README.mde3b65f85 KB
    Loading...
  6. 2026-03-11Update README.mdaafab155.1 KB
    Loading...
  7. 2026-03-11Update README.md22b1f085.2 KB
    Loading...
  8. 2026-03-11Update README.md0d832be5.1 KB
    Loading...
  9. 2026-03-11Update README.md9a8a3d85.3 KB
    Loading...
  10. 2026-03-11Update README.md3ccc3f95.1 KB
    Loading...
  11. 2026-03-11Update README.mdd73f0d35.1 KB
    Loading...
  12. 2026-03-11Update README.md1a938784.5 KB
    Loading...
  13. 2026-03-11Update README.md47984824.5 KB
    Loading...
  14. 2026-03-11initial commit4d0976d28 B
    Loading...

Discussions 2 threads

  1. 2026-04-06Image to Video pooropen3 💬#2
    Loading...
  2. 2026-03-12Difference with the previous Heretic "x"open3 💬#1
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration