← back to catalog · registered 2026-10-11 13:59

WaveCut/Gemma4-26B-A4B-QAT-Uncensored-HauhauCS-Balanced-MTP-NInfer-v3

WaveCut 26B GGUF MoE second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/WaveCut%2FGemma4-26B-A4B-QAT-Uncensored-HauhauCS-Balanced-MTP-NInfer-v3"
Response includes
  • classification m-uncensored
  • files 6
  • author_summary 5 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M-U
Primary method

Uncensored (method unknown)

No other method signals detected in this model.
Confidence
LOW
Why this label 3 signals
Weak or ambiguous signals. Best guess based on catalog patterns; treat as tentative and check the evidence below.
  • 'uncensored' in name/tags but no 'abliterated' marker
  • method not identifiable from author declaration alone
  • may be DPO fine-tune, prompt engineering, or unknown technique
Refusal direction extraction

No specific extraction method could be identified for this model. The producer either did not document it or used a proprietary pipeline.

What is a refusal direction? →
Downloads · 30-day
0
Likes
0
Model age
today
created 2026-10-11

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
gemma
Tags
ninfer gguf gemma4 qat moe mtp speculative-decoding vision multimodal uncensored rtx-3090 image-text-to-text

Related

Total size
0 B
Files
6
Quantizations
1
Registered
2026-10-11 13:59
Last updated on HF
2026-10-11 15:09

Files by quantization

Auxiliary files 6 files 15.9 GB
Gemma4-26B-A4B-QAT-Uncensored-HauhauCS-Balanced-MTP-ninfer-v3.ninfer 15.9 GB f198ebce download
Gemma4-26B-A4B-QAT-Uncensored-HauhauCS-Balanced-MTP-ninfer-v3.ninfer.conversion.json 233 KB a832b4d1 download
README.md 8.24 KB ed93c3da download
.gitattributes 1.59 KB cb331a00 download
NOTICE 1.32 KB 25af735d download
SHA256SUMS 135 B 7f5cf6ea download

README current version from Hugging Face


license: gemma
base_model:

  • HauhauCS/Gemma4-26B-A4B-QAT-Uncensored-HauhauCS-Balanced-MTP
    base_model_relation: quantized
    library_name: ninfer
    pipeline_tag: text-generation
    tags:
  • ninfer
  • gguf
  • gemma4
  • qat
  • moe
  • mtp
  • speculative-decoding
  • uncensored
  • rtx-3090

Gemma4 26B-A4B QAT Uncensored (HauhauCS Balanced) with MTP, NInfer v3 artifact

A NInfer v3 artifact of HauhauCS's
Gemma4-26B-A4B-QAT-Uncensored-HauhauCS-Balanced-MTP
(Gemma 4 26B-A4B it QAT with refusals removed, the Balanced variant, Q4_K_M) together with the
drafter the release ships, for NInfer-all, the master branch of
iamwavecut/ninfer-all. Every matrix keeps the release's
ggml blocks byte for byte, each in its own format, so nothing is requantized; the engine multiplies
the blocks in place and drafts with the release's MTP head (--spec mtp). Stock NInfer builds
refuse this file: the Gemma 4 family and the gguf_* formats exist only in that line, mixed-format
Gemma GGUFs and the drafter from commit
03bd4006
on. The family was developed and measured on the RTX 3090; the RTX 40 and RTX 50 builds compile
it, unmeasured.

It runs the text model (the release's vision projector is not included): one to eight concurrent
requests with prompt-prefix reuse, structured output, tool calls in Gemma's markup returned as
OpenAI tool calls, and reasoning in its own channel when a request turns thinking on.

What is inside

component representation
attention, 25 sliding-window and 5 full-attention blocks Q, K and output Q4_K; V Q6_K (13 blocks) or Q4_K (12); a block whose Q, K and V differ in format keeps each its own
dense MLP gate and up Q4_K (one parent), down Q8_0 (14 blocks) or Q5_0 (16)
128 routed experts per block [gate; up] Q4_K, down Q8_0 (14 blocks) or Q5_0 (16), expert-major
token table and tied head Q6_K
norms, router, router scale, expert scales, layer scalars FP32
drafter (mtp component) the release's mtp-gemma-4-26B-A4B-it.gguf (Q4_0 with its own token table; the release credits it to Unsloth's Gemma 4 GGUFs, and it differs from Unsloth's current QAT drafter file): Gemma 4's four-block assistant of width 1024 reading the model's own KV cache
tokenizer, chat template google/gemma-4-26B-A4B-it at revision 4d7ae498: Google's canonical chat template, newer than the one inside the release's GGUF
generation defaults the release card's recommendation: temperature 0.6, top-k 64, top-p 0.9, min-p 0.05; its repeat_penalty 1.1 has no NInfer equivalent (requests may set presence or frequency penalties)

One .ninfer file of 17,048,711,936 bytes (15.88 GiB; SHA-256
f198ebce11cfb925fbaa84f6831bd8759e3c570fe36d43b981e6c627e0f2be90), next to its conversion report,
SHA256SUMS and NOTICE: 705 parameters in Q4_K (133 objects), Q8_0 (28), Q5_0 (32), Q4_0 (19,
the drafter), Q6_K (14) and FP32 (416).

Conversion command, from the repository's tree at 03bd4006, with the release at revision
f9093662 (Gemma4-26B-A4B-QAT-Uncensored-HauhauCS-Balanced-Q4_K_M.gguf, sha256 3c131334…;
mtp-gemma-4-26B-A4B-it.gguf, sha256 62bd3af7…) and Google's text files at 4d7ae498:

mkdir -p hf
for f in config.json tokenizer.json tokenizer_config.json chat_template.jinja generation_config.json; do
  curl -fsSL -o hf/$f https://huggingface.co/google/gemma-4-26B-A4B-it/resolve/main/$f
done
hf download HauhauCS/Gemma4-26B-A4B-QAT-Uncensored-HauhauCS-Balanced-MTP \
  Gemma4-26B-A4B-QAT-Uncensored-HauhauCS-Balanced-Q4_K_M.gguf mtp-gemma-4-26B-A4B-it.gguf --local-dir .
# The release's sampling as the artifact's defaults.
jq '. + {do_sample: true, temperature: 0.6, top_k: 64, top_p: 0.9, min_p: 0.05}' \
  hf/generation_config.json > generation_config.json
python3 -m tools.convert --model hf --recipe gemma4_gguf --components text,mtp \
  --source gguf=Gemma4-26B-A4B-QAT-Uncensored-HauhauCS-Balanced-Q4_K_M.gguf \
  --source mtp=mtp-gemma-4-26B-A4B-it.gguf --resource generation_config.json=generation_config.json \
  --name Gemma4-26B-A4B-QAT-Uncensored-HauhauCS-Balanced-MTP \
  --out Gemma4-26B-A4B-QAT-Uncensored-HauhauCS-Balanced-MTP-ninfer-v3.ninfer --device cpu

The conversion takes under a minute.

Running

Four concurrent requests over 40K-token contexts with two drafts a round, from the
Docker image (an NVIDIA driver of the CUDA 13
branch, 580 or newer, and the NVIDIA Container Toolkit; images from sha-03bd40060ed1 on read
this file):

hf download WaveCut/Gemma4-26B-A4B-QAT-Uncensored-HauhauCS-Balanced-MTP-NInfer-v3 \
  Gemma4-26B-A4B-QAT-Uncensored-HauhauCS-Balanced-MTP-ninfer-v3.ninfer --local-dir models
docker run --rm --gpus all -p 8080:8080 --ulimit memlock=-1 \
  -v "$PWD/models:/models" -v ninfer-cache:/cache \
  ghcr.io/iamwavecut/ninfer-all serve /models/Gemma4-26B-A4B-QAT-Uncensored-HauhauCS-Balanced-MTP-ninfer-v3.ninfer \
  --model-id gemma4-hauhau --max-context 40960 --max-concurrency 4 --spec mtp --draft-tokens 2

The server answers OpenAI and Anthropic requests on http://localhost:8080/v1. A build of master
(its README; CMAKE_CUDA_ARCHITECTURES is
86, 89 or 120a, and 120a needs CUDA 13.1 or newer) takes the same flags:

ninfer-serve models/Gemma4-26B-A4B-QAT-Uncensored-HauhauCS-Balanced-MTP-ninfer-v3.ninfer \
  --model-id gemma4-hauhau --max-context 40960 --max-concurrency 4 --spec mtp --draft-tokens 2

The weights take 15.6 GiB on the GPU (0.25 GiB more for the drafter with --spec mtp) and the
prompt workspace about 1.1 GiB; a sequence holds 250 MiB for the sliding windows with --spec mtp
(200 MiB without) and 10 KiB per token of context for the full-attention blocks. The release card
describes the model as reasoning before it answers; NInfer follows the chat template, where thinking
is off unless a request asks (chat_template_kwargs: {"enable_thinking": true}).

Speculative decoding. The drafter proposes tokens that the model verifies in one pass, keeping
a draft only while it equals the request's own sample: the output is the one plain decoding gives,
and the model's behaviour is unchanged by the drafter. A request decoding alone takes the
configured drafts, two or three at once one draft each, four or more decode plainly.

Quality

The conversion keeps every block of the release byte for byte; the same build's kernels are
checked against FP64 oracles and its assistant path against an FP64 transcription of transformers'
Gemma 4 drafting (see the repository's
Gemma 4 notes). Perplexity
over eleven 8,192-token windows of a NInfer source text: 21.79 (Google's QAT Q4_0 release on the same
windows: 24.79). Before publication this file served greedy chat requests (code, a
Russian answer, a 13,600-token prompt) plainly and with two drafts (acceptance 0.63), and the
CLI. Reasoning benchmarks were not run.

Speed

October 2026, one RTX 3090 (24 GB) at 350 W, CUDA 12.8, a build of master at 03bd4006: prefill
6,113 tok/s for a 2,048-token prompt and 5,623 for 10,000; decode 206.2 tok/s from an empty context
and 182.7 after 10,000 tokens. Ten chat requests (code, explanation, summary, QA, translation, math,
a story, a Russian answer, a 13,600-token source excerpt), 384 greedy tokens each, one at a time,
decode tok/s:

all ten 13,600-token prompt
plain 200.3 179.3
--draft-tokens 2 (acceptance 0.69) 275.9 225.5
--draft-tokens 3 276.4

Credits and license

Uncensored weights, quantization and drafter GGUF: HauhauCS, whose release declares the Gemma license
(license: gemma), kept here; the drafter, per the release, comes from Unsloth's Gemma 4 GGUFs.
Base model, tokenizer and chat template: Google, Gemma 4, under the
Gemma 4 terms (Apache 2.0). Attribution notices
are collected in NOTICE. This repository only re-packs those weights into NInfer's
container.

Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration