license: gemma
base_model:
- HauhauCS/Gemma4-26B-A4B-QAT-Uncensored-HauhauCS-Balanced-MTP
base_model_relation: quantized
library_name: ninfer
pipeline_tag: text-generation
tags: - ninfer
- gguf
- gemma4
- qat
- moe
- mtp
- speculative-decoding
- uncensored
- rtx-3090
Gemma4 26B-A4B QAT Uncensored (HauhauCS Balanced) with MTP, NInfer v3 artifact
A NInfer v3 artifact of HauhauCS's
Gemma4-26B-A4B-QAT-Uncensored-HauhauCS-Balanced-MTP
(Gemma 4 26B-A4B it QAT with refusals removed, the Balanced variant, Q4_K_M) together with the
drafter the release ships, for NInfer-all, the master branch of
iamwavecut/ninfer-all. Every matrix keeps the release's
ggml blocks byte for byte, each in its own format, so nothing is requantized; the engine multiplies
the blocks in place and drafts with the release's MTP head (--spec mtp). Stock NInfer builds
refuse this file: the Gemma 4 family and the gguf_* formats exist only in that line, mixed-format
Gemma GGUFs and the drafter from commit03bd4006
on. The family was developed and measured on the RTX 3090; the RTX 40 and RTX 50 builds compile
it, unmeasured.
It runs the text model (the release's vision projector is not included): one to eight concurrent
requests with prompt-prefix reuse, structured output, tool calls in Gemma's markup returned as
OpenAI tool calls, and reasoning in its own channel when a request turns thinking on.
What is inside
| component | representation |
|---|---|
| attention, 25 sliding-window and 5 full-attention blocks | Q, K and output Q4_K; V Q6_K (13 blocks) or Q4_K (12); a block whose Q, K and V differ in format keeps each its own |
| dense MLP | gate and up Q4_K (one parent), down Q8_0 (14 blocks) or Q5_0 (16) |
| 128 routed experts per block | [gate; up] Q4_K, down Q8_0 (14 blocks) or Q5_0 (16), expert-major |
| token table and tied head | Q6_K |
| norms, router, router scale, expert scales, layer scalars | FP32 |
drafter (mtp component) |
the release's mtp-gemma-4-26B-A4B-it.gguf (Q4_0 with its own token table; the release credits it to Unsloth's Gemma 4 GGUFs, and it differs from Unsloth's current QAT drafter file): Gemma 4's four-block assistant of width 1024 reading the model's own KV cache |
| tokenizer, chat template | google/gemma-4-26B-A4B-it at revision 4d7ae498: Google's canonical chat template, newer than the one inside the release's GGUF |
| generation defaults | the release card's recommendation: temperature 0.6, top-k 64, top-p 0.9, min-p 0.05; its repeat_penalty 1.1 has no NInfer equivalent (requests may set presence or frequency penalties) |
One .ninfer file of 17,048,711,936 bytes (15.88 GiB; SHA-256f198ebce11cfb925fbaa84f6831bd8759e3c570fe36d43b981e6c627e0f2be90), next to its conversion report,SHA256SUMS and NOTICE: 705 parameters in Q4_K (133 objects), Q8_0 (28), Q5_0 (32), Q4_0 (19,
the drafter), Q6_K (14) and FP32 (416).
Conversion command, from the repository's tree at 03bd4006, with the release at revisionf9093662 (Gemma4-26B-A4B-QAT-Uncensored-HauhauCS-Balanced-Q4_K_M.gguf, sha256 3c131334…;mtp-gemma-4-26B-A4B-it.gguf, sha256 62bd3af7…) and Google's text files at 4d7ae498:
mkdir -p hf
for f in config.json tokenizer.json tokenizer_config.json chat_template.jinja generation_config.json; do
curl -fsSL -o hf/$f https://huggingface.co/google/gemma-4-26B-A4B-it/resolve/main/$f
done
hf download HauhauCS/Gemma4-26B-A4B-QAT-Uncensored-HauhauCS-Balanced-MTP \
Gemma4-26B-A4B-QAT-Uncensored-HauhauCS-Balanced-Q4_K_M.gguf mtp-gemma-4-26B-A4B-it.gguf --local-dir .
# The release's sampling as the artifact's defaults.
jq '. + {do_sample: true, temperature: 0.6, top_k: 64, top_p: 0.9, min_p: 0.05}' \
hf/generation_config.json > generation_config.json
python3 -m tools.convert --model hf --recipe gemma4_gguf --components text,mtp \
--source gguf=Gemma4-26B-A4B-QAT-Uncensored-HauhauCS-Balanced-Q4_K_M.gguf \
--source mtp=mtp-gemma-4-26B-A4B-it.gguf --resource generation_config.json=generation_config.json \
--name Gemma4-26B-A4B-QAT-Uncensored-HauhauCS-Balanced-MTP \
--out Gemma4-26B-A4B-QAT-Uncensored-HauhauCS-Balanced-MTP-ninfer-v3.ninfer --device cpu
The conversion takes under a minute.
Running
Four concurrent requests over 40K-token contexts with two drafts a round, from the
Docker image (an NVIDIA driver of the CUDA 13
branch, 580 or newer, and the NVIDIA Container Toolkit; images from sha-03bd40060ed1 on read
this file):
hf download WaveCut/Gemma4-26B-A4B-QAT-Uncensored-HauhauCS-Balanced-MTP-NInfer-v3 \
Gemma4-26B-A4B-QAT-Uncensored-HauhauCS-Balanced-MTP-ninfer-v3.ninfer --local-dir models
docker run --rm --gpus all -p 8080:8080 --ulimit memlock=-1 \
-v "$PWD/models:/models" -v ninfer-cache:/cache \
ghcr.io/iamwavecut/ninfer-all serve /models/Gemma4-26B-A4B-QAT-Uncensored-HauhauCS-Balanced-MTP-ninfer-v3.ninfer \
--model-id gemma4-hauhau --max-context 40960 --max-concurrency 4 --spec mtp --draft-tokens 2
The server answers OpenAI and Anthropic requests on http://localhost:8080/v1. A build of master
(its README; CMAKE_CUDA_ARCHITECTURES is86, 89 or 120a, and 120a needs CUDA 13.1 or newer) takes the same flags:
ninfer-serve models/Gemma4-26B-A4B-QAT-Uncensored-HauhauCS-Balanced-MTP-ninfer-v3.ninfer \
--model-id gemma4-hauhau --max-context 40960 --max-concurrency 4 --spec mtp --draft-tokens 2
The weights take 15.6 GiB on the GPU (0.25 GiB more for the drafter with --spec mtp) and the
prompt workspace about 1.1 GiB; a sequence holds 250 MiB for the sliding windows with --spec mtp
(200 MiB without) and 10 KiB per token of context for the full-attention blocks. The release card
describes the model as reasoning before it answers; NInfer follows the chat template, where thinking
is off unless a request asks (chat_template_kwargs: {"enable_thinking": true}).
Speculative decoding. The drafter proposes tokens that the model verifies in one pass, keeping
a draft only while it equals the request's own sample: the output is the one plain decoding gives,
and the model's behaviour is unchanged by the drafter. A request decoding alone takes the
configured drafts, two or three at once one draft each, four or more decode plainly.
Quality
The conversion keeps every block of the release byte for byte; the same build's kernels are
checked against FP64 oracles and its assistant path against an FP64 transcription of transformers'
Gemma 4 drafting (see the repository's
Gemma 4 notes). Perplexity
over eleven 8,192-token windows of a NInfer source text: 21.79 (Google's QAT Q4_0 release on the same
windows: 24.79). Before publication this file served greedy chat requests (code, a
Russian answer, a 13,600-token prompt) plainly and with two drafts (acceptance 0.63), and the
CLI. Reasoning benchmarks were not run.
Speed
October 2026, one RTX 3090 (24 GB) at 350 W, CUDA 12.8, a build of master at 03bd4006: prefill
6,113 tok/s for a 2,048-token prompt and 5,623 for 10,000; decode 206.2 tok/s from an empty context
and 182.7 after 10,000 tokens. Ten chat requests (code, explanation, summary, QA, translation, math,
a story, a Russian answer, a 13,600-token source excerpt), 384 greedy tokens each, one at a time,
decode tok/s:
| all ten | 13,600-token prompt | |
|---|---|---|
| plain | 200.3 | 179.3 |
--draft-tokens 2 (acceptance 0.69) |
275.9 | 225.5 |
--draft-tokens 3 |
276.4 |
Credits and license
Uncensored weights, quantization and drafter GGUF: HauhauCS, whose release declares the Gemma license
(license: gemma), kept here; the drafter, per the release, comes from Unsloth's Gemma 4 GGUFs.
Base model, tokenizer and chat template: Google, Gemma 4, under the
Gemma 4 terms (Apache 2.0). Attribution notices
are collected in NOTICE. This repository only re-packs those weights into NInfer's
container.