← back to catalog · registered 2026-08-22 13:56

AEON-7/Gemma-4-E4B-DECKARD-HERETIC-Uncensored-NVFP4

AEON-7 Gemma 4.0B multimodal second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/AEON-7%2FGemma-4-E4B-DECKARD-HERETIC-Uncensored-NVFP4"
Response includes
  • classification m3
  • files 10
  • benchmarks 11 entries
  • hub_downloads_all_time 1,035
  • author_summary 32 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M3
Primary method

Layer-wise ablation

Applied on top of direct removal inherited from the base model.
Confidence
HIGH
Inherited from base model
Why this label 2 signals
Producer identity confirmed by naming conventions, tags or the model card. This label is very unlikely to change.
  • 'heretic' in model name (Heretic-produced)
  • Heretic uses layer-wise optimization (M3) with underlying direction removal (M1)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
HIGH
Why we say so
name contains 'heretic'; Heretic default extraction is difference-of-means (Arditi 2024)
Downloads · lifetime
1K
185 last 30d - stable
Likes
1
Model age
6mo ago
created 2026-04-13
Downloads over time
Now1.2K→from110↑963%
574638691.3K110 on Apr 151.2K on Oct 11AprMayJunJulAugSepOct
Apr 15 → Oct 11 · 65 snapshots · spans 179 days

Benchmarks

Benchmark Score Source
Entertainment 2.8 UGI
Hazardous 4.7 UGI
Natural Intelligence 34.42 UGI
Political lean -22.2% UGI
Sensitive-Info 32.29 UGI
SocPol 2.6 UGI
UGI 52.36 UGI
Willingness (10) 9.2 UGI
W10-Adherence 8.5 UGI
W10-Direct 10 UGI
Writing 38.68 UGI

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
gemma
Languages
en
Tags
transformers safetensors gemma4 image-text-to-text 4-bit aarch64 abliterated aeon aeon-7 agentic arm64 awq

Related

Total size
9.47 GB
Files
10
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-10-06 01:36

Files by quantization

Auxiliary files 10 files 9.50 GB
model.safetensors 9.47 GB 98d0341b download
tokenizer.json 30.7 MB cc8d3a0c download
chat_template.jinja 11.7 KB cbffd77e download
README.md 8.75 KB 8a0e43cb download
processor_config.json 6.86 KB c9cb71ba download
config.json 6.40 KB 10e30cf8 download
tokenizer_config.json 2.62 KB f07b8ede download
.gitattributes 1.53 KB 52373fe2 download
hf_quant_config.json 475 B e8d3d1dd download
generation_config.json 179 B e897d9e0 download

README current version from Hugging Face


license: gemma
library_name: transformers
pipeline_tag: text-generation
base_model: DavidAU/gemma-4-31B-it-The-DECKARD-HERETIC-UNCENSORED-Thinking
tags:

  • 4-bit
  • aarch64
  • abliterated
  • aeon
  • aeon-7
  • agentic
  • arm64
  • awq
  • blackwell
  • chat
  • chunked-prefill
  • coding
  • conversational
  • deckard
  • dgx-spark
  • drafter
  • e4b
  • eagle
  • english
  • flashinfer-cutlass
  • fp4
  • fp8-kv-cache
  • function-calling
  • gb10
  • gemma
  • gemma-4
  • gemma4
  • google
  • gpu
  • grace-blackwell
  • heretic
  • instruct
  • long-context
  • modelopt
  • multimodal
  • native-fp4
  • nvfp4
  • nvidia
  • openai-api
  • openai-compatible
  • prefix-caching
  • production-ready
  • quantized
  • reasoning
  • refusal-removed
  • safetensors
  • sm_121a
  • speculative-decoding
  • text-generation
  • thinking
  • tool-calling
  • uncensored
  • unfiltered
  • vision
  • vision-language
  • vllm
  • weight-quantization
    model_type: gemma4
    quantization: nvfp4
    language:
  • en

Gemma 4 E4B DECKARD HERETIC Uncensored NVFP4

EAGLE speculative decoding drafter for Gemma 4 31B DECKARD HERETIC Uncensored NVFP4.

A 42-layer E4B (EAGLE for Blackwell) model quantized to NVFP4 AWQ using NVIDIA ModelOpt 0.42.0. Designed for EAGLE-based speculative decoding on NVIDIA DGX Spark (GB10, SM 12.1) and other Blackwell GPUs.

Model Details

Property Value
Architecture Gemma 4 (E4B EAGLE Drafter)
Target Model AEON-7/Gemma-4-31B-it-DECKARD-HERETIC-Uncensored-NVFP4
Layers 42 (35 sliding-window + 7 full-attention)
Hidden Size 2560
Attention Heads 8 (2 KV heads), head_dim=256, global_head_dim=512
Sliding Window 512 tokens
Max Context 131,072 tokens
Quantization NVFP4 AWQ (ModelOpt 0.42.0)
Model Size 9.6 GB
Vocabulary 262,144 tokens

Performance (DGX Spark)

Benchmarked on NVIDIA DGX Spark (GB10, SM 12.1, 128 GB unified memory) with 31B DECKARD AWQ_FULL target + this E4B drafter. 5 speculative tokens, 300 max tokens per request.

Concurrent Aggregate tok/s Per-Request tok/s Avg Latency (300 tok)
1 7.6 8.9 39.4s
2 21.7 10.8 27.7s
4 42.7 10.7 28.1s

Zero errors across all test runs. Throughput scales linearly with concurrency.

Quick Start

1. Download both models

pip install -U huggingface-hub

# Target model (31B)
huggingface-cli download AEON-7/Gemma-4-31B-it-DECKARD-HERETIC-Uncensored-NVFP4 \
  --local-dir ~/models/deckard-31b

# This drafter model (E4B)
huggingface-cli download AEON-7/Gemma-4-E4B-DECKARD-HERETIC-Uncensored-NVFP4 \
  --local-dir ~/models/e4b-drafter

2. Get the patched vLLM files

Three patches are required for Gemma 4 speculative decoding. Download from the GitHub repo:

for f in eagle_patched.py serving_chat_patched.py modelopt_patched.py; do
  curl -LO https://raw.githubusercontent.com/AEON-7/Gemma-4-31B-DECKARD-HERETIC-Uncensored-NVFP4/main/$f
done

3. Launch with Docker Compose

services:
  vllm:
    image: ghcr.io/aeon-7/vllm-spark-gemma4-nvfp4-awq:latest
    container_name: vllm-deckard-31b-spec
    restart: unless-stopped
    network_mode: host
    volumes:
      - ~/models/deckard-31b:/models/deckard
      - ~/models/e4b-drafter:/models/e4b-drafter
      - ./modelopt_patched.py:/usr/local/lib/python3.12/dist-packages/vllm/model_executor/layers/quantization/modelopt.py
      - ./serving_chat_patched.py:/usr/local/lib/python3.12/dist-packages/vllm/entrypoints/openai/chat_completion/serving.py
      - ./eagle_patched.py:/usr/local/lib/python3.12/dist-packages/vllm/v1/spec_decode/eagle.py
    environment:
      - VLLM_TEST_FORCE_FP8_MARLIN=1
      - VLLM_MARLIN_USE_ATOMIC_ADD=1
      - VLLM_ALLOW_LONG_MAX_MODEL_LEN=1
      - TORCH_MATMUL_PRECISION=high
      - PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True
    command:
      - bash
      - -c
      - |
        exec vllm serve /models/deckard \
          --served-model-name deckard-31b \
          --quantization modelopt \
          --dtype auto \
          --kv-cache-dtype fp8 \
          --tensor-parallel-size 1 \
          --max-model-len 131072 \
          --max-num-seqs 4 \
          --gpu-memory-utilization 0.65 \
          --trust-remote-code \
          --host 0.0.0.0 --port 8000 \
          --enable-chunked-prefill \
          --enable-prefix-caching \
          --enable-auto-tool-choice \
          --tool-call-parser gemma4 \
          --reasoning-parser gemma4 \
          --speculative-config '{"method":"draft_model","model":"/models/e4b-drafter","num_speculative_tokens":5,"quantization":"modelopt"}'
    ipc: host
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: all
              capabilities: [gpu]

On the DGX Spark's unified memory keep --gpu-memory-utilization at 0.6-0.7; above ~0.8 the shared CPU+GPU pool page-thrashes. With EAGLE speculative decoding the verify buffers are not counted by the fraction, so leave extra headroom (0.65 here). Discrete-VRAM GPUs can run higher.

4. Test

curl http://localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deckard-31b",
    "messages": [{"role": "user", "content": "Explain quantum entanglement."}],
    "max_tokens": 200
  }'

Required vLLM Patches

Three patches to vLLM 0.19.1 are required for speculative decoding with Gemma 4. All are available in the target model GitHub repo.

Patch What it fixes
eagle_patched.py Removes multimodal spec decode guard, adds Gemma4 model whitelist, supports multi-group KV cache (heterogeneous head_dim=256/512)
serving_chat_patched.py Fixes non-streaming reasoning parser — <|channel> tokens stripped by skip_special_tokens=True
modelopt_patched.py NVFP4_AWQ quant_algo support, AWQ pre_quant_scale handling, FP8 NaN scrubbing

Heterogeneous Attention

This E4B drafter mirrors the Gemma 4 heterogeneous attention design:

  • 35 sliding-window layers — head_dim=256, window of 512 tokens, default RoPE (theta=10000)
  • 7 full-attention layers — head_dim=512, global attention, proportional RoPE (theta=1M, partial_rotary_factor=0.25)

This creates two distinct KV cache groups, handled by the eagle_patched.py multi-group KV cache fix.

Related Models

Model Type Size Link
Gemma 4 31B DECKARD AWQ_FULL (target) Dense NVFP4 20.5 GB HuggingFace | GitHub
Gemma 4 31B DECKARD SVDQuant Dense NVFP4 20.9 GB HuggingFace
SuperGemma4 26B MoE MoE NVFP4 15.3 GB HuggingFace
vLLM AWQ Container Docker — GHCR

License

This model inherits the Gemma license from Google.


☕ Support the work

If this release has been useful, tips are deeply appreciated — they go directly toward more compute, more models, and more open releases.

₿ Bitcoin (BTC)
QR
bc1q09xmzn00q4z3c5raene0f3pzn9d9pvawfm0py4
Ξ Ethereum (ETH)
QR
0x1512667F6D61454ad531d2E45C0a5d1fd82D0500
◎ Solana (SOL)
QR
DgQsjHdAnT5PNLQTNpJdpLS3tYGpVcsHQCkpoiAKsw8t
ⓜ Monero (XMR)
QR
836XrSKw4R76vNi3QPJ5Fa9ugcyvE2cWmKSPv3AhpTNNKvqP8v5ba9JRL4Vh7UnFNjDz3E2GXZDVVenu3rkZaNdUFhjAvgd

Ethereum L2s (Base, Arbitrum, Optimism, Polygon, etc.) and EVM-compatible tokens can be sent to the same Ethereum address.

README history 6 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-10-06Add Patreon support sectioncd6d3039.3 KB
    Loading...
  2. 2026-07-15Recipe: gpu-util 0.6-0.7 on DGX Spark unified memory (>~0.8 thrashes the shar...2d49bc18.8 KB
    Loading...
  3. 2026-06-21tags: expand to maximally-searchable set (+28 tags, union with existing)5287c078.5 KB
    Loading...
  4. 2026-05-31Tip jar: single left-aligned QR column (fix narrow-viewport clipping)1d295398.2 KB
    Loading...
  5. 2026-05-01Add tip jar block (BTC/ETH/SOL/XMR with QR codes)42a6da18.3 KB
    Loading...
  6. 2026-04-13Add model card with speculative decoding documentation50ee3686.7 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration