← back to catalog · registered 2026-08-22 13:56

YuYu1015/Huihui-Gemma-4-E2B-it-abliterated-NVFP4

YuYu1015 Gemma 3.4B second-order
Your rig guess connected
? Why do I need an app?
Reading your rig…

This is a rough estimate. Install the free app - we'll show exact numbers.

Reading real hardware from your app right now. Numbers below are exact.

Below is the per-quantization compatibility for this model.

curl -H "Authorization: Bearer $ABL_KEY" \
     "https://abliteration.org/api/v1/models/YuYu1015%2FHuihui-Gemma-4-E2B-it-abliterated-NVFP4"
Response includes
  • classification m1
  • files 9
  • hub_downloads_all_time 484
  • author_summary 23 models
  • readme_text full
10 credits · hourly refresh · ~4 KB payload Get an API key →
Abliteration classifier · v1.0.0
M1
Primary method

Direct removal

No other method signals detected in this model.
Confidence
MEDIUM
Why this label 3 signals
Method inferred from partial signals - repository name, related files, or tag patterns. Producer identity not confirmed; label may sharpen or shift as we gather more evidence.
  • 'abliterated' in name/tags
  • is_gguf=0 (base model)
  • no specific method indicators - defaulting to M1 (most common)
Refusal direction extracted via
Extraction technique

Difference-of-means

Confidence
MEDIUM
Why we say so
primary_method=M1; difference-of-means is the reference extraction for M1/M3 (Arditi 2024)
Downloads · lifetime
484
27 last 30d - cooling
Likes
0
Model age
6mo ago
created 2026-04-11
Downloads over time
Now496→from265↑87%
253342431519265 on Apr 15496 on Oct 11AprMayJunJulAugSepOct
Apr 15 → Oct 11 · 65 snapshots · spans 179 days

Genealogy 0 direct forks

Full fork graph →

This model's place in the market. Above: what it was derived from. Below: the tree of everything derived from it.

Metadata

License
gemma
Languages
en zh
Tags
transformers safetensors gemma4 image-text-to-text nvfp4 4-bit quantized abliterated dgx-spark vllm modelopt text-generation

Related

Total size
7.27 GB
Files
9
Quantizations
1
Registered
2026-08-22 13:56
Last updated on HF
2026-04-12 18:09

Files by quantization

Auxiliary files 9 files 7.30 GB
model.safetensors 7.27 GB ca996bad download
tokenizer.json 30.7 MB 88e71407 download
config.json 14.1 KB 29945769 download
chat_template.jinja 11.6 KB afb1d517 download
hf_quant_config.json 7.95 KB 82c14032 download
README.md 5.63 KB e1bd1741 download
tokenizer_config.json 2.62 KB 59dd4b62 download
.gitattributes 1.53 KB 52373fe2 download
generation_config.json 203 B 92b5abfd download

README current version from Hugging Face


license: gemma
base_model:

  • huihui-ai/Huihui-gemma-4-E2B-it-abliterated
    base_model_relation: quantized
    library_name: transformers
    pipeline_tag: text-generation
    tags:
  • safetensors
  • gemma4
  • nvfp4
  • 4-bit
  • quantized
  • abliterated
  • dgx-spark
  • vllm
  • modelopt
    language:
  • en
  • zh

Huihui-gemma-4-E2B-it-abliterated-NVFP4

English | 繁體中文


English

NVFP4 quantization of huihui-ai/Huihui-gemma-4-E2B-it-abliterated, quantized using NVIDIA ModelOpt with NVFP4_MLP_ONLY strategy (only MLP layers quantized, attention preserved in higher precision).

Model Details

Item Value
Architecture Dense, Per-Layer Embeddings (PLE), ~2.3B effective parameters
Base model google/gemma-4-E2B-it
Fine-tuned by huihui-ai (abliteration)
Quantized by YuYu1015
Model size ~7.4 GB (NVFP4)
Context length Up to 128,000 tokens
Multimodal Vision + Audio supported

Quantization Details

Item Value
Method NVIDIA ModelOpt v0.42.0
Scheme NVFP4 (E2M1 + FP8 per-group scaling, group size 16)
Strategy NVFP4_MLP_ONLY — only MLP/FFN layers quantized, all attention layers preserved
Calibration dataset abisee/cnn_dailymail
Calibration samples 512
Hardware NVIDIA DGX Spark (GB10, 128GB unified memory)

Layers Preserved in Higher Precision

Layer Reason
self_attn.* (all layers) Attention layers preserved for accuracy (MLP_ONLY strategy)
lm_head Output head
vision_tower.* Vision encoder
audio_tower.* Audio encoder
multi_modal_projector.* Multimodal projection
embed_tokens Input embeddings

Serving with vLLM

vllm serve /path/to/model \
    --quantization modelopt \
    --served-model-name gemma-4-e2b \
    --trust-remote-code \
    --gpu-memory-utilization 0.90 \
    --max-model-len 32768 \
    --enable-prefix-caching \
    --enable-chunked-prefill \
    --language-model-only

DGX Spark (SM121) Compatibility Notes

  • NVFP4 on SM121 falls back to W4A16 (native W4A4 path not available, missing cvt.e2m1x2 instruction)
  • Use --quantization modelopt (not compressed-tensors)
  • --language-model-only skips vision/audio encoder profiling for text-only inference
  • Clear page cache before starting on UMA: sudo sh -c 'echo 3 > /proc/sys/vm/drop_caches'

Safety Warning

This model has safety filtering removed (abliterated) and may generate inappropriate content. Users are solely responsible for all consequences arising from its use.

Credits


繁體中文

huihui-ai/Huihui-gemma-4-E2B-it-abliterated 的 NVFP4 量化版本,使用 NVIDIA ModelOpt 的 NVFP4_MLP_ONLY 策略量化(僅量化 MLP 層,attention 保留高精度)。

模型資訊

項目 數值
架構 Dense,Per-Layer Embeddings (PLE),約 2.3B 有效參數
基礎模型 google/gemma-4-E2B-it
微調者 huihui-ai(abliteration)
量化者 YuYu1015
模型大小 ~7.4 GB(NVFP4)
Context 長度 最高 128,000 tokens
多模態 支援視覺 + 音訊

量化詳情

項目 數值
方法 NVIDIA ModelOpt v0.42.0
方案 NVFP4(E2M1 + FP8 逐群縮放,群組大小 16)
策略 NVFP4_MLP_ONLY — 僅量化 MLP/FFN 層,所有 attention 層保留高精度
校準資料集 abisee/cnn_dailymail
校準樣本數 512
量化硬體 NVIDIA DGX Spark(GB10, 128GB 統一記憶體)

保留高精度的層

層 原因
self_attn.*(所有層) Attention 層保留以確保精度(MLP_ONLY 策略)
lm_head 輸出頭
vision_tower.* 視覺編碼器
audio_tower.* 音訊編碼器
multi_modal_projector.* 多模態投影層
embed_tokens 輸入嵌入

DGX Spark (SM121) 相容性說明

  • NVFP4 在 SM121 上會退回 W4A16(原生 W4A4 路徑不可用,缺少 cvt.e2m1x2 指令)
  • 使用 --quantization modelopt(非 compressed-tensors)
  • --language-model-only 跳過視覺/音訊編碼器 profiling,加速純文字推理
  • UMA 架構啟動前請先清除 page cache:sudo sh -c 'echo 3 > /proc/sys/vm/drop_caches'

安全警告

此模型已移除安全過濾機制(abliterated),可能產生不當內容。使用者須自行承擔所有風險與法律責任。

致謝

README history 3 versions

The author's README evolved over time. Click a version to see its content at that point.

  1. 2026-04-12Update README.mdfe2e5075.6 KB
    Loading...
  2. 2026-04-11Update README.mda7ad8d06.7 KB
    Loading...
  3. 2026-04-11Create README.mddf3d93a6.7 KB
    Loading...
Catalog is the map. Apps are the tools.

Run models on your own machine, not in the cloud.

Every model page has an "Open in Abliteration" button that hands the model directly to the first-party desktop client, at the quantization your rig can actually run. No API keys, no subscription, no prompt leakage.

Open in Abliteration